Scale By The Bay 2018: Robert J Neal, Paul Cho, Scaling Bayesian Experimentation
In the past few years companies industry wide have noted the limitations of traditional null hypothesis significance testing (NHST) for online experimentation. In particular, statistical problems like multiple comparisons and peeking have been difficult to solve while still being able to make fast business and product decisions. Bayesian methods provide an alternative to overcome these problems, but are often avoided because of worries about their complexity and computational intensity. We will talk about three challenges with Bayesian statistics for experimentation and how big data, tools like Spark, and a little statistical ingenuity can help us address them. The three challenges we will discuss are (1) coming up with priors for experimentation in a world of big data, (2) building a fast Bayesian computation pipeline that is generalizable to all of the metrics your organization cares about, and (3) overcoming computational inefficiencies when using these statistical methods in a real-time experimentation environment. To accomplish (2) we use bootstrapping and for (3) we will talk about some of the challenges and solutions to making it computationally efficient. In the past few years Internet-based companies have noted the limitations of traditional null hypothesis significance testing (NHST) for large-scale, online experimentation. In particular, statistical problems like multiple comparisons and peeking have been difficult to solve. Bayesian methods provide an alternative to overcome these problems, but are often avoided because of worries about their complexity…
Connections
2 relationships