Why experiments, and why they are hard here
A model tells you what is likely; only an experiment tells you what a change caused. Operators that run disciplined experimentation programmes make better product, marketing and retention decisions than those that ship changes and watch the dashboard, because the dashboard cannot separate the change from the season, the sporting calendar, a competitor's promotion and luck. This lesson covers how to run experiments in an environment with some specific difficulties: extreme variance in the outcome metric, heavy-tailed customers, a product whose results are random by design, and regulatory limits on what may be tested.
The basic design
A controlled experiment assigns customers at random to a treatment (the change) or a control (the status quo), exposes them, and compares an outcome metric between the groups. Random assignment makes the groups comparable, so the difference in outcome is attributable to the change.
The design decisions:
Unit of randomisation. Almost always the customer, not the session or the bet, because customers are exposed repeatedly and their outcomes are correlated. Randomise by customer identifier with a stable hash so that a customer sees the same variant every time.
Population. Which customers are eligible: new registrations, active customers in a market, customers with a particular product history. The narrower the population, the smaller the sample and the longer the test.
Metric. The primary metric the experiment is designed to move, chosen in advance, plus guardrail metrics that must not degrade (complaints, harm scores, withdrawal times) and secondary metrics for interpretation.
Duration. Long enough to cover the customer's natural cycle (a weekly bettor needs weeks) and the sporting calendar's variation, and fixed in advance to avoid stopping when the result looks good.
Sample size. Computed before the test from the metric's variance, the minimum effect worth detecting and the required power. In gambling this calculation delivers bad news, and the next section is about why.
The variance problem
Gambling revenue per customer is extremely skewed: most customers contribute little, a few contribute enormously, and the outcome of any customer's play over a test period includes the randomness of the games themselves. The variance of net revenue per customer is huge, and the sample size needed to detect a modest effect on it can exceed the customer base.
The remedies, in rough order of value:
Choose a lower-variance metric. Deposits are usually less variable than net revenue, because they are not set directly by game outcomes, although a winning or losing run still changes what a customer deposits next. Activity metrics (sessions, active days, retention) are lower still. Where the business question is about revenue, test on a proxy that predicts it and confirm on revenue over a longer window.
Winsorise or cap. Cap each customer's contribution at a high percentile so a single whale cannot swing the result. Microsoft reported that at Bing, capping revenue per user at $10 a week cut the metric's skewness from 18 to about 5 and let the same sample detect a change 30% smaller. Report both capped and uncapped, and be honest that the capped result excludes the tail.
Stratify and use covariates. Assign within strata (by historical value, product, market) and adjust the analysis for pre-period behaviour (CUPED and similar variance-reduction techniques, which use each customer's own history to reduce noise). The Microsoft team that introduced CUPED in 2013 reported reducing variance by about 50% on Bing, the same statistical power with half the users or half the duration. The gain depends on how well past behaviour predicts the outcome, so it is large for established customers and little or nothing for new registrations, who have no pre-period history to draw on.
Sequential and Bayesian designs. Sequential methods, such as always-valid p-values, allow monitoring during the test without inflating false positives, which repeatedly checking an ordinary p-value does; Bayesian designs report probabilities of improvement rather than a binary significance verdict. Both are more useful to a product team, and more honest about uncertainty, provided the stopping rules are set before the test.
Accept longer tests. Sometimes there is no substitute for time.
Interference and contamination
Customers talk, and gambling customers talk about promotions. A bonus offered to the treatment group leaks to the control through forums and social media; a price boost visible on the site to some customers is screenshotted by others. Marketplace-style products (poker, exchanges) have direct interference: what one group does changes the experience of the other. The mitigations are cluster randomisation (by market or by time period), holdout markets, and switchback designs (alternating treatment and control over time for everyone), each with its own analysis; a switchback, for instance, has to be designed around carryover, the time a treatment keeps affecting the outcome after it is switched off.
What may be tested
Regulation constrains experimentation in ways product teams from other industries do not expect. Tests that vary responsible gambling messaging, limit prompts, or the friction of protective tools are ethically and often legally constrained: a control group deprived of a protective feature is a group the operator has chosen not to protect, and where the feature is a licence requirement, such as the British rule that operators must prevent marketing and the take-up of new bonus offers where strong indicators of harm have been identified, withholding it is a breach. Testing protection is not itself the problem: the same British code requires operators to take all reasonable steps to evaluate the effectiveness of their approach, for example by trialling and measuring impact. Tests that vary bonus terms have to stay within the promotional rules in every market the test runs: in Britain, for example, wagering requirements have been capped at ten times, and offers mixing betting with casino or other products banned, since 19 January 2026. Tests on pricing have to respect fairness obligations. And a test whose treatment increases harm scores must stop, whatever its effect on revenue.
The workable position: protective features are tested only in the direction of more protection (does a stronger prompt work better than the standard one, never versus none), and every experiment carries the harm score as a guardrail with a pre-agreed stopping rule.
Reading a result
A significant result on the primary metric is the beginning of the analysis, not the end. Before shipping: check the guardrails; check the effect across segments (an average gain that is a loss for high-value customers, or a gain concentrated in a group the harm model flags, is a different decision); check for novelty and primacy effects by looking at the effect over time within the test; check for sample ratio mismatch (the groups are not the size randomisation should have produced, a symptom that something in assignment, exposure, logging or analysis is broken, and one Microsoft found in about 6% of its experiments); and check that the pre-period behaviour of the groups was balanced. A result that survives all of that is a result.
The programme, not the test
A single experiment is a fact. An experimentation programme is a culture: a platform that makes randomisation and analysis routine, a register of every test with its hypothesis, design and result, a review that stops bad tests and challenges good ones, and a habit of testing the things the organisation believes rather than only the things it doubts. Operators with such programmes find that most changes do not move the metric they were meant to (at Microsoft, only about one third of ideas tested improved the metrics they were designed to improve), that some move it the wrong way, and that the ones that work are rarely the ones the loudest person predicted. That is the value.
The last lesson takes the models and the experiments and asks who is responsible for them.