Skip to content
iGaming Times

Independent industry intelligence in your inbox. We will email you a link to confirm your subscription, and every newsletter carries a one-click unsubscribe link.

Lesson 5 of 6 · 18 min

Experimentation Under Variance and Constraint

Controlled experiment design, why gambling variance makes sample sizes impossible and the remedies, interference and contamination, what may ethically and legally be tested, reading a result properly, and running a programme.

Fact-checked 23 September 2026 by iGaming Times editorial team · 8 sources

In this lesson

  • Design an experiment with the right unit, population, metric, duration and sample size
  • Reduce variance with metric choice, capping, stratification and covariate adjustment
  • Handle interference with cluster, holdout and switchback designs
  • State the constraints on testing protective features and promotions
  • Check guardrails, segments, novelty, sample ratio and balance before shipping a result

Why experiments, and why they are hard here

A model tells you what is likely; only an experiment tells you what a change caused. Operators that run disciplined experimentation programmes make better product, marketing and retention decisions than those that ship changes and watch the dashboard, because the dashboard cannot separate the change from the season, the sporting calendar, a competitor's promotion and luck. This lesson covers how to run experiments in an environment with some specific difficulties: extreme variance in the outcome metric, heavy-tailed customers, a product whose results are random by design, and regulatory limits on what may be tested.

The basic design

A controlled experiment assigns customers at random to a treatment (the change) or a control (the status quo), exposes them, and compares an outcome metric between the groups. Random assignment makes the groups comparable, so the difference in outcome is attributable to the change.

The design decisions:

Unit of randomisation. Almost always the customer, not the session or the bet, because customers are exposed repeatedly and their outcomes are correlated. Randomise by customer identifier with a stable hash so that a customer sees the same variant every time.

Population. Which customers are eligible: new registrations, active customers in a market, customers with a particular product history. The narrower the population, the smaller the sample and the longer the test.

Metric. The primary metric the experiment is designed to move, chosen in advance, plus guardrail metrics that must not degrade (complaints, harm scores, withdrawal times) and secondary metrics for interpretation.

Duration. Long enough to cover the customer's natural cycle (a weekly bettor needs weeks) and the sporting calendar's variation, and fixed in advance to avoid stopping when the result looks good.

Sample size. Computed before the test from the metric's variance, the minimum effect worth detecting and the required power. In gambling this calculation delivers bad news, and the next section is about why.

The variance problem

Gambling revenue per customer is extremely skewed: most customers contribute little, a few contribute enormously, and the outcome of any customer's play over a test period includes the randomness of the games themselves. The variance of net revenue per customer is huge, and the sample size needed to detect a modest effect on it can exceed the customer base.

The remedies, in rough order of value:

Choose a lower-variance metric. Deposits are usually less variable than net revenue, because they are not set directly by game outcomes, although a winning or losing run still changes what a customer deposits next. Activity metrics (sessions, active days, retention) are lower still. Where the business question is about revenue, test on a proxy that predicts it and confirm on revenue over a longer window.

Winsorise or cap. Cap each customer's contribution at a high percentile so a single whale cannot swing the result. Microsoft reported that at Bing, capping revenue per user at $10 a week cut the metric's skewness from 18 to about 5 and let the same sample detect a change 30% smaller. Report both capped and uncapped, and be honest that the capped result excludes the tail.

Stratify and use covariates. Assign within strata (by historical value, product, market) and adjust the analysis for pre-period behaviour (CUPED and similar variance-reduction techniques, which use each customer's own history to reduce noise). The Microsoft team that introduced CUPED in 2013 reported reducing variance by about 50% on Bing, the same statistical power with half the users or half the duration. The gain depends on how well past behaviour predicts the outcome, so it is large for established customers and little or nothing for new registrations, who have no pre-period history to draw on.

Sequential and Bayesian designs. Sequential methods, such as always-valid p-values, allow monitoring during the test without inflating false positives, which repeatedly checking an ordinary p-value does; Bayesian designs report probabilities of improvement rather than a binary significance verdict. Both are more useful to a product team, and more honest about uncertainty, provided the stopping rules are set before the test.

Accept longer tests. Sometimes there is no substitute for time.

Interference and contamination

Customers talk, and gambling customers talk about promotions. A bonus offered to the treatment group leaks to the control through forums and social media; a price boost visible on the site to some customers is screenshotted by others. Marketplace-style products (poker, exchanges) have direct interference: what one group does changes the experience of the other. The mitigations are cluster randomisation (by market or by time period), holdout markets, and switchback designs (alternating treatment and control over time for everyone), each with its own analysis; a switchback, for instance, has to be designed around carryover, the time a treatment keeps affecting the outcome after it is switched off.

What may be tested

Regulation constrains experimentation in ways product teams from other industries do not expect. Tests that vary responsible gambling messaging, limit prompts, or the friction of protective tools are ethically and often legally constrained: a control group deprived of a protective feature is a group the operator has chosen not to protect, and where the feature is a licence requirement, such as the British rule that operators must prevent marketing and the take-up of new bonus offers where strong indicators of harm have been identified, withholding it is a breach. Testing protection is not itself the problem: the same British code requires operators to take all reasonable steps to evaluate the effectiveness of their approach, for example by trialling and measuring impact. Tests that vary bonus terms have to stay within the promotional rules in every market the test runs: in Britain, for example, wagering requirements have been capped at ten times, and offers mixing betting with casino or other products banned, since 19 January 2026. Tests on pricing have to respect fairness obligations. And a test whose treatment increases harm scores must stop, whatever its effect on revenue.

The workable position: protective features are tested only in the direction of more protection (does a stronger prompt work better than the standard one, never versus none), and every experiment carries the harm score as a guardrail with a pre-agreed stopping rule.

Reading a result

A significant result on the primary metric is the beginning of the analysis, not the end. Before shipping: check the guardrails; check the effect across segments (an average gain that is a loss for high-value customers, or a gain concentrated in a group the harm model flags, is a different decision); check for novelty and primacy effects by looking at the effect over time within the test; check for sample ratio mismatch (the groups are not the size randomisation should have produced, a symptom that something in assignment, exposure, logging or analysis is broken, and one Microsoft found in about 6% of its experiments); and check that the pre-period behaviour of the groups was balanced. A result that survives all of that is a result.

The programme, not the test

A single experiment is a fact. An experimentation programme is a culture: a platform that makes randomisation and analysis routine, a register of every test with its hypothesis, design and result, a review that stops bad tests and challenges good ones, and a habit of testing the things the organisation believes rather than only the things it doubts. Operators with such programmes find that most changes do not move the metric they were meant to (at Microsoft, only about one third of ideas tested improved the metrics they were designed to improve), that some move it the wrong way, and that the ones that work are rarely the ones the loudest person predicted. That is the value.

The last lesson takes the models and the experiments and asks who is responsible for them.

Key terms

CUPED
Controlled-experiment Using Pre-Experiment Data: a variance-reduction technique, introduced by Microsoft in 2013, that adjusts experiment outcomes using each customer’s pre-period behaviour.
Winsorising
Capping each customer’s contribution at a high percentile to limit the effect of outliers.
Switchback design
An experiment alternating treatment and control over time for all customers, used where groups interfere; its design must allow for carryover between periods.
Guardrail metric
A metric that must not degrade during an experiment, with a pre-agreed stopping rule.
Sample ratio mismatch
Groups whose sizes differ from what randomisation should produce, a symptom that assignment, exposure, logging or analysis is broken and that the result cannot be trusted.

Key takeaways

  • The dashboard cannot separate the change from the season, the calendar, a competitor and luck; only an experiment can.
  • Net revenue per customer has enormous variance; test on lower-variance proxies and confirm on revenue over time.
  • Covariate adjustment using each customer’s own history can cut required samples substantially, by about half in Microsoft’s published CUPED results, though not for new customers with no history.
  • Protective features are tested only in the direction of more protection.
  • Most changes do not move the metric they were meant to; that is the value of the programme.

Sources

The legislation, regulator material and research this lesson was checked against.

  1. Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data (CUPED), WSDM 2013, Deng, Xu, Kohavi and Walker, Microsoft, accessed 2026-09-23
  2. Seven Rules of Thumb for Web Site Experimenters, KDD 2014, Kohavi, Deng, Longbotham and Xu, accessed 2026-09-23
  3. Online Controlled Experiments at Large Scale, KDD 2013, Kohavi, Deng, Frasca, Walker, Xu and Pohlmann, Microsoft, accessed 2026-09-23
  4. Diagnosing Sample Ratio Mismatch in Online Controlled Experiments: A Taxonomy and Rules of Thumb for Practitioners, KDD 2019, Fabijan et al., Microsoft, Booking.com and Outreach.io, accessed 2026-09-23
  5. Always Valid Inference: Bringing Sequential Analysis to A/B Testing, Johari, Pekelis and Walsh (arXiv), accessed 2026-09-23
  6. Design and Analysis of Switchback Experiments, Bojinov, Simchi-Levi and Zhao (arXiv), accessed 2026-09-23
  7. Licence Conditions and Codes of Practice, SR code 3.4.3 Remote customer interaction and SR code 5.1.1 Rewards and bonuses, Gambling Commission, accessed 2026-09-23
  8. Gambling promotions to be safer and simpler, Gambling Commission, accessed 2026-09-23

Check your understanding

3 questions · answer them all, then check.

  1. 1. An experiment on a new bonus shows a significant revenue gain. Before shipping, the first check is:

  2. 2. Why randomise by customer rather than by bet?

  3. 3. A team wants to test removing the deposit-limit prompt at registration to measure its effect on conversion. This is:

Sign in to track your progress through the course.

Cookie Preferences

Choose which cookies you want to accept. Essential cookies are required for the website to function properly.

Required

Necessary for the website to function. Cannot be disabled.

Help us understand how visitors interact with our website.

Used to deliver relevant advertisements and track ad performance.

Remember your preferences and settings for a better experience.