Skip to content
iGaming Times

Independent industry intelligence in your inbox. We will email you a link to confirm your subscription, and every newsletter carries a one-click unsubscribe link.

Lesson 6 of 7 · 17 min

Product Measurement and Experimentation

Choosing metrics that survive contact with a skewed distribution, running tests that answer questions, and knowing what not to measure.

Fact-checked 23 September 2026 by iGaming Times editorial team · 9 sources

In this lesson

  • Select product metrics that reflect value rather than activity
  • Design experiments accounting for the specific statistical properties of gambling data
  • Recognise when a metric improvement reflects a real gain and when it reflects displacement
  • Apply harm-relevant evaluation alongside commercial evaluation

The measurement problem in one sentence

The outcomes that matter in gambling product take months to observe, and product teams make decisions weekly.

Everything difficult about measurement in this sector follows from that gap. Retention, lifetime contribution and whether a customer's relationship with the operator ends well are the things worth optimising, and none of them is available at the point a decision is made. So teams use proxies, and proxies have a systematic bias.

The bias in proxies

Proxy metrics measure immediate response: conversion, click-through, session length, deposits per session. They are available quickly, they respond to changes, and they favour a particular kind of change.

A design that produces immediate engagement will improve them. A design that produces durable value may not, or may even worsen them in the short term. And a design that increases immediate engagement at the cost of long-term retention will show as a clear success for as long as anyone looks.

This is not a hypothetical. Several product patterns in this sector improve proximate metrics and harm the relationship: aggressive promotional interruption, friction removed from repeated deposits, discovery narrowed to maximise immediate play, and communications increased in frequency. Each converts well and each is measurable as a win. Britain's technical standards already restrict part of this territory: gambling products must not actively encourage customers to chase their losses, increase their stake or increase the amount they have decided to gamble, and the Commission's guidance on that requirement says funds taken into a product should not be topped up without the customer choosing to do so on each occasion. Nor is the divergence unique to gambling. Microsoft reported that an experiment bug which showed Bing users very poor search results raised distinct queries per user by more than 10% and revenue per user by more than 30%, a short-term win that would have eroded long-term value.

The responses are practical rather than clever.

Hold changes for delayed evaluation. Where a change can be assessed again at ninety days against retention rather than only at two weeks against conversion, do so, and treat the second assessment as the real one.

Validate proxies against outcomes. Establish which proximate metrics actually correlate with downstream contribution for this operator's customers, since the assumption that conversion predicts value is testable and is sometimes wrong.

Use paired metrics. Any conversion measure should be paired with a quality measure, so that increased volume of lower value is visible rather than counted as improvement.

Treat immediate improvement with suspicion when the mechanism is unclear. A change that improved a number without anyone being able to explain why usually has an explanation nobody has looked for.

Skew and what it does to tests

The statistical point made in the Operations Strategy course applies with particular force to product experimentation.

Revenue in gambling is concentrated in a small minority of customers. In a GambleAware-commissioned study of nearly 140,000 British online gambling accounts active between July 2018 and June 2019, the top 10% of gamblers by amount staked delivered 79% of operator revenue. Random allocation does not reliably balance that concentration between test groups. A single customer of unusual value landing in one group can produce an apparent effect entirely unrelated to the change being tested.

The practical disciplines are specific.

Check balance before interpreting. Compare groups on pre-test value, not only on size. The same pre-test data can also reduce noise: Microsoft's CUPED method, which adjusts metrics using each customer's pre-experiment behaviour, cut variance by about 50% on Bing, equivalent to running a test with half the users or for half the duration.

Use trimmed or median measures alongside means. If an effect appears in the mean and vanishes in the median, it is being carried by a small number of customers. Capping extreme values is a related technique: when Microsoft capped revenue per user at $10 a week, the metric's skewness fell from 18 to 5.3 and the same sample could detect a change 30% smaller.

Remove the largest few and re-run. If the effect does not survive, it was not an effect.

Analyse the high-value tail separately. These customers behave differently and a change may affect them differently, which is worth knowing rather than averaging away.

Require larger samples and longer durations than conventional guidance suggests, because gambling revenue is heavily skewed. Standard sample size formulas assume the average is normally distributed, and skewed metrics break that assumption: one published rule of thumb puts the minimum at 355 times the square of the skewness coefficient for each variant, which for a Bing revenue metric with a skewness of about 18 meant 114,000 users.

Do not stop early on a favourable result. This is a common error in commercial experimentation. Checking a conventional test continuously and stopping when it first looks significant makes its p-values unreliable: research on A/B testing found that the false positive rate can easily increase fivefold, so effects found this way often fail to replicate. Fix the duration in advance, or use a sequential method designed for continuous monitoring.

Displacement

A common false positive in product measurement is a change that captures activity rather than creating it.

A new discovery rail shows strong engagement. Customers clicked it, played the games it surfaced, and generated revenue. The rail appears successful. The question nobody asks is whether those customers would have played something else through another route, in which case the rail has redistributed activity rather than added any.

The same applies to cross-vertical bridges, promotional surfaces, notification campaigns and recommendation changes. Each can show excellent local metrics while total activity is unchanged.

Detecting displacement requires measuring at the level where it would show: total session activity, total revenue, total deposits, rather than the performance of the new surface. A rail that performs well while total play is flat is a rail that moved existing behaviour.

This matters because displacement-driven wins accumulate into a product full of surfaces that each demonstrably work and that collectively achieved nothing, while consuming screen space and engineering capacity that could have gone elsewhere.

Guardrails and harm-relevant evaluation

The evaluation framework in this sector needs a dimension that most product measurement does not have.

Standard practice defines a primary metric and a set of guardrails, which are measures that guard against the primary metric giving a wrong signal by checking that other important dimensions are not moving the wrong way while it improves. In most industries guardrails cover things like page performance, error rates, revenue and customer satisfaction.

Here they should also cover harm. The measures below overlap with the indicators British remote operators must already monitor under Social Responsibility Code Provision 3.4.3, which include customer spend, patterns of spend, time spent gambling and use of gambling management tools, and whose guidance lists increasing session length and escalating deposit levels among the signs of harm.

Spend distribution. Did the improvement come from broad modest increases or from a small group whose spend rose sharply? These look identical in aggregate and are entirely different findings.

Effect on customers with risk indicators. Did spend among customers displaying markers of harm increase disproportionately? If so, the change should not ship regardless of its aggregate performance.

Session length distribution. Did the change extend the tail of very long sessions?

Deposit frequency escalation. Did the change increase the rate of repeated deposits within a session?

Tool usage. Did the change reduce the use of limit-setting or other protective tools, which would indicate they became harder to find or less prominent. In Britain that is also a compliance question, since financial limit facilities must be provided via a direct link on the homepage and be clearly visible and accessible on deposit pages or through a direct link from them.

A change improving revenue through any of these mechanisms has produced a result that a purely commercial framework will approve and that the operator should decline. Building the guardrails into the standard evaluation means the question gets asked automatically rather than depending on someone raising it.

Things not to test

A short and important category.

Some experiments should not be run because the finding, if positive, would be one the operator should not act on. Testing whether a design increases spend among customers displaying harm indicators is the clearest case. Testing whether reducing the prominence of responsible gambling tools improves engagement is another. Testing the effect of promotional messaging timed to moments of loss is a third. In Britain the last two would also run into the technical standards on the visibility of limit-setting tools and on products that encourage customers to chase losses.

In each case the test is technically straightforward, the finding would likely be commercially positive, and running it would be indefensible.

The useful heuristic is the one from the Operations Strategy course: if the test design would be uncomfortable to describe publicly, that discomfort is information about the test rather than about public opinion. It is worth noticing before approval rather than after.

There is also a category of test that is fine to run and requires care in interpretation, namely anything where the population affected is small and vulnerable. A change affecting a small number of very high-spending customers may produce a statistically clear result that should still be examined for who those customers are.

Novelty and decay

A final practical caution.

Changes frequently produce a short-term response simply because they are different. Customers notice something new, engage with it, and the effect decays as it becomes familiar. Tests run over short periods capture the novelty and mistake it for a durable effect.

The response is duration. A test that runs long enough for the novelty to decay measures the sustained effect, which is the one that matters. Where duration is constrained, the pattern within the test period is informative: an effect that is large initially and declining throughout may be novelty, and one that is stable is more likely to be real. Read a trend with care, though: early estimates are noisy, and Microsoft's experimenters found that most suspected novelty and primacy effects were not real but a statistical artefact, so confirm a trend over a longer window before acting on it.

The same applies in reverse to changes that disrupt familiarity. A redesign can perform worse initially as experienced customers adapt, known as a primacy effect, and a test stopped early may find it harmful when it is not, although Microsoft reported that primacy effects which reverse an initial result were rare in its experiments.

What to actually track

To close, a compact set of measures that serve product work in this sector better than the usual dashboard.

Journey completion rates by stage, market and device, as covered in the journeys lesson.

Contact rate, since support volume is a direct measure of the friction the product generates.

Cohort contribution over time, which is the fundamental value measure.

Retention curves by acquisition source and by lifecycle stage.

Catalogue breadth in casino, indicating whether discovery is opening or narrowing.

Spend distribution and its stability, which serves both commercial and harm-relevant purposes.

Tool adoption, meaning what proportion of customers set limits and use protective features, which indicates whether the tooling is findable. Take-up alone can flatter: in the Patterns of Play study 21.5% of account holders set a deposit limit during the year, but a significant proportion set limits so high they were unlikely to constrain anything, so look at the level of limits as well as their number.

Performance and reliability, experienced constantly and measured rarely.

That set is shorter than most product dashboards and covers more of what determines whether the product is working. The longer dashboards are usually longer because nobody removed anything, which is the same accumulation problem that afflicts registration forms.

Qualitative evidence

A dimension that quantitative measurement cannot supply and that product teams in this sector use less than they should.

Usability testing watching real people attempt the core journeys reveals problems that funnel data identifies as a drop-off without explaining. A registration form losing customers at a particular field will show as a number; watching five people encounter that field explains why.

Support contact analysis is the cheapest qualitative source available and is largely ignored by product teams. Every contact is a customer explaining precisely where the product diverged from what they expected, categorised and quantified, as covered in the Customer Service course. Product teams that read support transcripts periodically find things no dashboard shows.

Complaint themes serve the same purpose at higher intensity, since a complaint indicates the customer cared enough to escalate.

Review and forum content shows what customers say about the operator publicly, which is both unfiltered and consequential given how much it influences acquisition.

Staff feedback, particularly from support and VIP teams, captures patterns those staff observe daily and that appear nowhere in product data.

The general point is that quantitative measurement tells you what is happening and rarely tells you why. In a sector where the funnel is long, the population is skewed and the outcomes are slow, the qualitative sources are frequently the faster route to a correct diagnosis.

Building the measurement foundation

A closing practical note, since much of this lesson assumes capability that many operators lack.

The prerequisites are consistent with those identified throughout these courses. Event instrumentation at sufficient granularity across the product. Identity resolution across devices and sessions. Attribution that persists from acquisition through the customer's life. Cost attribution at customer level, so contribution rather than only revenue can be measured. Agreed definitions applied consistently. And an experimentation platform capable of allocating customers, holding allocations stable, and computing results correctly including the trimmed and segmented views described above.

Operators lacking these can measure activity and cannot measure value, which means their product decisions are made on proxies with no way to validate them. Building the foundation is unglamorous, takes time, and determines whether everything else in this lesson is available.

Metrics that mislead

A catalogue of the measures most likely to produce wrong conclusions in this sector, with what to use instead.

Average revenue per user. Describes nobody in a skewed distribution. Use deciles, the median, and the proportion of revenue from the top few percent.

Session length. Longer is not better. A long session may indicate enjoyment or may indicate a customer who cannot stop, and the metric cannot distinguish them. Pair it with spend pattern and with the distribution of very long sessions.

Conversion rate in isolation. Improves when quality falls. Pair with downstream value.

Click-through on recommendations. Measures attention, not usefulness. Follow through to session outcome and retention.

Total registrations. Rewards volume regardless of whether those customers verify, deposit or play. Measure through to first deposit and beyond.

Bonus uptake. High uptake indicates an attractive offer, not an effective one. Measure incremental contribution against customers who did not receive it.

Aggregate revenue growth. Can be produced entirely by acquiring more customers of declining quality, as covered in the metrics lesson of iGaming Basics. Use cohort contribution.

Deposit count. Increases when customers deposit repeatedly within a session, which is a pattern warranting examination rather than celebration.

Net promoter or satisfaction in aggregate. Likely to be distorted by outcome, since customers who have just lost may rate the product poorly regardless of its quality. Segment by outcome.

Uptime. Necessary and insufficient, since a product that is available and slow is experienced as broken.

The common thread is that each of these measures activity, and activity is not value. The correction in every case is to follow through to what the activity produced and to look at how it is distributed.

Interpreting a result

A short protocol for what to do when a test returns a positive result, since the moment of a favourable finding is when scrutiny is weakest.

Check the mechanism. Can we explain why this worked? A result nobody can explain is more likely to be an artefact than an insight. Experimenters call this Twyman's law: any figure that looks interesting or different is usually wrong.

Check the distribution. Did this improve outcomes broadly, or is it carried by a small group? Remove the largest customers and re-run.

Check the guardrails. Did anything worsen, particularly the harm-relevant measures.

Check for displacement. Did total activity increase, or did this surface capture activity from elsewhere?

Check duration effects. Is the effect stable across the test period, or declining, which may indicate novelty and should be confirmed over a longer window.

Check segments. Does this help everyone, or help one group while harming another? Aggregate improvement can conceal a group made worse off.

Check the counterfactual. What was happening in the control group, and is there any reason it was affected by something else during the period.

Applying this to every positive result sounds burdensome and takes very little time once habitual. It also prevents the accumulation of changes that individually tested well and collectively made the product worse, which is a recognisable end state and one this sector has examples of.

Where to start

For a product function whose measurement is currently limited to a dashboard of activity metrics, the sequence that produces the fastest improvement.

Instrument the core journeys properly. Registration, verification, deposit and withdrawal, with each step and each error state captured. This alone usually reveals a specific loss nobody knew about.

Establish cohort contribution reporting. Grouping customers by acquisition period and following contribution over time is the foundation for every value question, and most operators do not have it readily available.

Define the metric set and remove the rest. A short set of measures that reflect value, applied consistently, beats a comprehensive dashboard nobody reads.

Add the guardrails. Spend distribution, effect on customers with risk indicators, session length tail. These need to exist before they are needed.

Validate one proxy. Take a proximate metric the team relies on and check whether it actually predicts downstream contribution for this operator. The answer is informative either way.

Read support transcripts. Cheapest qualitative source available and almost never used by product teams.

Each of these is available without new tooling and without a data science function. They establish the conditions under which the more sophisticated work in this lesson becomes possible, and skipping them produces sophisticated analysis of unreliable data.

Key terms

Proxy metric
A measurable indicator used in place of an outcome that is slow or difficult to observe, such as conversion standing in for lifetime value.
Displacement
An apparent improvement produced by moving activity from elsewhere rather than creating anything new.
Trimmed metric
A measure calculated with extreme values excluded, used to reduce the influence of outliers in a skewed distribution.
Guardrail metric
A measure monitored during a test to check that the primary metric is not giving a misleading signal, for example by detecting harm, slower performance or lost revenue while the primary metric improves.
Novelty effect
A short-term response to a change simply because it is different, which decays and can be mistaken for a durable improvement. Its opposite, the primacy effect, is an initial dip while customers adjust to a change.

Key takeaways

  • Proxy metrics are necessary because the outcomes that matter are slow, and they systematically favour changes that produce immediate response over durable value.
  • Revenue skew breaks conventional test design, and any result should be checked for whether it survives removing the largest few customers.
  • Displacement is a common false positive in product measurement, where a feature captures activity that would have happened anyway.
  • Guardrail metrics should include harm-relevant measures, because a change can improve every commercial metric through a mechanism the operator should not want.
  • Some questions should not be tested, and recognising them is part of the discipline rather than an external constraint on it.

Sources

The legislation, regulator material and research this lesson was checked against.

  1. Customer interaction guidance for remote gambling licensees (formal guidance under SR Code 3.4.3), Gambling Commission, accessed 2026-09-23
  2. Remote gambling and software technical standards, RTS 12: Financial limits, Gambling Commission, accessed 2026-09-23
  3. Remote gambling and software technical standards, RTS 14: Responsible product design, Gambling Commission, accessed 2026-09-23
  4. Patterns of Play: Extended Executive Summary Report (Forrest, McHale et al., June 2022, prepared for GambleAware), NatCen Social Research and University of Liverpool, accessed 2026-09-23
  5. Seven Rules of Thumb for Web Site Experimenters (Kohavi, Deng, Longbotham, Xu, KDD 2014), ACM SIGKDD / Microsoft ExP, accessed 2026-09-23
  6. Trustworthy Online Controlled Experiments: Five Puzzling Outcomes Explained (Kohavi et al., KDD 2012), ACM SIGKDD / Microsoft ExP, accessed 2026-09-23
  7. Always Valid Inference: Bringing Sequential Analysis to A/B Testing (Johari, Pekelis, Walsh), arXiv, accessed 2026-09-23
  8. Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data (Deng et al., WSDM 2013), ACM WSDM / Microsoft ExP, accessed 2026-09-23
  9. Data-Driven Metric Development for Online Controlled Experiments: Seven Lessons Learned (Deng and Shi, KDD 2016), ACM SIGKDD / Microsoft ExP, accessed 2026-09-23

Check your understanding

3 questions · answer them all, then check.

  1. 1. Why do proxy metrics systematically favour certain kinds of change?

  2. 2. What is displacement in product measurement?

  3. 3. Why should guardrail metrics include harm-relevant measures?

Sign in to track your progress through the course.

Cookie Preferences

Choose which cookies you want to accept. Essential cookies are required for the website to function properly.

Required

Necessary for the website to function. Cannot be disabled.

Help us understand how visitors interact with our website.

Used to deliver relevant advertisements and track ad performance.

Remember your preferences and settings for a better experience.

Product Measurement and Experimentation | Product Innovation