The question behind every finding
Someone observes that customers who use a feature spend more, or that customers who received an offer generated revenue, or that a change was followed by an improvement.
In each case the implied claim is causal: the feature increases spend, the offer generated the revenue, the change produced the improvement.
In each case the observation supports a considerably weaker claim, and the gap between what was observed and what is being asserted is where most analytical error in commercial settings occurs.
This lesson is about closing that gap.
Selection, the dominant confound
The specific problem in this industry is that groups almost always formed themselves.
Customers who used a feature chose to use it, and the kind of customer who explores features is a more engaged customer who would have spent more anyway.
Customers who accepted an offer were sufficiently engaged to open the message and act on it.
Customers in a high-value segment are there because they spent, so finding that the segment spends more is circular.
Customers who contacted support had a problem, which makes them different from those who did not in ways that affect everything downstream.
Customers who set a deposit limit did so for reasons, and comparing their subsequent spend to those who did not conflates the tool's effect with whatever prompted them to use it.
In every case the comparison measures the difference between kinds of customer rather than the effect of the thing. Statistical adjustment for observed differences helps and does not solve it, because the characteristic that caused the selection may not have been measured.
The practical discipline is to ask, of any comparison, how the groups came to be different. Where the answer is that they selected themselves, the comparison is descriptive rather than causal, and it should be presented as such.
Randomisation
The only method that handles unobserved confounders, which is why it is worth the cost.
Assigning subjects to groups at random makes the groups comparable in expectation on everything, whether or not it was measured. That is the property adjustment cannot replicate.
In a gambling operator this usually means a holdout: a randomly selected proportion of a qualifying population that does not receive a treatment, whose subsequent behaviour is the baseline.
The design requirements, developed in the CRM material and worth restating here in analytical terms.
Random assignment from the qualifying population, verified rather than assumed, since implementation bugs in assignment are common and invalidate everything.
Assignment stability, so a customer assigned to holdout remains in it rather than being reselected each period.
Contamination control, meaning the holdout does not receive the treatment through another route, which in an operator with several communication channels requires checking.
Pre-period comparison, confirming the groups were comparable before the treatment. If they differ materially beforehand, the randomisation did not work and the comparison is compromised.
Adequate size and duration, which given skew means larger and longer than standard calculators indicate.
Pre-registered analysis, meaning the metric, the period and the decision rule agreed before the result arrives.
Skew and test design
The statistical consequences of this data's distribution, which are substantial enough to warrant their own treatment.
Variance is high, which means detecting a given effect requires a larger sample than in most consumer contexts.
Random assignment does not reliably balance value. With revenue concentrated in a small minority, a single high-value customer landing in one group can produce a difference unrelated to the treatment. This is the single most common source of spurious test results in this industry.
The mean is unstable and results reported on it alone should be treated as provisional.
The design responses.
Check balance on pre-period value before interpreting.
Report median and trimmed measures alongside the mean. An effect in the mean and absent in the median is being carried by a small number of observations.
Remove the largest few and re-run. If it does not survive, it was not an effect.
Analyse the high-value tail separately, since the treatment may genuinely affect them differently and averaging conceals it.
Consider stratified assignment, randomising within value bands so that high-value customers are distributed evenly by construction rather than by chance.
Run longer. Both to accumulate sample and to allow novelty effects to decay.
Do not stop early on a favourable result, which is the most common error in commercial experimentation and reliably produces effects that fail to replicate.
When randomisation is unavailable
Many questions cannot be randomised: a market entry, a platform change, a regulatory implementation, a price change applied to everyone. Quasi-experimental methods can support causal claims with stated assumptions.
Difference in differences. Compare the change in a treated group against the change in an untreated comparison group over the same period. This removes anything affecting both, subject to the assumption that the two would otherwise have moved in parallel. That assumption is testable in the pre-period, and should be tested rather than asserted.
Interrupted time series. Model the pre-change trend and compare the post-change observations to its projection. This handles a treatment applied to everyone and depends on the assumption that the pre-existing trend would have continued, which is weaker over longer horizons.
Regression discontinuity. Where a treatment applies above a threshold, compare subjects just above and just below. Those subjects are similar in everything except the treatment, which makes the comparison strong locally, and the finding applies only near the threshold.
Matching. Constructing a comparison group resembling the treated group on observed characteristics. This addresses observed confounding and leaves unobserved confounding intact, which given selection effects is frequently the substantial part.
Instrumental approaches, where something influences treatment assignment without directly affecting the outcome, which is elegant, rare and demanding to justify.
Each method rests on assumptions. The discipline is to state them, test them where possible, and present the finding with the assumptions attached rather than as an unqualified causal claim.
Confounds that recur here
A catalogue specific to this industry.
Seasonality and the sporting calendar, which dominate short comparisons and require period alignment.
Promotional overlap, where a test period contained a campaign affecting one group more than another.
Regulatory change, which alters the product mid-test.
Payday cycles, visible in deposit patterns.
Cohort composition, where a test population includes customers of different tenures whose behaviour differs systematically.
Platform and release effects, where a deployment during a test changed something for everyone.
Novelty, where a change produces a temporary response that decays.
Survivorship, where analysis of customers still present excludes those who left, which is where the information about churn actually sits.
Attribution instability, where the link between customer and source changed during the period.
Each is checkable and each has produced wrong conclusions somewhere.
Tests that should not be run
A category that belongs in an analytical course rather than only in an ethics one.
Some experiments are technically straightforward, would likely produce a commercially positive finding, and should not be conducted.
Testing whether a design increases spend among customers displaying harm indicators. Testing whether reducing the prominence of protective tools improves engagement. Testing the effect of promotional messaging timed to follow losses. Testing whether removing a friction that exists for protective reasons increases deposits.
In each case the finding, if positive, is one the operator should not act on, which means running the test produces knowledge with no legitimate application and a record that would be difficult to explain.
The heuristic from the Operations Strategy course applies: if a test design would be uncomfortable to describe publicly, that discomfort is information about the test.
There is a related category requiring care rather than prohibition. Tests affecting small vulnerable populations may produce statistically clear results that warrant examination of who those subjects were before the finding is applied.
Presenting causal findings
To close, how to communicate a result honestly.
State the design. Randomised, quasi-experimental or observational, since the strength of the claim follows from it.
State the comparison. What was compared to what, and how the groups came to be different.
State the assumptions, particularly for quasi-experimental work.
State the effect size, not only significance, since a statistically clear effect may be commercially trivial.
Show the distribution, given everything in the previous lesson.
State the uncertainty, as a range rather than a point where possible.
State what would falsify it, which is the discipline that distinguishes a finding from a claim.
Distinguish what was measured from what is inferred. An observational finding presented as causal is the most common overreach in commercial analysis, and it is usually unintentional, occurring in the summary rather than in the analysis.
A worked confounding example
To demonstrate how these errors arise in practice, a realistic case.
An operator observes that customers who set a deposit limit generate more revenue over the following six months than customers who do not. The finding circulates, and a proposal follows to promote limit-setting more widely on the basis that it increases revenue.
The finding is real and the causal interpretation is almost certainly wrong.
The selection. Customers who set limits did so for reasons. The population includes people who are engaged enough to explore account settings, people who are managing a budget deliberately, and people who intend to keep playing over a long period. All of those characteristics predict higher subsequent revenue independently of the limit.
The comparison group includes customers who never engaged with the product enough to find the settings, customers who deposited once and left, and customers who were not thinking about their play at all.
What is actually being compared is engaged customers who plan against a general population containing many who barely started.
A better comparison would match on prior activity, tenure and engagement, which would narrow the gap substantially and would still leave unobserved differences.
A causal test would randomise the prompt to set a limit and compare those who received it against those who did not, measuring outcomes across the whole assigned group rather than only among those who acted on it.
The likely finding from that design is that prompting limit-setting has a modest effect on revenue in either direction and a meaningful effect on the protective outcomes it is actually for.
The general lesson is that a comparison between people who did something and people who did not is almost never a causal comparison, and that the more interesting the finding, the more likely selection explains it.
Practical experimentation in an operator
A note on running these in practice, since the constraints are organisational as much as statistical.
Getting a holdout accepted is the main obstacle, because it means deliberately withholding something from customers. The argument that works is the one from the CRM material: the objection presupposes the treatment works, which is what the holdout exists to establish.
Keeping the holdout clean requires coordination, since a customer excluded from one programme may be included in another affecting the same outcome.
Duration discipline is difficult when a result appears early, and the practice of stopping when the answer looks favourable is the single most common source of results that do not replicate.
Pre-registration, meaning the metric, period and decision rule agreed in writing beforehand, is the mechanism that prevents reinterpretation afterwards.
Documenting the negatives matters, because tests that found nothing are informative and are routinely discarded, which means organisations repeat them.
A small permanent holdout on major ongoing programmes is the most practical arrangement, giving continuous measurement without a discrete test for each change.
Sequencing matters where several changes are proposed at once, since simultaneous changes to the same population cannot be attributed individually.
Modelling and its causal limits
A note on predictive models, since they are frequently treated as answering causal questions and do not.
A model predicting churn identifies customers likely to leave. It does not establish what would prevent them leaving, and the features it found predictive are not necessarily levers.
A model predicting value identifies customers likely to be worth more. It does not establish that treating them differently will change that.
A model identifying customers likely to be experiencing harm is finding a pattern in behaviour. What intervention would help is a separate question requiring separate evidence.
The distinction matters because model outputs get used as though they were causal. A feature appearing prominently in a churn model is frequently interpreted as a cause of churn and targeted for improvement, when it may be a symptom, a marker of a customer type, or correlated with something unmeasured.
The practical disciplines.
Distinguish prediction from explanation explicitly when presenting model output.
Do not read feature importance as causal, since it describes what predicts in the presence of the other features rather than what drives anything.
Test interventions separately, since the fact that a model identifies a group does not establish that any particular action helps them.
Validate against outcomes, meaning what actually happened to the people the model identified, rather than against a proxy.
Watch for feedback, since a model whose predictions drive actions will subsequently be trained on data those actions shaped.
That last point is worth emphasising. A churn model driving retention offers will be retrained on a population where the customers it flagged received offers, which changes what it learns. Without care the model ends up predicting who received an offer rather than who was going to leave.
The honest summary
To close this lesson, what an analyst should actually claim.
Most findings in commercial analysis are associations. That is not a failure; associations are useful, they direct attention, and they support decisions where the cost of being wrong is low.
The failure is presenting an association as a cause, which happens most often not in the analysis but in the summary, where hedged language gets compressed into a claim.
The honest positions available, in order of strength.
We observed that X and Y occur together, which is a description.
X predicts Y, which is a stronger statement about the association's reliability and still not causal.
Among comparable groups differing in X, Y differed, which is a matched comparison and controls for what was observed.
In a randomised comparison, changing X changed Y, which is a causal claim and the only one that fully earns the word.
An analyst who is precise about which of these they are asserting will occasionally disappoint someone who wanted a stronger statement. They will also not be the person whose finding was acted on and turned out to have been selection all along, which happens regularly and is the reason the distinction is worth the discomfort.