Evidence is not free
Operators talk about being data-driven as though evidence were costless and its absence merely a failure of will. Analysis consumes analyst time, engineering time, and the calendar time during which a decision is not being made. Testing consumes traffic and delays implementation.
Because it is not free, the useful question is not whether to use evidence but where to spend it. The answer follows from the reversibility framework in the first lesson of this course.
Reversible decisions should generally be tested rather than analysed. A promotional structure, a page layout, a message or a method ordering can be tried and reversed, and doing so produces better information than a study would.
Irreversible decisions warrant analysis, because there is no opportunity to learn from doing. Platform selection, market entry, acquisition and organisational structure all fall here, and the evidence available will be imperfect, which is precisely why judgement is required.
The common failure is inverting this: extensive deliberation over choices that could have been tested in a fortnight, and rapid commitment to decisions that will constrain the operator for a decade.
Fix the data before analysing it
Most analytical problems in gambling operators are not analytical. They are attribution and definition problems, and analysis built on them produces confident wrong answers.
Attribution. As covered in the Affiliate Marketing course, the link between a customer and their acquisition source is frequently lost downstream, which makes channel-level value analysis impossible. The same applies to cost attribution: bonus, payment and servicing costs held only in aggregate permit revenue analysis and not contribution analysis, and contribution is the figure that matters.
Definitions. Active player, first time depositor, net revenue, churn and cohort each have several plausible definitions, and operators routinely run more than one simultaneously without reconciling them. Marketing's active count and finance's active count differing is not a rounding issue; it means the two functions are describing different populations.
Identity. The same person appearing as several customers across brands, devices or accounts distorts every per-customer measure.
Timing. Reporting periods that do not align across systems produce discrepancies that get investigated repeatedly and are not errors.
The practical requirement is unglamorous: agree definitions once, document them, build attribution that persists, and treat the reporting layer as infrastructure. Operators that skip this spend their analytical capacity reconciling numbers rather than answering questions.
Experiments and their constraints
Controlled comparison is the strongest tool available for reversible decisions, and this sector imposes real constraints on it.
Regulatory limits on differential treatment. Offering different terms, prices or promotions to different customers is constrained in some jurisdictions, and testing that involves materially different value propositions may not be permissible. This does not prevent testing presentation, messaging or flow, and it does limit testing of commercial terms.
Ethical limits. Some tests should not be run regardless of legality. Testing whether a change increases spend among customers displaying harm indicators is not a legitimate experiment, and a test design that would be uncomfortable to describe publicly warrants examination before it warrants approval.
Revenue skew. This is the technical constraint that catches people out. With revenue concentrated in a small minority of customers, random allocation does not reliably balance value between test groups. A single high-value customer landing in the treatment group can produce an apparent effect entirely unrelated to the change being tested.
The responses are to use measures less sensitive to outliers, such as median or trimmed metrics, alongside means; to analyse the high-value tail separately from the main population; to require larger samples and longer durations than conventional guidance suggests; and to be extremely cautious about results driven by a small number of customers, which is checkable by removing the top few and seeing whether the effect survives.
Long feedback loops. The outcomes that matter, retention and lifetime contribution, take months to observe. Tests measured on immediate conversion frequently favour changes that harm long-term value, and this is one of the more consistent sources of error in this sector.
Seasonality and events. Sporting calendars and promotional periods create variation large enough to swamp test effects, which means test periods must either span them or be interpreted with that variation accounted for.
Where testing does not apply
Some questions cannot be tested and require different methods.
Market entry cannot be trialled. Analysis, comparison with analogous markets and explicit assumption-testing are what is available.
Platform selection cannot be run in parallel meaningfully. Reference checking, structured evaluation and honest assessment of switching cost are the tools.
Organisational change cannot be A/B tested. It can be phased and reviewed.
Regulatory decisions are not optional and are not subject to commercial testing.
For these, the useful discipline is making assumptions explicit. A decision resting on an assumption that a market will produce a given level of revenue at a given acquisition cost should record that assumption, so that it can be checked against reality afterwards and so that the decision can be revisited if it proves wrong. Decisions whose assumptions were never written down cannot be learned from, because nobody can establish afterwards what was believed.
The failure modes
Several patterns recur in how operators use data, and each is worth recognising.
Analysis to justify rather than to decide. The test is counterfactual: would a contrary result change anything? If not, the analysis is documentation and the effort was wasted. This is extremely common and is usually visible in the framing of the request.
Optimising the measurable. Attention flows to what is reported, which in this sector means revenue and conversion. Structural decisions, retention quality and risk exposure generate no dashboard and receive correspondingly less scrutiny.
Metric fixation without mechanism. Improving a number without understanding what produced it produces gains that reverse. A conversion improvement achieved by removing a step that existed for a reason will show well and cost more later.
Aggregates concealing segments. Established throughout these courses. Averages in a skewed distribution describe nobody.
Confusing correlation with cause. Customers who use a feature spend more; therefore the feature increases spend. Almost always this is selection, since engaged customers use more features. Distinguishing requires either an experiment or a genuinely careful design.
Survivorship in retention analysis. Analysing the characteristics of long-retained customers to identify what drives retention systematically ignores everyone who left, which is where the information about churn actually sits.
Precision beyond the data. Reporting figures to a granularity the underlying measurement cannot support, which conveys false confidence and invites decisions the evidence does not carry.
Decisions that survive seniority
The organisational aspect matters as much as the analytical one, because analysis is only useful if it can change a decision.
Practices that help are reasonably simple. Stating the question before the analysis, so that the answer cannot be reverse-engineered. Recording the decision, including what was decided, why, on what evidence, what was uncertain, and what would cause it to be revisited. Separating the analyst from the advocate, so that the person producing the evidence is not the person whose proposal it supports. Making disconfirming evidence welcome, which is a cultural matter demonstrated by how the last piece of inconvenient analysis was received rather than by anything stated. And revisiting decisions against their recorded assumptions, which is the only mechanism by which an organisation learns.
The decision record is the single most valuable of these and the least used. Six months after a market entry, a platform selection or a structural change, the ability to see what was believed at the time, what was uncertain, and what would have counted as evidence against, converts a past decision into information. Without it, the organisation relies on recollection, which reliably reconstructs past reasoning to match present outcomes.
Proportion
A closing caution against the opposite error.
An organisation that requires evidence for every decision becomes slow, and slowness is itself a cost that no analysis captures. Many decisions in this industry are made under genuine uncertainty where the evidence will not resolve the question, and pretending otherwise produces delay rather than accuracy.
The practical balance is to spend analytical effort where reversibility is low and consequences are large, to test where testing is cheap, and to accept that a substantial proportion of decisions will be made on judgement supported by partial evidence. That is not a failure of rigour. It is the ordinary condition of operating a business, and the operators that handle it well are the ones that distinguish clearly between the decisions where they can know and the decisions where they must choose.
Designing a test properly
For the decisions where testing is available, some practical discipline that improves the proportion of tests producing usable answers.
State the question and the decision rule first. What are we trying to learn, and what result would cause us to do what. A test without a pre-stated decision rule invites interpretation after the fact, which is how ambiguous results become confirmation.
Estimate the effect size worth detecting. A test powered to detect a large effect will not detect a small one, and small effects are frequently what is available. If the sample cannot support detecting an effect worth acting on, the test should not be run.
Determine duration before starting. Stopping a test when the result looks favourable is one of the most common errors in commercial experimentation, and it reliably produces effects that do not replicate.
Check allocation balance. Given the revenue skew described above, verify that groups are comparable on value before attributing differences to the treatment.
Measure the outcome that matters, not the proximate one. A change that improves deposit conversion and reduces retention is not an improvement, and measuring only the first will find it to be one.
Look for harm as well as benefit. In this sector, a change that increases play warrants examination of who is playing more. A test that improves aggregate revenue by increasing spend among customers displaying risk indicators has produced a result that should not be implemented.
Record the result including the failures. Tests that found nothing are informative and are routinely discarded, which means organisations repeat them.
Building the analytical function
A brief organisational note, since analysis quality depends on how the function is arranged.
Analytics fully centralised as a service tends to become a request queue, disconnected from the decisions it should inform, producing work that arrives after the decision was made. Analytics fully embedded in functions tends to produce inconsistent methods and no shared standards, with each team's numbers differing.
The arrangement that generally works places infrastructure, definitions and standards centrally, with analysts embedded in the functions and markets they serve while retaining a professional line to the centre. This mirrors the compliance structure described earlier in this course and for the same reason: local proximity is required for relevance, central connection is required for consistency and independence.
The independence point matters here too. An analyst whose reporting line runs entirely to the person whose proposals they assess is under the same structural pressure as a compliance officer in that position. It does not mean the analysis will be wrong; it means the arrangement makes it harder to be right.
Questions worth asking of any analysis
A short checklist that catches most of the failure modes described above.
What question was this designed to answer, and was that stated before the work began?
Would a different result have changed the decision? If not, this is documentation.
What is the comparison group, and is it genuinely comparable?
Is this an average across a skewed distribution? If so, what does the distribution look like.
Is the effect driven by a small number of observations? Remove the largest few and check whether it survives.
Could the causation run the other way, or could both be caused by something else?
Who is excluded from this analysis? Churned customers, declined transactions, abandoned sessions and customers who never registered are all invisible in most operator data and are frequently where the answer sits.
What is the measurement error, and is the precision being reported consistent with it?
What would we expect to see if this conclusion were wrong? And have we looked.
None of these requires statistical training to ask, and asking them routinely improves decision quality more than any additional analytical capacity would.
The data the industry does not have
A closing observation about the limits of what any single operator can know.
An operator sees its own customers' activity with complete precision and sees nothing else. It does not see what those customers do with competitors, which means it cannot assess total gambling spend, cannot know whether a customer who reduced activity stopped gambling or moved, and cannot assess affordability against the full picture. It does not see the customers it failed to acquire, or the deposits that failed before reaching its systems. It does not see why churned customers left, only that they did.
These blind spots shape what analysis is capable of. Retention analysis conducted on an operator's own data cannot distinguish a customer who quit gambling from one who switched, and those are different findings with different implications. Affordability assessment based on activity with one operator understates total exposure, which is the cross-operator visibility problem covered in the Payment Operations course.
The practical response is partly to acquire what can be acquired, through survey research, market data and, where frameworks permit, external financial data with consent. It is also to be explicit about the limits when drawing conclusions, since analysis that presents an operator-level finding as a market-level one is overreaching in a way that is easy to miss.
The broader point is that being data-driven does not mean the data answers the question. It frequently means knowing precisely which part of the question the data can address and being honest about the rest.
Judgement is not a failure of rigour
One last point, because courses on evidence tend to imply that judgement is what remains when analysis has failed.
Many of the most consequential decisions in this sector cannot be resolved by evidence, and treating them as though they could produces delay and false confidence rather than accuracy. Whether a regulatory regime will tighten. Whether a market's competitive intensity will moderate. Whether an acquired team will stay. Whether a platform supplier will keep investing. These are judgements about the future made under genuine uncertainty, and no amount of analysis converts them into facts.
What good practice provides is not certainty but better-structured judgement: the assumptions made explicit, the range of outcomes considered, the disconfirming evidence sought, the decision recorded so it can be examined later, and the trigger identified that would cause a rethink.
An operator that does this is not more likely to be right about any individual decision. It is considerably more likely to notice when it was wrong, which over time matters more.