An unusual data position
Gambling operators hold data that most consumer businesses would envy.
Every transaction is recorded. Every session, every stake, every outcome, every deposit and withdrawal, every communication sent and every response. The customer is identified, verified and linked to their activity from the first interaction onwards. There is no offline component to lose, no cash to reconcile, and no gap between behaviour and record.
That completeness means the constraint on analysis quality in this industry is rarely data availability. It is almost always something else: definitions that disagree, attribution that breaks, questions that were not worth asking, and conclusions drawn from averages of skewed distributions.
This course is about those problems rather than about statistical technique, because they are where the errors actually occur.
The questions worth answering
Analysis in an operator supports a limited set of decisions, and it is worth being explicit about them.
Where are we losing customers, and why? Funnel drop-off, churn causes, and the upstream problems generating both.
Which customers are worth what? Value distribution, cohort contribution, and the difference between customers acquired through different routes.
Is our acquisition working? Cost per customer by channel, and whether the customers acquired are worth what was paid.
Does this change do anything? Product, promotional and process changes assessed against a comparison.
What is our exposure? Concentration by market, by supplier, by customer segment, and the risks that do not appear in performance reporting.
Who needs attention? Customers displaying indicators requiring intervention, which is an analytical capability as much as a compliance one.
What is actually happening? The descriptive layer, which is necessary and is not where the value concentrates.
A question that does not map to one of these is usually a request for a number rather than a request for a decision, and distinguishing them early saves a great deal of effort.
Descriptive, diagnostic and predictive
A useful hierarchy, because most operators are heavily weighted towards the first.
Descriptive answers what happened. Revenue by month, active players by market, deposits by method. Necessary for running a business and, on its own, not decision-relevant, because it does not indicate what to do.
Diagnostic answers why. Why did revenue fall in that market, why did that cohort retain worse, why did acceptance drop on that route. This is where most decision-relevant work sits and where most operators are thinnest.
Predictive answers what will happen. Churn likelihood, expected value, risk of harm. Valuable, dependent on the previous two being sound, and frequently attempted before the foundations support it.
Prescriptive answers what to do, which in practice means a human deciding with the previous three in front of them.
The common failure is an organisation with extensive descriptive reporting, minimal diagnostic capability, and an ambition to build predictive models. The models produce outputs nobody trusts because the definitions underneath them are inconsistent, and the effort would have been better spent on the diagnostic layer.
The skew problem
The single most important characteristic of gambling data, and the one that invalidates the most analysis.
Almost every commercial distribution in this industry is heavily skewed. Revenue concentrates in a small proportion of customers. So does deposit volume, session time, and lifetime value. Game play concentrates in a small proportion of titles. Traffic concentrates in a small proportion of affiliates.
The consequences are pervasive.
Means describe nobody. An average revenue per user figure in a distribution where a small minority generates most of the revenue describes a statistical artefact, not a customer.
Comparisons are unstable. Two randomly selected groups will differ substantially in mean value simply because of where the largest customers landed.
Test results are unreliable without checking whether an effect survives removing the largest observations.
Aggregates conceal. A total that is stable may contain a growing number of low-value customers offsetting a declining number of high-value ones.
Segment sizes mislead. A segment containing a small proportion of customers may contain most of the value.
The correction is distributional thinking as the default. Report the median alongside the mean. Show deciles. Show the distribution rather than a summary of it. Check whether a finding survives the removal of outliers. Ask what proportion of the effect comes from what proportion of the observations.
None of this is sophisticated. All of it is skipped routinely, and it is the most common source of confident wrong conclusions in this industry.
Actionability
The filter worth applying before work begins.
A question is worth answering if the answer would change a decision. That test eliminates a substantial proportion of analytical requests, and applying it is uncomfortable because the requests come from people who want the number.
The useful conversation is to ask what the requester would do differently depending on the answer. Where the answer is nothing, the request is for reassurance, for a slide, or for confirmation of a decision already taken, and it should be recognised as such.
Where the answer is something, the analysis has a purpose and the purpose usually clarifies what is actually needed, which is frequently narrower and faster than the original request.
The related discipline is stating the decision rule in advance. If an analysis is intended to determine whether to continue an activity, agreeing beforehand what result would mean stopping prevents the finding being reinterpreted after it arrives.
Analytical debt
A concept borrowed from software and directly applicable.
Every shortcut accumulates. A metric defined slightly differently in two reports. An attribution link that silently stops working after a platform change. A transformation whose logic exists only in one person's memory. A segment definition that drifted without documentation.
Individually each is minor. Collectively they produce an environment where two analyses of the same question return different answers, nobody can establish which is right, and the resulting disagreement consumes more time than the original question warranted.
The symptoms are recognisable. Meetings spent reconciling numbers rather than discussing them. Analysts unable to explain why a figure differs from another source. Reports whose logic nobody can trace. Reluctance to change anything because the dependencies are unknown.
The remedies are unglamorous: documented definitions, tested attribution, versioned logic, and a single agreed source for each core metric. They compete against delivering the next analysis, which is why the debt accumulates.
The practical argument for addressing it is that an operator with reliable foundations answers questions in hours that an operator with analytical debt answers in days, after a reconciliation exercise, with a caveat about which source was used.
What analysts actually need to know
To close, the capabilities that matter in this role, roughly in order of value.
Understanding the business. An analyst who does not know how the operator makes money, what the verticals do or why customers churn will produce technically correct answers to the wrong questions. This is the largest differentiator and the least discussed.
Distributional thinking, given everything above.
Causal reasoning, meaning knowing when a comparison supports a causal claim and when it does not, which the experimentation lesson covers.
Definitional rigour, since most disagreement is definitional rather than analytical.
Data lineage, meaning knowing where a number came from and what happened to it on the way.
Communication, since analysis that does not reach a decision-maker in a form they can act on has not done anything.
Technical skill, which matters and is further down this list than most analysts expect, because the constraint in this industry is rarely the ability to compute something.
The remaining lessons work through the data landscape, metric definitions, the techniques that suit this data, causal inference, reporting, and the quality and governance problems that undermine all of it.
Working with the rest of the business
Analysis produces value only when it reaches a decision, which makes the relationship with other functions a substantial part of the role.
Commercial and marketing generate most requests and generate them as questions about numbers rather than as decisions. The productive response is to establish the decision behind the request, which usually narrows and clarifies it.
Product needs measurement designed before a change ships rather than assembled afterwards, which requires analytics involved at planning rather than at review.
Finance owns the authoritative revenue figures and frequently defines them differently from everyone else, which is the definitional divergence covered in the metrics lesson.
Compliance and safer gambling need detection, monitoring and assurance analysis, and this is where analytical work has the most direct consequences for customers. It is also where the constraints described in the Product Innovation course apply, particularly around what a model may optimise for and what features it may use.
Trading in a sportsbook operation runs its own analytical function with specialist requirements, and the boundary between it and general analytics should be clear.
Technology owns the pipelines and the environment, and an analytics function without a working relationship there is limited to querying what already exists.
The organisational point, made in the Operations Strategy course, is that analytics fully centralised as a request queue becomes disconnected from decisions, and analytics fully embedded produces inconsistent methods. The arrangement that works places standards and infrastructure centrally with analysts embedded in the functions they serve.
Two failure modes
The recognisable ways an analytics function goes wrong.
The report factory produces a large volume of scheduled reporting, responds to requests as they arrive, and never establishes whether any of it changes anything. It is busy, it is measured on output, and its influence is minimal because it answers questions rather than shaping them.
The isolated laboratory builds sophisticated models disconnected from the operating business, using data whose reliability it has not established, answering questions nobody asked. It produces impressive work that nobody adopts.
The functions that avoid both share characteristics: they understand the business well enough to identify the questions that matter, they invest in foundations before technique, they are present when decisions are made rather than after, and they are willing to say that an analysis will not support the conclusion someone wants to draw from it.
That last characteristic is the one that determines whether the function is worth having, and it depends on the organisational conditions described in the Leadership course rather than on anything an analyst controls.
What good looks like
To close, the characteristics of an analytics function that is working.
It knows the business. Analysts can explain how the operator makes money, what each vertical does, and why customers leave, without reference to a dashboard.
It shapes questions rather than only answering them. Requests arrive as questions about numbers and leave as questions about decisions.
Its foundations are sound. Definitions are agreed, attribution works, identity resolves, and lineage is traceable.
It thinks distributionally. Medians and deciles appear alongside means as a matter of course, and results are checked for whether they survive the removal of outliers.
It is honest about uncertainty. Findings carry the caveats they warrant, and the function says when data cannot support a conclusion.
It measures what matters. Contribution rather than revenue, retention rather than activity, incremental effect rather than response.
It reaches decisions. Analysis arrives before the decision rather than after it, in a form the decision-maker can act on.
It says no. To requests that would not change anything, and to conclusions the evidence does not support.
The last is the hardest and is the clearest indicator of the rest, because a function that has never declined a request or contradicted a preferred conclusion is producing documentation rather than analysis.
The questions nobody asks
A closing observation about what is missing from most analytical agendas.
Requests arriving at an analytics function are shaped by what the organisation currently attends to, which means they cluster around revenue, acquisition and campaign performance. Those are legitimate and they are not the only questions that matter.
The questions that rarely arrive and usually should.
What proportion of our revenue comes from the top few percent of customers, and is it rising? A concentration measure that is both a commercial risk indicator and a protective one.
Are successive cohorts worth more or less than their predecessors? Which distinguishes growth from deterioration and is invisible in aggregate revenue.
Which of our markets contribute after fully loaded costs? Frequently a different list from the one that contributes on revenue.
Where are we concentrated? By supplier, by channel, by market, by customer segment. These risks materialise suddenly and appear in no performance report.
What did that change actually cause? Rather than what happened after it.
How long do new customers survive, and has that changed? The leading indicator of most subsequent problems.
What happened to the customers who displayed risk indicators? Which is a compliance question, an analytical one, and the one the Law and Compliance course identifies as determining enforcement outcomes.
An analytics function can raise these without being asked. Doing so is how the function moves from answering questions to shaping what the organisation looks at, which is where its influence actually comes from.
Where to start in a new role
For an analyst joining an operator, the sequence that produces credibility fastest.
Learn the business before the data. Spend the first fortnight understanding how the operator makes money, what each vertical does and where the pressure points are. Analysts who invert this produce technically sound work aimed at the wrong questions.
Establish the foundations. Identity resolution, authoritative sources, and which figures people trust, as the next lesson describes.
Find one broken thing and fix it. Every operator has a definition that conflicts, an attribution link that stopped working, or a report that has been wrong for months. Finding and fixing one establishes more credibility than any amount of new analysis.
Answer one question that changes something. Small and consequential beats large and interesting.
Build the metric dictionary if there is not one, which the metrics lesson covers and which is the highest-return infrastructure available.
Read the support transcripts, repeatedly, because they contain the diagnostic information that transactional data does not.
The general principle is that analytical credibility in this industry comes from understanding the business and from producing reliable answers to questions that mattered, in that order. Technical capability is necessary and is rarely the constraint.
A note on tooling
Briefly, because the question arises and the answer is less interesting than people expect.
The tooling in this industry is unremarkable: a warehouse, a transformation layer, a visualisation tool, a notebook environment, and whatever statistical or modelling capability the work requires. The specific products vary and the differences between them matter considerably less than the foundations described in this course.
An operator with excellent tooling and inconsistent definitions produces conflicting answers quickly. One with modest tooling and reliable foundations produces correct answers at a reasonable pace, which is the better position.
The tooling decisions that do matter are whether the semantic layer enforces definitions rather than leaving each analyst to apply them; whether transformation logic is versioned and reviewable rather than existing as ad hoc queries; whether identity resolution is maintained centrally rather than reimplemented; and whether access controls reflect the sensitivity of the underlying data.
Those are architectural rather than product choices, and an operator that gets them right can change tools without difficulty. One that has not will find that its analysis is embedded in whichever tool it happens to use, which is a dependency it did not choose.