What operators hold
The data landscape in a gambling operator divides into recognisable layers, each with its own origin and reliability characteristics.
Account data. Registration details, verification status and evidence, contact preferences, self-exclusion and limit settings, and the customer's declared circumstances. This originates from the customer and from verification providers. It is generally accurate on the verified elements and can be stale on the declared ones, since a customer's circumstances change and their record does not.
Financial data. Deposits, withdrawals, payment methods, transaction outcomes and balances. This is the most reliable layer, because it is reconciled against actual money movement and because errors surface quickly. It is also the layer where timing differences are most consequential, since authorisation, capture and settlement occur at different moments.
Gameplay data. Every stake, every outcome, every session, every game launched. This is by far the most voluminous layer and it originates from game providers and platform systems rather than from the operator's own records. Its reliability is generally good and its completeness depends on integration quality, particularly where content arrives through aggregators.
Betting data. For sportsbook, every selection, price, stake and settlement, plus the event data underlying them. Settlement corrections and voids mean this layer changes retrospectively more than others do.
Communication data. What was sent, through which channel, and what the customer did with it. Reliability varies by channel, and email engagement data in particular has become less dependable as privacy features have changed how opens are recorded.
Support data. Contacts, categories, transcripts and outcomes. Structurally reliable and analytically underused, since the transcripts contain the richest available information about why customers are unhappy and are almost never examined.
Compliance data. Verification records, monitoring alerts, interactions, interventions and their outcomes. Critical and frequently held in systems disconnected from the analytical environment.
Marketing data. Campaigns, spend, attribution and channel performance, with the attribution caveats described in the Affiliate Marketing course.
Third-party data. Verification providers, payment providers, affiliate platforms, game suppliers and any external enrichment. Each arrives on its own schedule in its own format with its own definitions.
Identity resolution
The foundation, and where errors do the most damage because they propagate silently.
Identity resolution links activity to the correct customer. It sounds trivial and is not, because the same person may appear as several records.
Multiple accounts. Deliberate or accidental, including customers who registered twice and forgot.
Multiple brands. A group operating several brands may have the same person as a customer of each, and whether those are linked determines whether spend can be aggregated.
Multiple devices. Activity across a phone, a tablet and a laptop must resolve to one person, and browser-based tracking makes this harder than it appears.
Separate verticals. Poker and bingo frequently run on separate platforms, and if the identity link is not maintained the customer's activity is fragmented.
Household sharing, where several people use one account, which is a compliance concern as well as an analytical one.
The consequences of imperfect resolution are severe and hard to detect. Customer counts are overstated. Value per customer is understated. Retention appears worse than it is, because a returning customer under a new record counts as a new one and a churned one. Cohort analysis is corrupted. And affordability assessment sees a fraction of a person's actual spend.
The practical guidance for an analyst is to establish how identity resolution works before trusting any customer-level figure, and to test it, since operators frequently believe it works better than it does.
Reliability by source
A working assessment of what to trust.
Financial transactions are the most reliable, being reconciled and consequential.
Gameplay records are generally reliable, with completeness depending on integrations and with occasional gaps when a supplier feed fails.
Account and verification records are reliable at the point of capture and may be stale afterwards.
Compliance records are reliable where the systems capture them properly and are frequently incomplete on the outcome side, since what an intervention achieved is often not recorded.
Communication delivery is reliable; communication engagement is increasingly not, for the privacy reasons noted.
Attribution is the least reliable layer in most operators, for the reasons set out in the Affiliate Marketing course, and analysis depending on it should carry that caveat explicitly.
Third-party enrichment varies enormously and should be assessed individually.
Self-reported data from customers, including declared occupation and income, is the least reliable of all and is nonetheless used in affordability assessment, which is worth knowing when interpreting it.
Latency and timing
An operational characteristic that produces a large proportion of apparent data errors.
Different sources update at different times. A gameplay feed may be near-real-time. Financial settlement may run overnight. Third-party data may arrive daily. Compliance systems may update on their own schedule.
Combining sources without accounting for this produces discrepancies that get investigated as errors when they are timing.
The practical requirements are to know each source's latency and update cycle; to define reporting periods consistently and apply them across sources; to be explicit about the as-at moment of any figure; and to distinguish genuine discrepancy from timing before investigating.
A related issue is retrospective change. Sportsbook settlements can be corrected. Chargebacks arrive weeks later. Bonus adjustments are applied after the fact. Compliance actions may void activity. A figure that was correct when reported may not be correct now, which means historical reports and current queries of the same period can legitimately disagree.
Analysts should know which figures are stable and which are subject to revision, and reporting should indicate it.
Blind spots
What an operator cannot see, which bounds what any analysis can conclude.
Competitor activity. An operator sees its own customers' spend and nothing else. It cannot know a customer's total gambling spend, cannot distinguish a customer who stopped gambling from one who moved, and cannot assess affordability against the full picture. This is the cross-operator visibility problem described in the Payment Operations course, and it limits every conclusion about customer behaviour.
Prospects who did not register. Everyone who arrived and left before creating an account is invisible beyond whatever web analytics captures, which means the top of the funnel is measured differently and less reliably than the rest.
Failed deposits before they reach the operator. Some payment failures occur before the transaction is recorded, which understates the deposit friction described in the Payment Operations course.
Why customers left. Behaviour before departure is visible; motivation is not. Support transcripts and complaints provide partial evidence and the majority of churned customers said nothing.
Life circumstances. Financial changes, health, relationships and everything else determining a person's gambling that has nothing to do with the operator.
Offline context. Whether a customer also gambles in retail, whether they have accounts elsewhere, and what their overall financial position is.
The practical implication is that analysis should be explicit about what it cannot see. A retention analysis conducted on operator data cannot distinguish a customer who quit gambling from one who switched, and those are different findings with different implications.
Building the analytical environment
A brief note on infrastructure, since it determines what is possible.
A warehouse or equivalent consolidating sources into one queryable environment, since analysis across systems is otherwise impractical.
Defined transformation logic, versioned and documented, so that derived measures are reproducible.
A semantic layer where core metrics are defined once and used consistently, which addresses the definitional divergence covered in the next lesson.
Identity resolution maintained as infrastructure rather than reimplemented per analysis.
Access controls, since this data includes verification documents, financial evidence and risk inferences, and the data protection considerations from the Law and Compliance course apply directly.
Documentation covering what each table contains, where it came from and what its known limitations are.
Operators that built this find analysis fast and reliable. Operators that did not find their analysts spending most of their time assembling and reconciling data, which is the most common condition and the largest single drag on analytical productivity in this industry.
Getting to know a new environment
A practical sequence for an analyst arriving at an operator, since the first weeks determine how much subsequent work is wasted.
Find out how identity resolution works. Ask, then test. Take a handful of customers and establish whether their activity across devices, brands and verticals resolves correctly. This is the foundation and it is frequently weaker than anyone believes.
Establish the authoritative source for revenue. There will be several candidates. Find out which one finance uses and why the others differ.
Trace one number end to end. Pick a figure from a report and follow it back through the transformations to the source. This reveals the lineage, the assumptions and the places where logic is undocumented.
Find the gaps. Which sources are not in the warehouse, which are stale, and which have known reliability problems that people work around without documenting.
Ask what nobody trusts. Every operator has figures that circulate and that experienced people quietly discount. Finding out which and why is faster than discovering it independently.
Read support transcripts. An hour of this teaches more about what customers actually experience than a month of transactional analysis.
Sit with the operating teams. Trading, CRM, payments and compliance each have a working understanding of the data that is not written down anywhere.
Check the definitions. Which measures have agreed definitions, which have several, and which have none.
Doing this before producing analysis prevents the standard first-month failure, which is delivering a confident answer built on an attribution link that has been broken since a platform migration two years ago.
Data and compliance
A dimension analysts frequently encounter late, and one where the obligations are real.
The data described in this lesson includes identity documents, financial evidence, complete transaction histories, behavioural records and, where models produce them, inferences about whether a customer may be experiencing gambling harm.
The consequences for analytical work, drawing on the Law and Compliance course.
Access should be proportionate. Broad analyst access to verification documents and financial evidence is difficult to justify, and access controls should reflect the sensitivity rather than the convenience.
Purpose matters. Data collected under a legal obligation for verification is not automatically available for marketing analysis, and the lawful basis for each processing purpose should be established rather than assumed.
Risk inferences are sensitive. An output indicating that a customer may be experiencing harm relates to their health, and it must not become a marketing input. Systems should make that structurally impossible rather than prohibiting it in policy.
Retention applies to analytical copies. Data extracted into an analytical environment inherits the retention obligations of its source, and analytical environments accumulate extracts that outlive their justification.
Automated decisions carry obligations. A model output that materially affects a customer engages the requirements described in the data protection lesson, including explainability and a route to contest.
The practical position for an analyst is that the environment they work in is holding unusually sensitive material about identifiable people, that the obligations are not the compliance team's problem alone, and that a working understanding of them prevents building things that cannot be deployed.
Common data problems
A catalogue of the failures that recur, since recognising them saves diagnostic time.
Attribution that stopped working. A platform change, a tracking update or a consent change breaks the link between acquisition source and customer, and it is invisible until someone asks a channel-level question.
Duplicate customers from imperfect identity resolution, inflating counts and deflating per-customer values.
Missing gameplay from a supplier feed that failed silently, producing a gap that looks like a behaviour change.
Timezone inconsistency, where sources record in different zones and daily aggregations disagree at the boundaries.
Currency handling, where conversion is applied at different points or different rates across sources.
Retrospective adjustments applied to historical periods, so a report run today for last month differs from the one run last month.
Test and internal accounts included in production figures, which at small volumes materially distorts them.
Bonus funds counted as deposits or vice versa, which distorts both deposit and revenue measures.
Voided activity included or excluded inconsistently.
Soft-deleted records that remain in tables and are picked up by queries that did not exclude them.
Each of these has produced confidently wrong analysis somewhere. The general defence is to sanity-check magnitudes against known values, to compare a derived figure against an authoritative source before relying on it, and to be suspicious of any result that is surprising in a convenient direction.
Beyond operator data
A closing note on the sources that address the blind spots.
Market research and panel data provides estimates of category size, competitive share and player behaviour across operators. It is directional rather than precise and it is the only route to questions operator data cannot answer.
Published competitor reporting for listed operators provides revenue, market mix and, increasingly, safer gambling metrics, with the comparability caveats set out in the Operations Strategy course.
Regulator publications provide market-level data in several jurisdictions, including participation and, in some, aggregate operator returns.
Affiliate and comparison content indicates how the operator is described publicly and what customers say, which is qualitative evidence with commercial consequences.
App store and review data provides sentiment and version-level feedback.
Open banking data, with customer consent, provides the financial picture operator data lacks, with the data protection considerations described in the Payment Operations course.
Survey research, conducted directly, answers motivational questions no behavioural data can.
The general point is that operator data answers what customers did with this operator and nothing else. Questions about the market, about competitive position, about total customer spend and about why people behave as they do require sources from outside, and analysis that presents an operator-level finding as a market-level one is overreaching in a way that is easy to miss.
What the landscape determines
A closing summary of why this lesson comes before the technique.
The analysis available to an operator is bounded by what it holds, how reliably it holds it, and whether the pieces connect.
Cohort contribution analysis requires identity resolution, cost attribution at customer level and persistent acquisition attribution. Many operators lack at least one.
Channel performance requires attribution that survives the customer's whole life, which is the most commonly broken link.
Cross-vertical customer value requires the verticals to be integrated, which for poker and bingo frequently they are not.
Affordability and harm monitoring requires complete visibility of a customer's activity, which fragmented systems prevent.
Incremental measurement requires the ability to hold out a group and track them, which requires persistent assignment.
Any of it over multi-year horizons requires definitions that have been stable or versioned.
The practical consequence is that an analyst asked for something the environment cannot support should say so rather than producing an approximation presented as an answer. The approximation will be used as though it were reliable, and the gap will surface later in a decision that depended on it.
Establishing what the environment can actually support, before promising what it will deliver, is the most useful thing an analyst does in their first month.