Skip to content
iGaming Times

Independent industry intelligence in your inbox. We will email you a link to confirm your subscription, and every newsletter carries a one-click unsubscribe link.

Lesson 1 of 6 · 16 min

The Data and the Questions

What an operator holds and what it does not, the five families of questions, the feature layer and why relative features win, gambling-specific data traps, and the stack in outline.

Fact-checked 23 September 2026 by iGaming Times editorial team · 7 sources

In this lesson

  • Describe the event streams and customer record an operator holds and the blind spots in them
  • Name the five model families and the conflict between value and harm
  • Design features relative to the customer’s own baseline
  • Recognise settlement timing, bonus accounting, verification changes and multiple identities as pipeline traps

What an operator actually has

An online gambling operator holds one of the richest behavioural datasets in consumer business. Every deposit, every bet, every spin, every session, every login, every support contact, every bonus, every limit set and every withdrawal is recorded with a timestamp and tied to an identified individual: in Britain, for example, operators must obtain and verify a customer's name, address and date of birth before that customer is permitted to gamble (licence condition 17). Because every spin and every bet is a separate event, even a modest customer base produces millions of records. That data is the raw material for every model in this course, and the first discipline is knowing what it contains and what it does not.

The event streams. Wagers (stake, product, market or game, odds or RTP, outcome, settlement time), transactions (deposits, withdrawals, method, success or failure, bonus credits and debits), sessions (device, channel, duration, pages or games), and customer relationship events (registrations, verification outcomes, contacts, complaints, limits, exclusions, marketing sent and opened).

The customer record. Verified identity attributes (age, address, and in some markets the results of affordability checks, such as Britain's financial vulnerability check against public records of bankruptcy and county court judgments once net deposits pass £150 in a rolling 30 days), registration source and affiliate, product preferences, limits and responsible gambling status, risk classifications.

What is missing. The customer's play with other operators, their financial situation beyond what they have disclosed, and their intent. Every model is built on a partial view of a person, and the models that go wrong are usually the ones that forget it.

The questions data science answers

The work divides into five families, and the rest of this course takes them in turn.

Value. What is a customer worth, now and in future? Lifetime value prediction drives acquisition spend, bonus allocation and VIP management.

Retention. Who is about to leave, and can we change it? Churn prediction and the interventions it triggers are among the largest commercial applications.

Harm. Who is being harmed, or is at risk? Responsible gambling risk models are now a regulatory expectation and the most consequential models an operator runs: in Great Britain, remote licensees must monitor customer activity for indicators of harm from the point an account is opened and act on strong indicators through automated processes (social responsibility code 3.4.3).

Integrity. Who is not who they seem, or not doing what they seem? Fraud, bonus abuse, multi-accounting, money laundering, collusion, and, on the sports side, suspicious betting.

Product. What should we show this customer, and does the change we made work? Recommendation, personalisation and experimentation.

The families share infrastructure, features and governance, and they conflict: the value model wants to identify the customers who will lose the most, and the harm model wants to protect them. Managing that conflict is a design responsibility, not an afterthought, and lesson three returns to it.

The feature layer

Models are built on features: quantities computed from the raw events that describe a customer's behaviour over windows of time. A mature operator maintains a feature store with a large library of them, recomputed daily or in near real time. The ones that recur across every family:

  • Deposit frequency, amount, trend and method mix
  • Net loss over windows (day, week, month) and its trend
  • Session count, duration, time-of-day distribution
  • Product mix and its change
  • Bet size relative to the customer's own history
  • Bonus uptake and wagering behaviour
  • Withdrawal requests, cancellations and reversals
  • Limit changes, especially increases
  • Contact with support, and its sentiment
  • Days since last activity, and the customer's own inter-session distribution

Feature engineering is where domain knowledge enters. A feature like "ratio of this week's net loss to the customer's median weekly loss over the prior three months" carries more signal than either number alone, because it captures escalation relative to the customer's own baseline. The best features in gambling are almost always relative to the individual, not to the population. Regulators ask for both views: the Gambling Commission's guidance lists escalation in deposit levels and increasing session length alongside amounts spent compared with other customers among the indicators of harm operators must use, and notes that a higher percentage of overnight gamblers were found to be problem gamblers.

Data quality and the traps

Gambling data has specific pitfalls that catch data scientists arriving from other industries:

Settlement timing. A bet placed on Saturday may settle on Sunday, or, for an ante-post bet, months later. Revenue attributed to the wrong period distorts every window-based feature.

Bonus accounting. Wagering with bonus funds is not the same as wagering with cash, and models that conflate them misjudge both value and risk.

Product heterogeneity. A slot spin and an accumulator bet are both wagers, with completely different variance and margin. Aggregating them naively hides the signal.

Verification changes. A customer's attributes change as verification progresses; features must be computed against what was known at the time, not what is known now, or every backtest suffers lookahead. This is a form of data leakage, where information not available at prediction time produces overly optimistic performance estimates; feature stores address it by generating point-in-time correct feature sets.

Voids and corrections. Settled bets get resettled, deposits get reversed, accounts get merged. Pipelines that are not idempotent produce features that drift; the standard advice is that a pipeline task should produce the same outcome on every re-run, for example by upserting rather than inserting.

Multiple identities. The same person may hold several accounts, legitimately across brands or illegitimately within one; models that treat accounts as people are wrong in both directions. In Britain, operators must identify separate accounts held by the same individual and base customer interaction decisions on behaviour across all of them (social responsibility code 3.9.1).

Infrastructure in outline

The typical stack: event streams from the platform into a warehouse or lakehouse; a feature store computing daily and streaming features; a model registry with versioning; batch scoring for value and churn, streaming or near-real-time scoring for harm and fraud; a decision layer that turns scores into actions (a message, a bonus, a limit, a block, a review); and a monitoring surface for drift and performance. The decision layer is where the regulatory obligations attach, and lesson six is about governing it.

The stance of this course

Everything that follows assumes the models will be used on real people whose money, and sometimes whose welfare, depends on the output. That imposes standards a recommendation engine for a streaming service does not carry: explainability to a regulator, fairness across customer groups, the ability to reconstruct why a decision was made long afterwards (British operators must record the rationale for decisions taken after a financial vulnerability check, and the Money Laundering Regulations 2017, which apply to casinos among other businesses, require records sufficient to reconstruct transactions for at least five years after the business relationship ends), and a clear line between what the model suggests and what a human decides. In Britain, where automated processes act on strong indicators of harm, the licensee must manually review their operation in each customer's case and allow the customer to contest the decision. The next lesson starts with the models the business most wants, value and retention, and the one after with the models the regulator most wants.

Key terms

Feature store
A system that manages the behavioural quantities models are built on, keeping point-in-time correct history for training and serving current values to live models, updated daily or in near real time.
Decision layer
The component that turns model scores into actions: a message, a bonus, a limit, a hold, a review.
Lookahead
Using information in training or backtesting that was not available at the time of the decision, a form of data leakage that makes performance look better than it will be in production.
Idempotent pipeline
A data pipeline that produces the same result when re-run, so resettlements and reversals do not cause drift.

Key takeaways

  • Every model is built on a partial view of a person; the models that go wrong forget it.
  • The best features in gambling are relative to the individual, not the population.
  • Features must be computed against what was known at the time or every backtest suffers lookahead.
  • The value model wants to find the customers who will lose the most; the harm model wants to protect them.
  • The decision layer is where regulatory obligations attach.

Sources

The legislation, regulator material and research this lesson was checked against.

  1. Licence Conditions and Codes of Practice: licence condition 17, social responsibility codes 3.4.3, 3.4.4 and 3.9.1, Gambling Commission, accessed 2026-09-23
  2. Customer interaction guidance for remote gambling licensees, requirement 5: indicators of harm, Gambling Commission, accessed 2026-09-23
  3. The Money Laundering, Terrorist Financing and Transfer of Funds (Information on the Payer) Regulations 2017, regulation 40: record-keeping, legislation.gov.uk, accessed 2026-09-23
  4. The Money Laundering, Terrorist Financing and Transfer of Funds (Information on the Payer) Regulations 2017, regulation 8: application, legislation.gov.uk, accessed 2026-09-23
  5. Feast documentation: introduction, Feast (open-source feature store project), accessed 2026-09-23
  6. Common pitfalls and recommended practices: data leakage, scikit-learn, accessed 2026-09-23
  7. Best practices: creating a task, Apache Airflow, accessed 2026-09-23

Check your understanding

3 questions · answer them all, then check.

  1. 1. Which feature is likely to carry the most signal about escalation?

  2. 2. A feature uses the customer’s current verification status to score a decision made last year. The problem is:

  3. 3. Why should slot spins and accumulator bets not be aggregated naively into one wager feature?

Sign in to track your progress through the course.

Cookie Preferences

Choose which cookies you want to accept. Essential cookies are required for the website to function properly.

Required

Necessary for the website to function. Cannot be disabled.

Help us understand how visitors interact with our website.

Used to deliver relevant advertisements and track ad performance.

Remember your preferences and settings for a better experience.