Skip to content
iGaming Times

Independent industry intelligence in your inbox. We will email you a link to confirm your subscription, and every newsletter carries a one-click unsubscribe link.

Lesson 3 of 6 · 20 min

Responsible Gambling Risk Models

What harm looks like in the data, the label problem and how to work around it, model choice under an explainability constraint, the decision layer regulators inspect, evaluation without ground truth, fairness and proportionality, affordability.

Fact-checked 23 September 2026 by iGaming Times editorial team · 9 sources

In this lesson

  • Map markers of harm to features in the event data and explain why relative features beat absolute ones
  • Describe the candidate labels for harm and their biases, and combine them defensibly
  • Choose between learned and rule-based models with monotonic constraints and a rule floor
  • Design the decision layer: tiers, marketing suppression, human review, feedback and timeliness
  • Evaluate a harm model with several lenses including specialist case audit

The most consequential model an operator runs

Responsible gambling risk models identify customers who are being harmed by their gambling or are at risk of it. They are the models regulators ask about first, the models whose failure features in many enforcement cases, and the models with the highest stakes for the individuals scored. In June 2026 Petfre (Gibraltar) Limited agreed to pay £900,000 after the Gambling Commission found it did not have sufficient automated processes to identify indicators of harm, or processes to take immediate automated action where strong indicators were identified. A data scientist building one is making decisions that affect whether a person in trouble gets an intervention. This lesson is about doing it properly.

What harm looks like in the data

Gambling harm has a behavioural signature that research has described consistently, and the markers of harm operators are expected to monitor map onto features the data already contains:

Escalation. Rising deposit frequency and amount, rising net loss, relative to the customer's own baseline. One of the most consistently reported signals: an early study of online betting accounts found that frequent, intensive betting with highly variable and increasing stakes in the first month marked out a small high-risk group, 73% of whom later closed their accounts because of gambling problems.

Chasing. Rapid re-deposits after losses, deposits at increasing size within a session, play continuing through a losing streak beyond the customer's usual stop point.

Loss of control indicators. Deposit limit increases, especially repeated ones; cancelled or reversed withdrawals, or regularly failing to withdraw winnings; play at unusual hours, especially late night; very long sessions; multiple payment methods, especially after one fails.

Financial stress indicators. Declined deposits, use of credit where permitted (Britain bans credit card payments for gambling), deposits that cluster around paydays, source-of-funds responses that show strain.

Disclosure. Support contacts mentioning money problems, family, work, or asking about self-exclusion and then not proceeding. Text signals are underused and powerful.

Post-exclusion behaviour. Attempts to reopen an account, new registrations matching an excluded identity, contact asking for the exclusion to be lifted.

Britain turns this list into an obligation: remote operators must monitor customer spend, patterns of spend, time spent gambling, gambling behaviour indicators, customer-led contact, use of gambling management tools and account indicators, and the Commission's guidance names high amounts around paydays, overnight play, changing deposit limits, failed deposits and multiple payment methods among the specific signals.

It matters just as much what harm does not look like: high spend alone is not harm if it is stable and affordable, and low spend is not safety if it is unaffordable for that customer. Relative and contextual features beat absolute ones.

Building the model

The hardest problem is the label. There is no clean ground truth for "this customer was harmed". The candidates each have problems:

  • Self-exclusion is a strong signal but late (the harm preceded it) and biased (many harmed customers never self-exclude).
  • Customer disclosure in support contacts is accurate but rare and unevenly recorded.
  • Screening outcomes from surveys (PGSI-style instruments) are the closest to ground truth but exist for a small sample. The Problem Gambling Severity Index is a nine-item screen developed for general-population surveys rather than clinical use, so even this label is a self-reported risk category, not a clinical diagnosis.
  • Operator interaction outcomes (a trained specialist assessed the customer as at risk) are available at scale and encode the operator's own judgement, with its biases.

Practical approaches combine them: train on the union of strong labels, weight them by confidence, validate against the survey sample, and treat the output as a risk score rather than a diagnosis. Regulators increasingly specify the markers themselves, and marker-based scores built on those lists are transparent by construction; the trade-off against a learned model is accuracy for explainability, and in this domain explainability wins more often than a data scientist would like.

Model choice follows: gradient-boosted trees on the feature layer for a learned model, with monotonic constraints so that more escalation never lowers the score; a weighted rule set for a transparent one; often both, with the rules as a floor the learned model cannot go below.

From score to action

A risk score does nothing. The decision layer attached to it is what regulators inspect:

The design principle: the model must be able to override the commercial models, and not the reverse. A harm score that suppresses marketing, blocks a VIP assignment and caps a bonus is doing its job; a harm score that is one input into a value-optimising decision is not.

Evaluation without a clean label

Since ground truth is partial, evaluation uses several lenses: precision and recall against the strong labels; calibration against the survey sample; the rate at which flagged customers subsequently self-exclude or disclose compared with unflagged customers of similar spend; the proportion of self-excluding customers the model had flagged beforehand (a recall measure that is easy to explain to a regulator); and audit of individual cases by specialists. The last is not optional. A sample of high and low scores reviewed by a person every month catches failure modes no metric shows. British operators must also check that their number of customer interactions is at least in line with the problem gambling rates the Commission publishes for the relevant activity, and be able to demonstrate the outcomes of evaluating their overall approach.

Fairness and proportionality

Harm models can discriminate. A model that learns that a demographic group is more likely to self-exclude will score that group higher and subject its members to more friction. Some of that reflects real base rates and some reflects biased labels (who gets contacted, who discloses). The obligations: measure score distributions and action rates across protected and proxy groups; remove features that are proxies for protected characteristics unless their inclusion is justified and documented, knowing that removing protected characteristics and obvious proxies does not on its own stop a model reproducing discrimination, so outcomes must also be compared across groups; and be able to explain to a regulator why the model treats groups as it does. Some differentiation is expected: Britain's regulator treats 18 to 24 year olds as vulnerable to gambling harm, and its planned financial risk assessments use lower thresholds for under-25s. Proportionality cuts the other way too: an intervention should match the risk, and a model that triggers account closure on moderate signals is failing customers as surely as one that misses severe ones.

Affordability

Affordability checks, required in some markets above spend thresholds, combine harm modelling with financial assessment: is this level of spend consistent with what is known of the customer's means? In Britain, remote operators must run a financial vulnerability check, a public-record search for bankruptcy, county court judgments and similar, once a customer's deposits minus withdrawals exceed £150 in a rolling 30 days; in July 2026 the Commission confirmed that document-free financial risk assessments from credit reference agencies will be introduced in stages: for customers aged 25 and over, starting with the largest operators at £5,000 net deposits in a rolling 24 hours and ending at £1,000 in 24 hours or £3,000 over 90 days, with the stage one timetable still to be confirmed. Data science contributes the spend trajectory and the threshold triggers; the financial assessment uses disclosed income, open banking data where the customer consents, credit reference data where permitted, and the operator's own judgement. The model's role is to make sure the check happens at the right time, not to replace it.

The regulator's questions

A regulator reviewing a harm model will ask: what signals does it use, how was it built, how is it validated, what actions does it drive, how quickly, who reviews the high scores, how are outcomes recorded, can you show me this customer's history and what the model said each day, and what did you do about it. An operator that can answer all of those has a defensible programme. The data scientist's job is to make each answer true. The next lesson turns to the other models with an adversary on the other side: fraud and integrity.

Key terms

Markers of harm
Behavioural signals research has associated with gambling harm: escalation, chasing, limit increases, cancelled withdrawals, late-night play, disclosure.
Monotonic constraint
A model constraint ensuring that more of a risk signal never lowers the score.
PGSI
The Problem Gambling Severity Index, a nine-item screen developed by Ferris and Wynne (2001) for general-population surveys and used in Britain's official gambling statistics. Survey samples scored on it are used to validate harm models.
Marketing suppression
Automatic removal of at-risk customers from promotional campaigns, bonus offers and VIP treatment. In Britain, operators must prevent marketing and new bonus take-up where strong indicators of harm are identified.
Affordability check
An assessment of whether a customer’s spend is consistent with their means, triggered by spend thresholds.

Key takeaways

  • Escalation relative to the customer’s own baseline is one of the most consistent signals of harm.
  • There is no clean ground truth for harm; treat the output as a risk score, not a diagnosis.
  • The model must be able to override the commercial models, and not the reverse.
  • Marketing suppression is not optional: in Britain it is a licence requirement once strong indicators of harm are identified.
  • A harm score that cannot be explained to the customer it affects cannot be defended.

Sources

The legislation, regulator material and research this lesson was checked against.

  1. Licence Conditions and Codes of Practice: SR code 3.4.3 remote customer interaction, 3.4.4 financial vulnerability check, 6.1.2 use of credit cards, Gambling Commission, accessed 2026-09-23
  2. Customer interaction guidance for remote gambling licensees (formal guidance under SR code 3.4.3), Gambling Commission, accessed 2026-09-23
  3. Problem gambling screens, Gambling Commission, accessed 2026-09-23
  4. Commission to introduce Financial Risk Assessments in staged approach (7 July 2026), Gambling Commission, accessed 2026-09-23
  5. QuinnBet (Gibraltar) Limited to pay £609,104 for regulatory failures (20 August 2026), Gambling Commission, accessed 2026-09-23
  6. Petfre (Gibraltar) Limited to pay £900,000 for regulatory failures (30 June 2026), Gambling Commission, accessed 2026-09-23
  7. Guidance on AI and data protection: what about fairness, bias and discrimination?, Information Commissioner's Office, accessed 2026-09-23
  8. How do gamblers start gambling: identifying behavioural markers for high-risk internet gambling (Braverman and Shaffer, European Journal of Public Health, 2012), PubMed, US National Library of Medicine, accessed 2026-09-23
  9. Monotonic constraints, XGBoost documentation, accessed 2026-09-23

Check your understanding

3 questions · answer them all, then check.

  1. 1. Why is self-exclusion an imperfect label for training a harm model?

  2. 2. A learned harm model is 3% more accurate than a transparent rule-based score. Which should drive customer restrictions?

  3. 3. Which recall measure do regulators find most persuasive for a harm model?

Sign in to track your progress through the course.

Cookie Preferences

Choose which cookies you want to accept. Essential cookies are required for the website to function properly.

Required

Necessary for the website to function. Cannot be disabled.

Help us understand how visitors interact with our website.

Used to deliver relevant advertisements and track ad performance.

Remember your preferences and settings for a better experience.

Responsible Gambling Risk Models | Player Data Science