The most consequential model an operator runs
Responsible gambling risk models identify customers who are being harmed by their gambling or are at risk of it. They are the models regulators ask about first, the models whose failure features in many enforcement cases, and the models with the highest stakes for the individuals scored. In June 2026 Petfre (Gibraltar) Limited agreed to pay £900,000 after the Gambling Commission found it did not have sufficient automated processes to identify indicators of harm, or processes to take immediate automated action where strong indicators were identified. A data scientist building one is making decisions that affect whether a person in trouble gets an intervention. This lesson is about doing it properly.
What harm looks like in the data
Gambling harm has a behavioural signature that research has described consistently, and the markers of harm operators are expected to monitor map onto features the data already contains:
Escalation. Rising deposit frequency and amount, rising net loss, relative to the customer's own baseline. One of the most consistently reported signals: an early study of online betting accounts found that frequent, intensive betting with highly variable and increasing stakes in the first month marked out a small high-risk group, 73% of whom later closed their accounts because of gambling problems.
Chasing. Rapid re-deposits after losses, deposits at increasing size within a session, play continuing through a losing streak beyond the customer's usual stop point.
Loss of control indicators. Deposit limit increases, especially repeated ones; cancelled or reversed withdrawals, or regularly failing to withdraw winnings; play at unusual hours, especially late night; very long sessions; multiple payment methods, especially after one fails.
Financial stress indicators. Declined deposits, use of credit where permitted (Britain bans credit card payments for gambling), deposits that cluster around paydays, source-of-funds responses that show strain.
Disclosure. Support contacts mentioning money problems, family, work, or asking about self-exclusion and then not proceeding. Text signals are underused and powerful.
Post-exclusion behaviour. Attempts to reopen an account, new registrations matching an excluded identity, contact asking for the exclusion to be lifted.
Britain turns this list into an obligation: remote operators must monitor customer spend, patterns of spend, time spent gambling, gambling behaviour indicators, customer-led contact, use of gambling management tools and account indicators, and the Commission's guidance names high amounts around paydays, overnight play, changing deposit limits, failed deposits and multiple payment methods among the specific signals.
It matters just as much what harm does not look like: high spend alone is not harm if it is stable and affordable, and low spend is not safety if it is unaffordable for that customer. Relative and contextual features beat absolute ones.
Building the model
The hardest problem is the label. There is no clean ground truth for "this customer was harmed". The candidates each have problems:
- Self-exclusion is a strong signal but late (the harm preceded it) and biased (many harmed customers never self-exclude).
- Customer disclosure in support contacts is accurate but rare and unevenly recorded.
- Screening outcomes from surveys (PGSI-style instruments) are the closest to ground truth but exist for a small sample. The Problem Gambling Severity Index is a nine-item screen developed for general-population surveys rather than clinical use, so even this label is a self-reported risk category, not a clinical diagnosis.
- Operator interaction outcomes (a trained specialist assessed the customer as at risk) are available at scale and encode the operator's own judgement, with its biases.
Practical approaches combine them: train on the union of strong labels, weight them by confidence, validate against the survey sample, and treat the output as a risk score rather than a diagnosis. Regulators increasingly specify the markers themselves, and marker-based scores built on those lists are transparent by construction; the trade-off against a learned model is accuracy for explainability, and in this domain explainability wins more often than a data scientist would like.
Model choice follows: gradient-boosted trees on the feature layer for a learned model, with monotonic constraints so that more escalation never lowers the score; a weighted rule set for a transparent one; often both, with the rules as a floor the learned model cannot go below.
From score to action
A risk score does nothing. The decision layer attached to it is what regulators inspect:
- Thresholds and tiers. Low, moderate, high, with a defined action for each: an automated safer-gambling message, a prompt to set limits, a trained-person contact, restrictions on bonuses and marketing, a mandatory affordability check, restrictions on play, closure.
- Marketing suppression. Customers above a threshold are removed from promotional campaigns and VIP treatment automatically. In Britain this is a licence requirement, not a design choice: licensees must prevent marketing and the take-up of new bonus offers where strong indicators of harm have been identified.
- Human review. High scores go to trained specialists who make and record a judgement. For the strongest signals, British rules put automation first: strong indicators of harm must be acted on through automated processes, whose operation is manually reviewed in each customer's case, and the customer must be able to contest an automated decision. The model prioritises and can trigger immediate protective action; the person reviews and decides what follows.
- Feedback. The outcome of every interaction (customer engaged, set a limit, declined contact, continued escalating) feeds back as a label.
- Timeliness. Harm can escalate within hours. In a 2026 British case, one customer's stakes escalated after a large win to more than £215,000 in a day, and this was not identified until a report was produced the following day; the Commission's guidance says timely action will in some cases mean automated, real-time measures. A daily batch score is a floor, not a target; near-real-time scoring for the fastest-moving signals (chasing within a session) is where mature operators are.
The design principle: the model must be able to override the commercial models, and not the reverse. A harm score that suppresses marketing, blocks a VIP assignment and caps a bonus is doing its job; a harm score that is one input into a value-optimising decision is not.
Evaluation without a clean label
Since ground truth is partial, evaluation uses several lenses: precision and recall against the strong labels; calibration against the survey sample; the rate at which flagged customers subsequently self-exclude or disclose compared with unflagged customers of similar spend; the proportion of self-excluding customers the model had flagged beforehand (a recall measure that is easy to explain to a regulator); and audit of individual cases by specialists. The last is not optional. A sample of high and low scores reviewed by a person every month catches failure modes no metric shows. British operators must also check that their number of customer interactions is at least in line with the problem gambling rates the Commission publishes for the relevant activity, and be able to demonstrate the outcomes of evaluating their overall approach.
Fairness and proportionality
Harm models can discriminate. A model that learns that a demographic group is more likely to self-exclude will score that group higher and subject its members to more friction. Some of that reflects real base rates and some reflects biased labels (who gets contacted, who discloses). The obligations: measure score distributions and action rates across protected and proxy groups; remove features that are proxies for protected characteristics unless their inclusion is justified and documented, knowing that removing protected characteristics and obvious proxies does not on its own stop a model reproducing discrimination, so outcomes must also be compared across groups; and be able to explain to a regulator why the model treats groups as it does. Some differentiation is expected: Britain's regulator treats 18 to 24 year olds as vulnerable to gambling harm, and its planned financial risk assessments use lower thresholds for under-25s. Proportionality cuts the other way too: an intervention should match the risk, and a model that triggers account closure on moderate signals is failing customers as surely as one that misses severe ones.
Affordability
Affordability checks, required in some markets above spend thresholds, combine harm modelling with financial assessment: is this level of spend consistent with what is known of the customer's means? In Britain, remote operators must run a financial vulnerability check, a public-record search for bankruptcy, county court judgments and similar, once a customer's deposits minus withdrawals exceed £150 in a rolling 30 days; in July 2026 the Commission confirmed that document-free financial risk assessments from credit reference agencies will be introduced in stages: for customers aged 25 and over, starting with the largest operators at £5,000 net deposits in a rolling 24 hours and ending at £1,000 in 24 hours or £3,000 over 90 days, with the stage one timetable still to be confirmed. Data science contributes the spend trajectory and the threshold triggers; the financial assessment uses disclosed income, open banking data where the customer consents, credit reference data where permitted, and the operator's own judgement. The model's role is to make sure the check happens at the right time, not to replace it.
The regulator's questions
A regulator reviewing a harm model will ask: what signals does it use, how was it built, how is it validated, what actions does it drive, how quickly, who reviews the high scores, how are outcomes recorded, can you show me this customer's history and what the model said each day, and what did you do about it. An operator that can answer all of those has a defensible programme. The data scientist's job is to make each answer true. The next lesson turns to the other models with an adversary on the other side: fraud and integrity.