Models as regulated decisions
A model that decides which customers get a bonus, which get a limit, which withdrawals are held and which accounts are flagged for harm is making regulated decisions at scale. Regulators have caught up with this: they ask how models are built, validated, governed and explained, and data protection law in Britain, the EU and elsewhere gives individuals rights over automated decisions that significantly affect them. This lesson sets out what a defensible model governance framework contains, and it applies to every model in the course.
The inventory
The first control is knowing what exists. A model inventory lists every model in production and in development: its purpose, its owner, the decisions it drives, the data it uses, its version, its last validation, its risk tier. Operators that cannot produce this list on request have, in the regulator's eyes, no governance. Banking supervisors put this first: the Prudential Regulation Authority's model risk statement opens with model identification and model risk classification. Risk tiering matters because the controls scale with it: a game recommendation model is low tier; a harm model or a withdrawal-hold model is high tier and carries every control below.
Documentation
Each model has a document, kept current, covering: the business problem and the decision the model supports; the data sources, features and their definitions; the training population and period; the modelling approach and its alternatives; the validation results, including performance across customer groups; known limitations; the decision thresholds and the actions attached; the monitoring in place; and the review history. Model cards, proposed by Mitchell and colleagues in 2018, are a lightweight standard for this. The test of adequate documentation is whether a competent person who has never seen the model could understand, from the document alone, what it does and whether it is doing it.
Validation and independent review
Before a high-tier model goes live, someone who did not build it reviews it: the data pipeline for leakage and lookahead, the label construction, the evaluation methodology, the fairness analysis, the threshold choice, and the decision layer. In larger operators this is a separate model risk function; in smaller ones it is a peer who is not on the project. The review is recorded, and its findings are addressed before deployment. Banking supervisors have expected this for years: US model risk guidance, first issued as SR 11-7 in 2011 and replaced by revised interagency guidance in April 2026, calls for effective challenge by people with sufficient independence to stay objective, and the PRA's statement, in force for UK banks with internal model approval since May 2024, makes independent model validation one of its five principles. Gambling regulators are moving the same way: in Britain, the remote customer interaction code (SR 3.4.3) requires operators to act on strong indicators of harm through automated processes, manually review their operation in each customer's case, let the customer contest any automated decision that affects them, and evaluate the effectiveness of their approach.
Explainability
Two audiences need explanations. The regulator needs to understand how the model works in general and why it treats customers as it does. The customer, and the specialist speaking to them, needs to know why this decision was made about this person. Both are served by:
- Preferring models that can be explained (constrained trees, additive models, transparent rule sets) for high-tier decisions, and accepting an accuracy cost where necessary.
- Producing per-decision explanations (which features drove this score, in plain language) and storing them with the decision.
- Keeping the full history: the model version, the feature values and the score for every decision, retained for as long as the regulator can ask about it.
- Being able to reproduce a past decision: rerun the model as it was, on the data as it was.
A harm score that cannot be explained to the customer it affects is a harm score the operator cannot defend. A withdrawal hold whose reason the operator cannot state is a complaint the operator will lose.
Fairness
Models learn from history, and history contains bias. The obligations across the course's models:
- Do not use protected characteristics as features unless the use is justified, lawful and documented (age is used in harm models, and the Equality Act 2010 permits age-based treatment shown to be a proportionate means of achieving a legitimate aim; ethnicity is not).
- Test for proxies: postcode, device, payment method and name can stand in for protected characteristics, and their effect must be measured. The ICO warns that removing protected characteristics from the inputs is unlikely to be enough, because features such as postcode can act as proxies for race.
- Measure outcomes across groups: score distributions, action rates, false positive and false negative rates. Disparities must be explained or removed.
- Consider the label: if the label reflects who the operator chose to contact, the model reproduces that choice.
- Re-test after every retrain.
The fairness question is sharpest in the commercial models, where a model optimising value may lawfully learn to favour groups that lose more, and the operator has to decide whether it wants a model that does that. Usually it should not, and the reason is not only ethics: a regulator reviewing acquisition targeting that skews toward vulnerable groups will see a harm case.
Data protection and automated decisions
Under the EU GDPR, individuals have the right not to be subject to a decision based solely on automated processing that has legal or similarly significant effects (Article 22); where such a decision is permitted because it is necessary for a contract or rests on explicit consent, the controller must provide safeguards, at least the right to human intervention, to express their view and to contest the decision. In the UK, the Data (Use and Access) Act 2025 replaced Article 22 of the UK GDPR, in force from 5 February 2026, with Articles 22A to 22D: solely automated significant decisions are permitted more widely, but not on special category data except under narrow conditions, and every one must carry safeguards: information about the decision, the chance to make representations, human intervention and the ability to contest it. Account closures, withdrawal holds and play restrictions are likely to be significant effects: the ICO's draft guidance gives an automated freeze on a bank account as an example of a decision that can be significant, because it affects the person's financial circumstances. The design consequences: a human in the loop for those decisions, a route for the customer to challenge, and transparency about the fact that automated processing is used. Privacy notices, data protection impact assessments for new models (the EU GDPR requires one for systematic and extensive profiling on which significant decisions are based), and data minimisation (using the features that are needed, not all the features that exist) are the routine obligations.
Monitoring in production
Every model degrades. Monitoring covers:
- Input drift: the feature distributions moving away from the training distribution, which happens with every product change, market entry and season.
- Output drift: the score distribution shifting, so that thresholds set on the old distribution act on a different population.
- Performance: where labels arrive (churn, fraud confirmations, self-exclusions), the model's accuracy over time.
- Decision volumes: how many customers are receiving each action, week by week, with alerts on jumps.
- Fairness metrics: re-computed on a schedule.
Each metric has an owner, a threshold and a response, and high-tier models have an automatic fallback (a conservative rule set, or human review of everything) when they breach.
The conflict, resolved in governance
The course has returned repeatedly to the conflict between the commercial models and the harm model. Governance is where it is resolved. The framework states, in writing, that the harm model's output constrains the commercial models: customers above a risk threshold are excluded from targeting, bonusing, VIP assignment and recommendation optimisation, automatically, and that exclusion cannot be overridden by commercial staff. In Britain part of this is already a licence requirement: operators must prevent marketing and the take-up of new bonus offers where strong indicators of harm have been identified. It states who can change the threshold and how. It states that experiments carry the harm score as a guardrail. And it states that the responsible gambling lead has authority over the decision layer for harm, independent of the commercial owners of the other models.
An operator with those statements, and evidence that they are followed, has a data science function a regulator can trust. An operator with excellent models and no such statements has built the machinery of an enforcement case. The whole discipline, in the end, is the same one that runs through every course on this hub: know what you are doing, record it, apply it consistently, and be able to explain it to the person it affects.