Skip to content
iGaming Times

Independent industry intelligence in your inbox. We will email you a link to confirm your subscription, and every newsletter carries a one-click unsubscribe link.

Lesson 2 of 6 · 18 min

Value and Retention Models

Lifetime value under extreme skew, cohort, probabilistic and supervised approaches, the early-life problem, churn as silence, uplift modelling for interventions, segmentation and the account manager question, personalisation with constraints.

Fact-checked 23 September 2026 by iGaming Times editorial team · 9 sources

In this lesson

  • Choose between cohort curves, buy-till-you-die models and supervised prediction for a given LTV decision
  • Handle skew by predicting log revenue, two-stage models or quantile regression
  • Define churn relative to the customer’s own rhythm and build a survival or classification model on engagement trend
  • Explain why uplift models, not churn scores, should target retention interventions
  • Constrain personalisation with harm-model inputs rather than a filter applied afterwards

Lifetime value: the model everything else is priced against

Customer lifetime value (LTV) is the expected net revenue from a customer over their relationship with the operator, discounted if the business is careful. It sets the ceiling on what the operator can pay to acquire a customer, the size of bonus it can afford to offer, and which customers get an account manager. Get it wrong and the marketing budget is misallocated in the direction the error points.

The naive approach, average revenue per user times average lifetime, fails for the reason that defines gambling: the distribution of value is extremely skewed. A small share of customers produces a large share of the revenue: on British Columbia's government-run PlayNow site, researchers found that 5% of players produced 46% of revenue and the most active 20% placed 82% of bets. The average describes nobody. Useful LTV models predict at the individual level and preserve the distribution.

Three approaches in practice:

Cohort curves. Group customers by acquisition month and channel, observe how each cohort's cumulative net revenue develops, and project the curve. Simple, robust, and the right tool for channel-level budgeting. It does not predict individuals.

Probabilistic models. The buy-till-you-die family (Pareto/NBD, BG/NBD and their extensions) models transaction frequency and an unobserved dropout as latent processes, and forecasts each customer's future transactions from two summaries of their history, recency and frequency; a companion gamma-gamma model of spend per transaction adds a customer-level forecast of spend, and the two together give an individual expected future value. They can be applied to gambling because deposit behaviour is the non-contractual repeat behaviour they were built for, where customers can transact at any time and the moment they stop is never observed, and they return a probability that each customer is still active rather than only a point estimate.

Supervised prediction. Gradient-boosted trees or neural networks trained to predict net revenue over the next N months from the feature layer. These win on accuracy for customers with enough history, lose on cold-start customers, and require careful handling of the skew: predicting log revenue, or a two-stage model (will they be active, and if so how much), or quantile regression so the model says something about the tail. One published version of the two-stage idea is the zero-inflated lognormal loss, which models the chance of zero future value and a heavy-tailed amount together.

A common arrangement is a cohort model for budgeting and a supervised model for individual decisions, reconciled against each other.

The early-life problem

The decisions with the most money attached (how much to bonus a new customer, whether to assign them an account manager, which channel to scale) have to be made before the customer has generated the data that predicts their value. Early-life LTV models use what is available in the first days: registration source and affiliate, first deposit size and method, first product, first-week activity pattern, device, and the verification attributes. These are surprisingly predictive, and they are also where the fairness questions arise: a model that learns that customers from a particular postcode or age band are worth more will allocate marketing accordingly, and that can be lawful and still indefensible. Lesson six returns to this.

Churn: the largest commercial application

Churn in gambling is not cancellation; it is silence. It is a non-contractual setting, in the modelling literature's term, one where the time at which customers become inactive is unobserved. A customer who has not deposited or played for a period is gone, and the operator finds out by their absence. Churn modelling therefore starts with defining the event: inactivity for N days, where N depends on the customer's own rhythm (a weekly bettor who misses three weeks has churned; a monthly one has not).

The model predicts, for each active customer, the probability of churning in the next window, from features that describe engagement trend: declining session frequency, longer gaps than the customer's own norm, falling deposit amounts, a withdrawal of the full balance, a product switch, a failed deposit, a losing streak beyond the customer's usual tolerance, a bonus that expired unused. Gradient-boosted models on these features perform well; survival models (predicting time to churn rather than churn in a window) perform better where the business needs to prioritise by urgency.

The model's output is only useful with an intervention attached, and the interventions have to be tested. A retention bonus sent to every customer with a high churn score wastes money on those who would have stayed and on those who will leave regardless; the customers worth intervening on are the ones whose behaviour the intervention changes. That is an uplift modelling problem: predicting not who will churn but who will stay because of the intervention. Uplift models need experiments (a randomised control group that receives nothing). A study published in the Journal of Marketing Research in 2018, using field data from a wireless services provider and a membership organisation, found that the customers at highest risk of churning are not necessarily the best targets: simulations on both datasets showed that targeting on the likely response to the intervention reduced churn more than targeting on risk.

Segmentation and the account manager question

Beyond individual scores, operators segment customers into groups that get different treatment: recreational, regular, high-value, at-risk, dormant. Clustering on the feature layer produces segments that describe behaviour; business rules on value and risk produce segments that drive treatment. The high-value segment is where account management sits, and it is where the value model and the harm model collide most directly: the customers a value model identifies as most valuable are, disproportionately, the customers a harm model identifies as most at risk. In one study of online gambling subscribers, between 38% and 67% of the "vital few" who accounted for most activity screened positive for gambling-related problems, against 24% to 35% of the rest. An operator that lets the value model assign account managers without the harm model's veto is building the fact pattern of a future enforcement case. In Britain, the Gambling Commission's guidance on high value customer schemes, which covers personal account management, sets minimum safeguards and says a licensee that cannot meet them should not operate such schemes.

Personalisation

Recommendation models suggest games, markets and content: collaborative filtering on play history, content-based similarity on game attributes, and contextual bandits that learn what to show by testing. They increase engagement, which is the point, and they also increase the responsibility: a recommendation engine that learns to show a losing customer the highest-variance slot because it maximises the next session's revenue is doing what harm-prevention rules are written to stop: in Britain, remote operators must monitor for indicators of harm from account opening and prevent marketing and the take-up of new bonus offers where strong indicators of harm have been identified. Personalisation models need constraints from the harm model as inputs, not as a filter applied afterwards.

Measuring the models

Value models are evaluated on calibration (does the predicted value match realised value across deciles, the decile chart used in the LTV literature) and on the decisions they drive (did the acquisition spend they justified pay back). Churn models are evaluated on precision and recall at the operating threshold, on lift over a baseline, and on the incremental retention the interventions produced in experiments. Every model has a business metric it is supposed to move, and the honest report shows the model's accuracy alongside that metric's movement, because a model can be accurate and useless.

The next lesson takes up the models the regulator asks about first.

Key terms

Lifetime value (LTV)
The expected net revenue from a customer over their relationship with the operator.
BG/NBD
The beta-geometric/negative binomial distribution model (Fader, Hardie and Lee, 2005) of repeat transactions and unobserved dropout, which forecasts each customer's future transactions from their recency and frequency.
Uplift model
A model predicting the change in outcome caused by an intervention, trained on randomised treatment and control groups.
Survival model
A model of time until an event, such as churn, rather than its occurrence in a window.
Contextual bandit
A learning algorithm that chooses what to show each user from information about the user and the options, balancing testing untried options against exploiting what has worked, and adapting from feedback.

Key takeaways

  • The average customer describes nobody; useful LTV models predict individually and preserve the distribution.
  • Early-life models are predictive and are where fairness questions arise.
  • The customers worth intervening on are the ones whose behaviour the intervention changes, who are not necessarily the ones at highest risk of churning.
  • The customers a value model finds most valuable are, disproportionately, the ones a harm model finds most at risk.
  • A model can be accurate and useless; report its accuracy alongside the business metric it was meant to move.

Sources

The legislation, regulator material and research this lesson was checked against.

  1. Licence Conditions and Codes of Practice, SR code 3.4.3: Remote customer interaction, Gambling Commission, accessed 2026-09-23
  2. High Value Customers: Industry guidance, Introduction, Gambling Commission, accessed 2026-09-23
  3. "Counting Your Customers" the Easy Way: An Alternative to the Pareto/NBD Model (Fader, Hardie and Lee, Marketing Science, 2005), INFORMS, author copy at brucehardie.com, accessed 2026-09-23
  4. Does Pareto rule Internet gambling? Problems among the "vital few" and "trivial many" (Tom, LaPlante and Shaffer, 2014), Journal of Gambling Business and Economics, accessed 2026-09-23
  5. Evidence brief: Proportion of revenue from problem gambling (2019), Gambling Research Exchange Ontario (GREO), accessed 2026-09-23
  6. The Gamma-Gamma Model of Monetary Value (Fader and Hardie, 2013), Bruce Hardie, London Business School, accessed 2026-09-23
  7. Ascarza's "Retention Futility: Targeting High-Risk Customers Might be Ineffective" wins 2018 award (Journal of Marketing Research, February 2018), American Marketing Association, accessed 2026-09-23
  8. A Deep Probabilistic Model for Customer Lifetime Value Prediction (Wang, Liu and Miao, 2019), arXiv, accessed 2026-09-23
  9. A Contextual-Bandit Approach to Personalized News Article Recommendation (Li, Chu, Langford and Schapire, 2010), arXiv, accessed 2026-09-23

Check your understanding

3 questions · answer them all, then check.

  1. 1. An operator sends a retention bonus to every customer with a churn probability above 0.7. The likely waste is on:

  2. 2. Why does average revenue per user times average lifetime fail as an LTV estimate in gambling?

  3. 3. A recommendation engine learns to show losing customers the highest-variance slot because it maximises next-session revenue. This is:

Sign in to track your progress through the course.

Cookie Preferences

Choose which cookies you want to accept. Essential cookies are required for the website to function properly.

Required

Necessary for the website to function. Cannot be disabled.

Help us understand how visitors interact with our website.

Used to deliver relevant advertisements and track ad performance.

Remember your preferences and settings for a better experience.