Skip to content
iGaming Times

Independent industry intelligence in your inbox. We will email you a link to confirm your subscription, and every newsletter carries a one-click unsubscribe link.

Lesson 2 of 6 · 18 min

Rating Models and Scoring-Process Models

Elo and its extensions, Poisson with the Dixon-Coles and time-weighting corrections, expected goals as an input, closing the last mile with lineups, and the governance a production model needs.

Fact-checked 23 September 2026 by iGaming Times editorial team · 12 sources

In this lesson

  • Explain the difference between a rating model and a scoring-process model and why the latter dominates in scoring sports
  • Tune an Elo K factor, home advantage and between-season regression against log loss
  • Fit a time-weighted Poisson model with the Dixon-Coles correction and justify each further parameter out of sample
  • Describe how expected goals and other process-level inputs reduce noise and what dependency they introduce
  • Set out versioning, shadow running and kill-switch controls for a live pricing model

Two families of model

A pricing model answers one question: given what is known before the event, what is the probability of each outcome? There are two broad families and a good desk runs both.

Rating models assign each competitor a strength number and convert the difference in strengths into a probability. Elo is the archetype. They are simple, robust, and work in any sport with head-to-head contests.

Outcome models describe the scoring process directly. Poisson models of goals in football are the archetype. They produce a full distribution over scorelines, which means one model prices the match result, the total goals, the handicap, both teams to score and the correct score consistently. That consistency is the reason they dominate in scoring sports: a book pricing each market with its own separate model will be arbitraged across markets by anyone who notices the inconsistency.

Elo, and why it is still used

Elo updates each competitor's rating after every result: winner gains, loser loses, and the size of the move depends on how surprising the result was. The expected score for A against B is 1 / (1 + 10^((R_B - R_A) / 400)), where a win scores 1, a draw 0.5 and a loss 0. The constants are conventional and arbitrary; what matters is the K factor, which sets how fast ratings react.

K is a bias-variance trade-off. High K tracks form quickly but is noisy; low K is stable but slow to notice a team has changed. The right K differs by sport and by phase of season, and it should be tuned by minimising log loss, a proper scoring rule, on historical data rather than chosen by feel. Extensions that matter in practice:

A rating model gives a win probability, or in three-outcome sports a win probability that then has to be split into win, draw and loss. That split is the weak point, and it is why football desks prefer outcome models.

Poisson and its corrections

The foundational football model, set out by Maher in 1982 and extended by Dixon and Coles, treats each team's goals as a Poisson random variable with a mean equal to its attacking strength times the opponent's defensive weakness times the league's average, with a home-advantage multiplier. Fit the strengths by maximum likelihood over a window of results and you have a scoreline distribution for every fixture.

Two corrections are standard and should be treated as mandatory:

Dixon and Coles (1997) found, in English league and cup matches from 1992 to 1995, that the independence assumption of plain Poisson underestimates 0-0 and 1-1 and overestimates 1-0 and 0-1. Their dependence parameter adjusts the low-scoring cells. Without it the draw price is wrong in a way that costs money on every match.

Time weighting, proposed in the same paper, downweights older results with an exponential decay so the fit reflects current strength. The decay rate is another parameter to tune on log loss: Dixon and Coles chose it by maximising the log-likelihood of match outcomes, and their optimum of 0.0065 per half-week implies a half-life of roughly a year. The right value for other leagues and eras has to be found the same way.

Beyond those, teams add what their data supports: a bivariate Poisson to capture the correlation between the two teams' goals, which also improves the prediction of draws, a negative binomial when the data is overdispersed, and covariates for rest days, travel, cup distraction and lineup changes. Each addition must earn its place on out-of-sample log loss. A model with fifteen parameters that fits last season perfectly and prices next season worse than a five-parameter model is a common and expensive mistake.

Expected goals and the input problem

Results are a noisy signal of strength: a team can dominate and lose. Expected goals (xG) models score each shot by its probability of becoming a goal given location, angle, assist type and situation, and sum over a match. Feeding xG rather than goals into the strength model reduces noise, and betting companies were among the metric's early adopters. Many models use a blend of xG and actual goals, because xG measures only the quality of chances and it is actual goals that decide matches.

This introduces a dependency: the quality of the price now rests on the quality of a third-party event feed and its shot model. A change in the provider's xG definition will move every strength rating in the book, and each provider's model is fitted to a particular body of historical shots: Opta's xG model is trained on nearly one million shots from 40 competitions between 2018-19 and 2021-22. Desks should version their inputs, monitor for drift in the input distribution, and be able to say exactly which feed version produced a given price.

The same pattern applies across sports. Tennis desks use point-level data and serve/return models rather than match results. Basketball desks use possession-based efficiency ratings. American football desks use play-level expected points. In every case the principle is the same: model the process that generates results rather than the results themselves, because the process has far more observations.

Lineups, news and the last mile

A model fitted on historical data knows nothing about tonight's team sheet. The gap between the model's opening price and a good price at kick-off is mostly information: injuries, rotations, weather, motivation in a dead rubber. There are two ways to close it.

The first is to build the information into the model: player-level ratings that sum to a team strength, so that a missing striker moves the price by a computed amount. This is the more rigorous route, and it is expensive, because player-level models need player-level data across every competition offered.

The second is trader judgement, applied as an adjustment to the model output. This is faster to build and harder to govern. The discipline is to log every manual adjustment with a reason, and to review them: a trader whose adjustments improve log loss is adding information; one whose adjustments do not should be adjusting less. Only the log can show which traders add real value and which adjustments are noise, and that finding is itself worth the logging.

Model governance

A pricing model is a production system and needs the controls of one:

  • Versioning. Every price should be traceable to a model version, a parameter set and an input snapshot.
  • Out-of-sample testing. The model is fitted on one window and scored on the next, never on the data it was fitted to.
  • Champion and challenger. A new model runs in shadow against the live one until it has beaten it on log loss over a meaningful sample, which for a weekly league can take a season.
  • Kill switches. When a model's live performance degrades past a threshold, prices revert to the market blend automatically.
  • Explanation. A trader should be able to see why the model priced a match as it did: the strengths, the home advantage, the adjustments. Black-box pricing is unmanageable when it goes wrong, and it will.

The next lesson turns to the other source of probability: the market itself, and how to know whether your model or the market is closer to the truth.

Key terms

Elo rating
A strength number per competitor updated after each result by an amount proportional to how surprising the result was; the K factor sets the speed.
Dixon-Coles correction
An adjustment to the independent Poisson model for the observed excess of 0-0 and 1-1 scorelines and deficit of 1-0 and 0-1, published by Dixon and Coles in 1997 in the same paper that also proposed exponential time weighting.
Expected goals (xG)
The sum over a match of each shot’s modelled probability of scoring; a lower-noise measure of team performance than goals.
Champion-challenger
Running a new model in shadow against the live one until it has beaten it on a proper scoring rule over a meaningful sample.
Log loss
The negative mean log probability assigned to outcomes that occurred; a proper scoring rule and the standard objective for fitting and comparing probability models.

Key takeaways

  • One scoring-process model prices every market in an event consistently; separate per-market models invite cross-market arbitrage.
  • Plain Poisson misprices the draw; Dixon-Coles and exponential time weighting are mandatory, everything beyond must earn its place on out-of-sample log loss.
  • Model the process that generates results (shots, points, possessions) rather than results, because the process has far more observations.
  • Manual adjustments are logged and reviewed; the log is how you tell which traders add real information and which adjustments are noise.
  • Every price is traceable to a model version, a parameter set and an input snapshot.

Sources

The legislation, regulator material and research this lesson was checked against.

  1. Modelling Association Football Scores and Inefficiencies in the Football Betting Market (Dixon and Coles, Applied Statistics 46(2), 1997), Royal Statistical Society, accessed 2026-09-23
  2. Analysis of sports data by using bivariate Poisson models (Karlis and Ntzoufras, The Statistician 52(3), 2003), Royal Statistical Society, accessed 2026-09-23
  3. Seasonal Home Advantage in English Professional Football, 1974 to 2018 (Peeters and van Ours, De Economist, 2021), Springer Nature, accessed 2026-09-23
  4. The impact of crowd effects on home advantage of football matches during the COVID-19 pandemic: a systematic review, PLOS One, accessed 2026-09-23
  5. Strictly Proper Scoring Rules, Prediction, and Estimation (Gneiting and Raftery, JASA, 2007), University of Washington (author copy), accessed 2026-09-23
  6. A Starting Point for Analyzing Basketball Statistics (Kubatko, Oliver, Pelton and Rosenbaum, JQAS, 2007), De Gruyter, accessed 2026-09-23
  7. nflWAR: A Reproducible Method for Offensive Player Evaluation in Football (Yurko, Ventura and Horowitz), arXiv (published in JQAS, 2019), accessed 2026-09-23
  8. Combining player statistics to predict outcomes of tennis matches (Barnett and Clarke, IMA Journal of Management Mathematics, 2005), Victoria University repository, accessed 2026-09-23
  9. The Glicko system (Mark E. Glickman), glicko.net, accessed 2026-09-23
  10. TrueSkill Ranking System, Microsoft Research, accessed 2026-09-23
  11. About World Football Elo Ratings, eloratings.net, accessed 2026-09-23
  12. What Is Expected Goals (xG)?, Opta Analyst (Stats Perform), accessed 2026-09-23

Check your understanding

3 questions · answer them all, then check.

  1. 1. Why do football desks prefer a scoring-process model to a rating model?

  2. 2. A fifteen-parameter model fits last season better than a five-parameter one and prices next season worse. What has happened?

  3. 3. What is the main risk of feeding expected goals into the strength model?

Sign in to track your progress through the course.

Cookie Preferences

Choose which cookies you want to accept. Essential cookies are required for the website to function properly.

Required

Necessary for the website to function. Cannot be disabled.

Help us understand how visitors interact with our website.

Used to deliver relevant advertisements and track ad performance.

Remember your preferences and settings for a better experience.

Rating Models and Scoring-Process Models