Two families of model
A pricing model answers one question: given what is known before the event, what is the probability of each outcome? There are two broad families and a good desk runs both.
Rating models assign each competitor a strength number and convert the difference in strengths into a probability. Elo is the archetype. They are simple, robust, and work in any sport with head-to-head contests.
Outcome models describe the scoring process directly. Poisson models of goals in football are the archetype. They produce a full distribution over scorelines, which means one model prices the match result, the total goals, the handicap, both teams to score and the correct score consistently. That consistency is the reason they dominate in scoring sports: a book pricing each market with its own separate model will be arbitraged across markets by anyone who notices the inconsistency.
Elo, and why it is still used
Elo updates each competitor's rating after every result: winner gains, loser loses, and the size of the move depends on how surprising the result was. The expected score for A against B is 1 / (1 + 10^((R_B - R_A) / 400)), where a win scores 1, a draw 0.5 and a loss 0. The constants are conventional and arbitrary; what matters is the K factor, which sets how fast ratings react.
K is a bias-variance trade-off. High K tracks form quickly but is noisy; low K is stable but slow to notice a team has changed. The right K differs by sport and by phase of season, and it should be tuned by minimising log loss, a proper scoring rule, on historical data rather than chosen by feel. Extensions that matter in practice:
- Home advantage as a fixed rating offset, fitted per competition and re-fitted regularly, because it changes over time: in English league football it has declined since the mid-1980s, and it fell in many leagues when matches were played without crowds during the pandemic, though the effect varied by country. Some public systems use one figure for everything: World Football Elo Ratings adds 100 points to the home side.
- Margin of victory scaling the update so a 5-0 moves ratings more than a 1-0, with the extra weight shrinking as the margin grows so blowouts do not dominate: World Football Elo Ratings raises K by a half for a two-goal win, by three-quarters for three and by a further eighth for each goal beyond that.
- Regression between seasons, pulling every rating part-way to the mean over the summer, because squads change and last season's rating is only partly informative.
- Glicko and TrueSkill add a rating uncertainty, which in Glicko grows while a competitor is inactive and shrinks with each result and in TrueSkill scales how far each result moves the rating, so a team with few recent results moves faster. Worth it in sports with sparse schedules; unnecessary in a 38-game league.
A rating model gives a win probability, or in three-outcome sports a win probability that then has to be split into win, draw and loss. That split is the weak point, and it is why football desks prefer outcome models.
Poisson and its corrections
The foundational football model, set out by Maher in 1982 and extended by Dixon and Coles, treats each team's goals as a Poisson random variable with a mean equal to its attacking strength times the opponent's defensive weakness times the league's average, with a home-advantage multiplier. Fit the strengths by maximum likelihood over a window of results and you have a scoreline distribution for every fixture.
Two corrections are standard and should be treated as mandatory:
Dixon and Coles (1997) found, in English league and cup matches from 1992 to 1995, that the independence assumption of plain Poisson underestimates 0-0 and 1-1 and overestimates 1-0 and 0-1. Their dependence parameter adjusts the low-scoring cells. Without it the draw price is wrong in a way that costs money on every match.
Time weighting, proposed in the same paper, downweights older results with an exponential decay so the fit reflects current strength. The decay rate is another parameter to tune on log loss: Dixon and Coles chose it by maximising the log-likelihood of match outcomes, and their optimum of 0.0065 per half-week implies a half-life of roughly a year. The right value for other leagues and eras has to be found the same way.
Beyond those, teams add what their data supports: a bivariate Poisson to capture the correlation between the two teams' goals, which also improves the prediction of draws, a negative binomial when the data is overdispersed, and covariates for rest days, travel, cup distraction and lineup changes. Each addition must earn its place on out-of-sample log loss. A model with fifteen parameters that fits last season perfectly and prices next season worse than a five-parameter model is a common and expensive mistake.
Expected goals and the input problem
Results are a noisy signal of strength: a team can dominate and lose. Expected goals (xG) models score each shot by its probability of becoming a goal given location, angle, assist type and situation, and sum over a match. Feeding xG rather than goals into the strength model reduces noise, and betting companies were among the metric's early adopters. Many models use a blend of xG and actual goals, because xG measures only the quality of chances and it is actual goals that decide matches.
This introduces a dependency: the quality of the price now rests on the quality of a third-party event feed and its shot model. A change in the provider's xG definition will move every strength rating in the book, and each provider's model is fitted to a particular body of historical shots: Opta's xG model is trained on nearly one million shots from 40 competitions between 2018-19 and 2021-22. Desks should version their inputs, monitor for drift in the input distribution, and be able to say exactly which feed version produced a given price.
The same pattern applies across sports. Tennis desks use point-level data and serve/return models rather than match results. Basketball desks use possession-based efficiency ratings. American football desks use play-level expected points. In every case the principle is the same: model the process that generates results rather than the results themselves, because the process has far more observations.
Lineups, news and the last mile
A model fitted on historical data knows nothing about tonight's team sheet. The gap between the model's opening price and a good price at kick-off is mostly information: injuries, rotations, weather, motivation in a dead rubber. There are two ways to close it.
The first is to build the information into the model: player-level ratings that sum to a team strength, so that a missing striker moves the price by a computed amount. This is the more rigorous route, and it is expensive, because player-level models need player-level data across every competition offered.
The second is trader judgement, applied as an adjustment to the model output. This is faster to build and harder to govern. The discipline is to log every manual adjustment with a reason, and to review them: a trader whose adjustments improve log loss is adding information; one whose adjustments do not should be adjusting less. Only the log can show which traders add real value and which adjustments are noise, and that finding is itself worth the logging.
Model governance
A pricing model is a production system and needs the controls of one:
- Versioning. Every price should be traceable to a model version, a parameter set and an input snapshot.
- Out-of-sample testing. The model is fitted on one window and scored on the next, never on the data it was fitted to.
- Champion and challenger. A new model runs in shadow against the live one until it has beaten it on log loss over a meaningful sample, which for a weekly league can take a season.
- Kill switches. When a model's live performance degrades past a threshold, prices revert to the market blend automatically.
- Explanation. A trader should be able to see why the model priced a match as it did: the strengths, the home advantage, the adjustments. Black-box pricing is unmanageable when it goes wrong, and it will.
The next lesson turns to the other source of probability: the market itself, and how to know whether your model or the market is closer to the truth.