The pricing problem, stated properly
A sportsbook price is a probability with a margin on top. Everything a quantitative trading team does is an attempt to get the probability right and to decide how much margin the market will bear. Those are two different problems, and a surprising amount of bad pricing comes from mixing them up: adjusting the probability when the margin should move, or the other way round.
Start with the conversion. A decimal price of 2.50 implies a probability of 1 / 2.50 = 0.40. Fractional 6/4 is the same thing (1 / (1 + 1.5)). American +150 is 100 / (150 + 100). None of this is difficult, but a pricing team works in one representation internally and it should be probability, not odds. Odds are for customers. Probabilities add, can be compared across markets, and expose the margin as a plain number.
Sum the implied probabilities of every outcome in a market and you get more than 1. On a football match priced 2.10, 3.40 and 3.60 the sum is 0.476 + 0.294 + 0.278 = 1.048. The 4.8% is the overround. It is close to, but not the same as, the book's theoretical margin: if money arrived in proportion to the prices, the book would pay back 1 / 1.048 = 0.954 of every unit staked, a margin of about 4.6% of turnover. It is the first number a trader looks at on a competitor's screen and the first number a model has to remove before the prices mean anything.
Removing the margin is a modelling choice
The naive way to strip the overround is to divide each implied probability by the sum. On the example above that gives 0.454, 0.281, 0.265. This is the multiplicative method and it is wrong in a specific, well-documented way: it assumes the margin is spread in proportion to probability, when in practice books load more margin onto longshots.
The favourite-longshot bias is one of the most documented findings in betting economics: first noted in horse racing in 1949, it has since been found in racetrack data around the world, and it holds in sports betting too. Across 84,230 European football matches from 2011/12 to 2021/22, bets in the lowest decile of implied probability lost 17% on average, against 2% in the highest decile. Bettors overvalue long odds, so books can shade them harder without losing volume. In a hypothetical market, offering 20.0 on a longshot whose fair price is 21.0 loses the book almost nothing in turnover and earns about 5% per unit staked, while shading a 1.25 favourite to 1.22 earns only about 2.4% and moves real money elsewhere. Any method that ignores this recovers probabilities that are too low for favourites and too high for longshots, because it leaves part of the longshots' extra margin in their fair probabilities: in a worked example by Clarke, Kovalchik and Ingram the multiplicative method gives a 1.15 favourite a fair probability of 0.696 where the power method gives 0.825.
Three alternatives are in common use:
- The power method raises each implied probability to an exponent k and solves for the k that makes the results sum to 1. It pulls margin off longshots harder than favourites.
- Shin's method models a market with a fraction of insider bettors and solves for that fraction; it was built for horse-race betting with bookmakers and first estimated on UK data, and produces the same shape of correction from an explicit story about why the shape exists. In a two-outcome market it gives the same answer as simply subtracting an equal share of the overround from each side.
- Odds-ratio and logit methods apply the correction in odds or log-odds space. The odds-ratio method, from Keith Cheung, solves for the single odds ratio between quoted and fair probabilities that makes the fair probabilities sum to 1, which leaves one parameter to fit per market.
Which one is right depends on the book you are reading. A team that regularly de-margins competitor prices should fit the method per competitor per sport, because each book has its own habit. The test is simple: take a season of that book's closing prices, de-margin with each method, and check which one's probabilities are best calibrated against results. When Clarke, Kovalchik and Ingram ran this kind of test in 2017 on tennis, greyhound and horse-racing prices, the power method beat the multiplicative method on every dataset and every measure, and generally matched or beat Shin's, though not on every measure. Calibration is covered in lesson 3; for now the point is that "the fair price" is not a fact you read off a screen but an estimate that depends on an assumption about how the margin was applied.
Margin is a decision about the market, not the event
Once the probability is settled, the margin is applied on top, and this is where a lot of trading judgement lives. The overround on a market is not a constant. It should reflect:
Confidence in the probability. A Premier League match with many models and a deep market can be priced tightly: across 51 bookmakers and an exchange, Premier League match-result odds in 2016/17 and 2017/18 implied an average overround of about 4%, and a book competing on price can choose to go lower. A third-division Norwegian fixture where the model is thin and the information is worse deserves a wider margin, perhaps 7% or more, because the book is being paid to carry the risk of being wrong.
Liquidity and information asymmetry. Markets that attract informed money need more margin, or lower limits, or both. In-play markets, where feed latency lets faster bettors act on information the book has not yet priced, justify more margin than pre-match for the same reason.
Competitive position. A recreational book can run wider than a sharp book because its customers are not price-shopping every bet. A book chasing acquisition through best-price advertising cannot.
The number of outcomes. Overround scales with the size of the market. In the same Premier League data the average overround was about 4% on the three-way result but about 12% on exact scorelines, and both are normal, because the margin per selection is what the trader controls and a scoreline market has many more selections.
The decision is usually encoded as a target margin per market type per competition tier, with the trader able to override. It should be reviewed with data: the right margin is the one that maximises expected profit given the elasticity of the customers, and elasticity can be measured by varying margin and watching turnover.
Where the margin goes: shading and shape
Applying a 5% overround to a three-way market still leaves the question of how to distribute it. A book that spreads it evenly across the three selections is applying a multiplicative margin and is leaving money on the table for the reasons above. The usual approach is a shape: less on the favourite, more on the draw and the outsider, with the shape fitted from the book's own turnover data so that the margin lands where the volume is least price-sensitive.
This is also where the book's liability position enters. If the book is already heavily exposed on the home side, it can shade the home price down and the away price up, moving the margin rather than the probability. A trader who "moves the line" in response to liability is doing exactly this, and a well-built pricing system makes the two levers separate: the model probability, which only new information should move, and the price shape, which liability and commercial considerations can move.
Blurring the two is a common and expensive failure in sportsbook pricing. When a book takes a large bet and moves the price, it should ask whether the bet was information (a sharp customer, a syndicate, a pattern that has been right before) or just volume. Information should move the probability. Volume should move the shape. A team that cannot tell the difference will either overreact to noise or underreact to signal, and both are expensive.
The reference model
Many quantitative desks keep two numbers for every selection: the model probability and the market probability, the latter being a de-margined consensus of the books they consider sharp. The price offered is a blend, weighted by how much the desk trusts its own model in that competition. In a top league the market weight is high; in a niche sport where the desk has an edge the model weight is high.
The blend weight is not a matter of taste. It can be estimated: run both probabilities against results over a long window, and find the mixing weight that minimises log loss. If the market always wins, the desk has learned that it does not have a model in that sport and should say so. If the model wins, the desk has an edge worth protecting with lower margin and higher limits, because that is where it can afford to take more.
What this lesson expects you to carry forward
Probability first, odds last. Strip margin with a method fitted to the book you are reading, not the naive one. Set margin as a decision about the market's confidence, liquidity and elasticity, not the event. Keep the probability lever and the shape lever separate and know which one a given bet should move. The rest of the course builds the probability side: where it comes from, how it is checked, and what it does in-play and across correlated markets.