The market as a model
The betting market aggregates every participant's information and every model in the world, weighted by how much money each is prepared to risk. The price at which a deep market closes is, empirically, among the most accurate forecasts of a sporting outcome available anywhere: studies repeatedly find that betting odds outperform statistical rating models and expert tipsters. Any desk that does not use it is throwing away its most valuable input; any desk that uses only it has no business being a desk.
The practical version of this: the closing line of a sharp book, de-margined, is the benchmark. "Sharp" means a book that welcomes informed money, runs low margins and high limits, and moves its price on what it takes. Its closing price incorporates everything the market learned up to the off. Recreational books, which restrict winning customers and shade toward their own liabilities, are less informative, though not useless.
A desk should maintain a consensus probability from the sharp books it can observe, weighted by how informative each has proved historically, and refresh it continuously. Even a crude consensus is powerful: in one study of 479,440 football matches from 2005 to 2015, the probability implied by the average closing price across up to 32 bookmakers tracked realised outcome frequencies closely. That consensus is the second input to the blend described in lesson 1.
Closing line value
If the closing line is the best estimate of truth, then the quality of any earlier price can be measured against it. A bet placed at 2.20 on a selection that closes at 2.00 was placed at a better price than the market's final view; the bettor got closing line value (CLV). Over a large sample, CLV is a better guide to future profit than past profit is, because profit is dominated by variance for hundreds of bets and CLV much less so. An analysis of 87,960 pairs of pre-closing and closing football prices from one sharp bookmaker, covering four seasons from 2012/13, found that the ratio of the price taken to the closing price predicted actual returns almost one for one, less the bookmaker's margin.
For the book this is the mirror image and it is the most useful customer metric a trading desk has. A customer who consistently beats the closing line is informed, whether or not they are currently winning. A customer who is winning but consistently gets worse than closing prices is running lucky and will give it back. Customer management built on CLV rather than P&L is both fairer and more accurate.
The same measure applies to the desk's own opening prices. A desk whose opening prices are systematically beaten by its own closing prices in a predictable direction has a bias in its opening model, and the direction tells it which way.
Calibration: the test that matters
A probability model is calibrated if, across all the events it priced at 30%, about 30% happened. That is a different property from accuracy: a model can be well calibrated and uninformative (always saying 50%), or sharp and badly calibrated (confident and wrong). A pricing model needs both, and calibration is the one that gets neglected.
The calibration check is simple. Bin predictions by probability (0 to 5%, 5 to 10%, and so on), count actual outcomes in each bin, and plot observed frequency against predicted. A calibrated model sits on the diagonal. The common failures are:
- Overconfidence: predictions above 70% happen less than predicted, predictions below 30% happen more. The model needs shrinking toward the mean.
- Underconfidence: the reverse, common in models that regress too hard.
- Longshot miscalibration: the tail bins are off because there are few observations and the model extrapolates poorly there. Longshots are also where the effective margin is widest, the favourite-longshot bias: in US horse racing from 1992 to 2001, bets on horses at 100/1 or longer lost about 61% of stakes, against 5.5% for backing the favourite in every race. So this error is also where it is most hidden.
Fixes are mechanical. Platt scaling fits a logistic regression from model output to outcome; isotonic regression fits a monotone step function; both are applied after the model, fitted on data the model was not trained on, and re-fitted periodically. Isotonic regression can correct any monotonic distortion but overfits when data is scarce, where Platt scaling does better, which matters in thin competitions and in the longshot tail. A desk that runs calibration correction should keep the raw and corrected outputs so it can tell whether the underlying model is drifting.
Scoring rules
Calibration plots are diagnostic; a desk also needs a single number to compare models and to track over time. Two proper scoring rules are standard:
Log loss (the negative mean log probability assigned to the outcome that happened) punishes confident mistakes heavily, without limit as the probability given to the actual outcome approaches zero. Minimising it is maximum likelihood estimation, which makes it the natural objective for fitting model parameters.
Brier score (the mean squared difference between the probability vector and the outcome indicator) is bounded and more interpretable, and, as Allan Murphy showed in 1973, decomposes into reliability (calibration), resolution and uncertainty terms that tell you where a model is gaining or losing. A single Brier number is not a calibration test on its own: a worse calibrated model with more discriminating power can score better, which is why the calibration plot is still needed.
Both should be tracked per sport, per competition and per market type, in a rolling window, against the market consensus as a baseline. The question every week is not "is the model good" but "is the model better than the market, and where". A desk that cannot answer that per competition does not know where its edge is, and will price everything the same way, which means it will lose margin where it has no edge and leave money on the table where it does.
Weighting money: reading the flow
The market is informative because money is, but not all money equally. A desk that treats every bet as equal information will be moved around by recreational volume and will miss the sharp bet that mattered. The refinement is to weight each bet by how informative that customer has proven to be, which brings CLV back in: the book already knows who beats the close.
The standard approach is a per-customer weight, updated over time, that governs how much the model probability moves when they bet. Sharp money moves the probability; recreational money moves only the shape (the margin distribution) and the liability. In effect the book runs an internal prediction market in which its best customers have the loudest voice, and it charges them for the privilege through lower margins and, where necessary, limits.
The ethical and commercial edge of this is worth stating. Restricting winning customers is common practice, and in one study bookmakers began to severely limit the researchers' accounts within months of their profitable betting. It is also usually permitted: Massachusetts rules, for example, expressly allow an operator to limit a patron's wager for reasons it considers necessary or appropriate. Regulators have started to demand transparency, though. Since 1 June 2026 Massachusetts has required licensed sportsbooks to give a limited patron timely notice, a specific explanation and the markets affected, and the Commission expects that notice within 48 hours. A book that restricts too aggressively also loses the information it needs and gains a reputation that costs it acquisition. Sharp books take the other route: accept the informed money, use it to fix the price, and make margin from the corrected price against everyone else. Which route a book takes is a strategic choice, not a trading one, but the trading desk lives with the consequences either way.
Detecting when you are wrong
The market will tell a desk when its price is wrong, if the desk is listening. The signals, in rough order of urgency:
- A price materially away from the sharp consensus at a time when the consensus is stable. The desk has either information or an error, and it should know which.
- A run of sharp money in one direction on one market. Something is known that the model does not know.
- Arbitrage volume: customers betting the desk's price against a rival's. This is pure signal that the desk is off.
- A drift in calibration over a rolling window in one competition. The model has stopped describing that competition.
Each of these should be an alert with a defined response, and the responses should be automated where the desk trusts the rule and escalated where it does not. The lesson on risk management picks this up; the next lesson is about the situation where all of this happens many times a minute: in-play.