Skip to content
iGaming Times

Independent industry intelligence in your inbox. We will email you a link to confirm your subscription, and every newsletter carries a one-click unsubscribe link.

Lesson 3 of 6 · 18 min

The Market as a Model: Closing Lines, Calibration and Scoring

Why the sharp closing line is the benchmark, closing line value as the best customer and desk metric, calibration plots and their fixes, log loss and Brier score, and weighting money by how informed it is.

Fact-checked 23 September 2026 by iGaming Times editorial team · 12 sources

In this lesson

  • Build a de-margined consensus from sharp books and explain why it is among the most accurate available predictors
  • Use closing line value to assess customers and the desk’s own opening prices
  • Produce and read a calibration plot and apply Platt or isotonic correction
  • Track log loss and Brier score per competition against the market baseline and interpret the gap
  • Weight incoming bets by customer informativeness so information moves probability and volume moves shape

The market as a model

The betting market aggregates every participant's information and every model in the world, weighted by how much money each is prepared to risk. The price at which a deep market closes is, empirically, among the most accurate forecasts of a sporting outcome available anywhere: studies repeatedly find that betting odds outperform statistical rating models and expert tipsters. Any desk that does not use it is throwing away its most valuable input; any desk that uses only it has no business being a desk.

The practical version of this: the closing line of a sharp book, de-margined, is the benchmark. "Sharp" means a book that welcomes informed money, runs low margins and high limits, and moves its price on what it takes. Its closing price incorporates everything the market learned up to the off. Recreational books, which restrict winning customers and shade toward their own liabilities, are less informative, though not useless.

A desk should maintain a consensus probability from the sharp books it can observe, weighted by how informative each has proved historically, and refresh it continuously. Even a crude consensus is powerful: in one study of 479,440 football matches from 2005 to 2015, the probability implied by the average closing price across up to 32 bookmakers tracked realised outcome frequencies closely. That consensus is the second input to the blend described in lesson 1.

Closing line value

If the closing line is the best estimate of truth, then the quality of any earlier price can be measured against it. A bet placed at 2.20 on a selection that closes at 2.00 was placed at a better price than the market's final view; the bettor got closing line value (CLV). Over a large sample, CLV is a better guide to future profit than past profit is, because profit is dominated by variance for hundreds of bets and CLV much less so. An analysis of 87,960 pairs of pre-closing and closing football prices from one sharp bookmaker, covering four seasons from 2012/13, found that the ratio of the price taken to the closing price predicted actual returns almost one for one, less the bookmaker's margin.

For the book this is the mirror image and it is the most useful customer metric a trading desk has. A customer who consistently beats the closing line is informed, whether or not they are currently winning. A customer who is winning but consistently gets worse than closing prices is running lucky and will give it back. Customer management built on CLV rather than P&L is both fairer and more accurate.

The same measure applies to the desk's own opening prices. A desk whose opening prices are systematically beaten by its own closing prices in a predictable direction has a bias in its opening model, and the direction tells it which way.

Calibration: the test that matters

A probability model is calibrated if, across all the events it priced at 30%, about 30% happened. That is a different property from accuracy: a model can be well calibrated and uninformative (always saying 50%), or sharp and badly calibrated (confident and wrong). A pricing model needs both, and calibration is the one that gets neglected.

The calibration check is simple. Bin predictions by probability (0 to 5%, 5 to 10%, and so on), count actual outcomes in each bin, and plot observed frequency against predicted. A calibrated model sits on the diagonal. The common failures are:

  • Overconfidence: predictions above 70% happen less than predicted, predictions below 30% happen more. The model needs shrinking toward the mean.
  • Underconfidence: the reverse, common in models that regress too hard.
  • Longshot miscalibration: the tail bins are off because there are few observations and the model extrapolates poorly there. Longshots are also where the effective margin is widest, the favourite-longshot bias: in US horse racing from 1992 to 2001, bets on horses at 100/1 or longer lost about 61% of stakes, against 5.5% for backing the favourite in every race. So this error is also where it is most hidden.

Fixes are mechanical. Platt scaling fits a logistic regression from model output to outcome; isotonic regression fits a monotone step function; both are applied after the model, fitted on data the model was not trained on, and re-fitted periodically. Isotonic regression can correct any monotonic distortion but overfits when data is scarce, where Platt scaling does better, which matters in thin competitions and in the longshot tail. A desk that runs calibration correction should keep the raw and corrected outputs so it can tell whether the underlying model is drifting.

Scoring rules

Calibration plots are diagnostic; a desk also needs a single number to compare models and to track over time. Two proper scoring rules are standard:

Log loss (the negative mean log probability assigned to the outcome that happened) punishes confident mistakes heavily, without limit as the probability given to the actual outcome approaches zero. Minimising it is maximum likelihood estimation, which makes it the natural objective for fitting model parameters.

Brier score (the mean squared difference between the probability vector and the outcome indicator) is bounded and more interpretable, and, as Allan Murphy showed in 1973, decomposes into reliability (calibration), resolution and uncertainty terms that tell you where a model is gaining or losing. A single Brier number is not a calibration test on its own: a worse calibrated model with more discriminating power can score better, which is why the calibration plot is still needed.

Both should be tracked per sport, per competition and per market type, in a rolling window, against the market consensus as a baseline. The question every week is not "is the model good" but "is the model better than the market, and where". A desk that cannot answer that per competition does not know where its edge is, and will price everything the same way, which means it will lose margin where it has no edge and leave money on the table where it does.

Weighting money: reading the flow

The market is informative because money is, but not all money equally. A desk that treats every bet as equal information will be moved around by recreational volume and will miss the sharp bet that mattered. The refinement is to weight each bet by how informative that customer has proven to be, which brings CLV back in: the book already knows who beats the close.

The standard approach is a per-customer weight, updated over time, that governs how much the model probability moves when they bet. Sharp money moves the probability; recreational money moves only the shape (the margin distribution) and the liability. In effect the book runs an internal prediction market in which its best customers have the loudest voice, and it charges them for the privilege through lower margins and, where necessary, limits.

The ethical and commercial edge of this is worth stating. Restricting winning customers is common practice, and in one study bookmakers began to severely limit the researchers' accounts within months of their profitable betting. It is also usually permitted: Massachusetts rules, for example, expressly allow an operator to limit a patron's wager for reasons it considers necessary or appropriate. Regulators have started to demand transparency, though. Since 1 June 2026 Massachusetts has required licensed sportsbooks to give a limited patron timely notice, a specific explanation and the markets affected, and the Commission expects that notice within 48 hours. A book that restricts too aggressively also loses the information it needs and gains a reputation that costs it acquisition. Sharp books take the other route: accept the informed money, use it to fix the price, and make margin from the corrected price against everyone else. Which route a book takes is a strategic choice, not a trading one, but the trading desk lives with the consequences either way.

Detecting when you are wrong

The market will tell a desk when its price is wrong, if the desk is listening. The signals, in rough order of urgency:

  1. A price materially away from the sharp consensus at a time when the consensus is stable. The desk has either information or an error, and it should know which.
  2. A run of sharp money in one direction on one market. Something is known that the model does not know.
  3. Arbitrage volume: customers betting the desk's price against a rival's. This is pure signal that the desk is off.
  4. A drift in calibration over a rolling window in one competition. The model has stopped describing that competition.

Each of these should be an alert with a defined response, and the responses should be automated where the desk trusts the rule and escalated where it does not. The lesson on risk management picks this up; the next lesson is about the situation where all of this happens many times a minute: in-play.

Key terms

Closing line
The price at which a market closes at the start of the event; at a sharp book, among the most accurate available forecasts of the outcome.
Closing line value (CLV)
The amount by which the price a bet was placed at beats the closing price, usually measured as the ratio of the two; consistently positive CLV over a large sample indicates an informed bettor.
Calibration
The property that events priced at a given probability occur at that frequency; checked by binning predictions and comparing observed rates.
Brier score
The mean squared difference between the probability vector and the outcome; Murphy (1973) showed it decomposes into reliability (calibration), resolution and uncertainty.
Platt scaling
A post-hoc calibration fix that fits a logistic regression from model output to outcome, on data the model was not trained on; it copes better with small samples than isotonic regression.

Key takeaways

  • The de-margined closing line of a sharp book is the benchmark; a desk that ignores it discards its best input, a desk that only uses it has no edge.
  • CLV predicts a customer’s future profit better than their past profit does, and it is the defensible basis for limits.
  • Calibration and sharpness are different properties; a pricing model needs both and calibration is the one that gets neglected, especially in the longshot tail.
  • The weekly question is not whether the model is good but where it beats the market, per competition, because that is where edge and limits should go.
  • Sharp books accept informed money and use it to fix the price; recreational books restrict it and lose the information. Both are choices with consequences.

Sources

The legislation, regulator material and research this lesson was checked against.

  1. The Betting Odds Rating System: Using soccer forecasts to forecast soccer (Wunderlich and Memmert, PLOS ONE, 2018), PLOS ONE via PubMed Central, accessed 2026-09-23
  2. Predicting Good Probabilities With Supervised Learning (Niculescu-Mizil and Caruana, ICML 2005), Cornell University, accessed 2026-09-23
  3. Strictly Proper Scoring Rules, Prediction, and Estimation (Gneiting and Raftery, Journal of the American Statistical Association, 2007), University of Washington, accessed 2026-09-23
  4. A New Vector Partition of the Probability Score (Murphy, Journal of Applied Meteorology, 1973), American Meteorological Society, accessed 2026-09-23
  5. 205 CMR 247.00: Uniform Standards of Sports Wagering, 247.08 Minimum and Maximum Wagers, Massachusetts Gaming Commission, accessed 2026-09-23
  6. 205 CMR 238.30: Acceptance of Sports Wagers, paragraph (11) notice of wagering limits, Massachusetts Gaming Commission, accessed 2026-09-23
  7. Explaining the Favorite-Longshot Bias: Is it Risk-Love or Misperceptions? (Snowberg and Wolfers, NBER Working Paper 15923, 2010), National Bureau of Economic Research, accessed 2026-09-23
  8. Beating the bookies with their own numbers (Kaunitz, Zhong and Kreiner, 2017 preprint), arXiv, accessed 2026-09-23
  9. Probability calibration (user guide), scikit-learn, accessed 2026-09-23
  10. Using Pinnacle.com's Closing Line to Predict Profits (Joseph Buchdahl, 2016), Football-Data, accessed 2026-09-23
  11. Massachusetts Wants a Good Reason for 'Limiting' Sports Bettors (26 February 2026), Covers, accessed 2026-09-23
  12. UPDATE: Starting Today, Massachusetts Sports Betting Operators Must Send Out Limitation Notices to Users (1 June 2026), SportsBettingDime, accessed 2026-09-23

Check your understanding

3 questions · answer them all, then check.

  1. 1. A customer is up 20% over 300 bets but consistently gets worse prices than the close. The right reading is:

  2. 2. A calibration plot shows predictions above 70% occurring less often than predicted. The model is:

  3. 3. Why should the desk score its model against the market consensus rather than against results alone?

Sign in to track your progress through the course.

Cookie Preferences

Choose which cookies you want to accept. Essential cookies are required for the website to function properly.

Required

Necessary for the website to function. Cannot be disabled.

Help us understand how visitors interact with our website.

Used to deliver relevant advertisements and track ad performance.

Remember your preferences and settings for a better experience.