How this works
How accurate are NFL prediction models?
The short answer
What the numbers mean
Why two in three is the ceiling for now
NFL games are close and short. Sixteen or seventeen games a season, one possession deciding a third of them, and injuries that change a team from week to week. The betting market absorbs a week of public and private information, price movement, sharp money, injury news, and still only reaches about 66–67%. That number is the honest benchmark for everything else: a model that cannot beat it is normal, and one that claims to beat it by a lot over many seasons deserves a very hard look.
One season proves almost nothing
Over a 272-game regular season, a picker whose true accuracy is 66.5% will land anywhere between about 61% and 72% on luck alone (one standard deviation is 2.9 points). The market itself ranged from 62.4% to 71.4% across seasons with no change in how good it was. So a model that went 70% last year has told you very little; a model that went 65% for ten years has told you a lot.
Accuracy is the wrong headline anyway
Two models can both go 65% while one says "55%" on every game and the other says "80%" on every game. The second is lying about what it knows. The better measure is calibration: when a model says 70%, do those teams win about 70% of the time? And the Brier score, which penalizes confident misses more than cautious ones (lower is better; a coin flip scores 0.250, the market's final price about 0.211, a good model near 0.215–0.220). A model that publishes probabilities and lets you check them is worth more than one that publishes a record.
Against the market margin is a different question
Picking winners and beating the point market margin are not the same skill. The market margin is set so that either side is about a coin flip, and a bettor has to win 52.4% of market margin picks just to cover the bookmaker's margin. Public models that beat the final market margin consistently are, in our experience of the public record, essentially nonexistent, and ours does not either. Any site showing a big against-the-market margin record without showing every pick before kickoff is asking you to take it on faith.
How to judge any model's claim
Five questions, in order:
1. Were the picks published before kickoff, and can I see all of them? A record built after the fact, or shown selectively, is not a record.
2. Over how many games? Anything under a few hundred is mostly noise. Sixteen seasons of backtest is about 4,000 games; that is where the noise gets small enough to mean something.
3. Was the backtest honest? "Walk-forward" means the model only ever used games that had already been played. A model tested on the same seasons it was tuned on will look better than it is.
4. What is the benchmark? Compare against the market's final price and against "pick the home team," on the same games. "65%" means nothing until you know the market got 66.5% on those games.
5. Are the probabilities calibrated? Ask for the table: stated probability by bucket versus how often those picks won. If the site cannot produce it, it has not checked.
Our own answers: picks stored before kickoff and never edited (every one on the weekly pages and Record); 4,142 backtest games walk-forward with a leak audit (Backtest); the market and the home-team baseline shown beside us; calibration published live and below. We are a normal model, 65.0% against the market's 66.5%, and we say so.
Season by season, 2010–2025
| Season | Games | Market pick at kickoff | Our model | Home team |
|---|---|---|---|---|
| 2010 | 240 | 66.7% | 61.7% | 54.6% |
| 2011 | 256 | 66.8% | 68.4% | 56.6% |
| 2012 | 255 | 64.3% | 63.1% | 57.3% |
| 2013 | 255 | 71.4% | 63.5% | 60.0% |
| 2014 | 255 | 66.7% | 69.4% | 56.9% |
| 2015 | 256 | 62.5% | 64.1% | 53.9% |
| 2016 | 253 | 64.4% | 64.8% | 57.7% |
| 2017 | 253 | 70.8% | 66.0% | 56.5% |
| 2018 | 254 | 66.1% | 68.9% | 60.2% |
| 2019 | 255 | 64.3% | 64.3% | 51.8% |
| 2020 | 255 | 67.5% | 67.1% | 49.8% |
| 2021 | 271 | 62.4% | 61.3% | 51.7% |
| 2022 | 269 | 66.2% | 62.5% | 56.1% |
| 2023 | 272 | 68.0% | 63.6% | 55.5% |
| 2024 | 272 | 71.3% | 70.2% | 53.3% |
| 2025 | 271 | 65.3% | 60.9% | 53.9% |
| All | 4,142 | 66.5% | 65.0% | 55.3% |
How often pick actually win
| Market said | Games | Pick won | Our pick won |
|---|---|---|---|
| 50–55% | 556 | 52.9% | 51.1% |
| 55–60% | 737 | 57.4% | 53.5% |
| 60–65% | 775 | 59.6% | 58.3% |
| 65–70% | 615 | 68.5% | 66.3% |
| 70–75% | 596 | 74.7% | 74.3% |
| 75–80% | 447 | 78.7% | 78.5% |
| 80%+ | 415 | 86.3% | 86.3% |
Where these numbers come from
Game results, market's final price and play-by-play from the open-source nflverse project; our backtest replays 2010–2025 week by week with the model retrained only on games already played, checked by an automated leak audit. Our live picks are stored before kickoff and graded after the final, with the market pick at kickoff graded beside them on the same games. Every term here, win probability, calibration, Brier score, market's final price, walk-forward, is defined in the glossary. Method, in the level of detail we publish, is on How this works. This page is updated automatically whenever the backtest or the live record changes.