About · Accuracy & Methodology
Honest benchmarks for a noisy market
We benchmark the performance of betting models — ours included — to separate real signal from marketing. Everything on this page is measured the hard way: on held-out seasons the model never trained on, against the closing line.
How accurate are we?
Honest answer: our projected margin matches the closing line — it does not beat it, and neither does anything else we track. What we add is calibration: probabilities that mean what they say, tie the devigged moneyline market, and beat the public consensus. Held-out evaluation, generated 2026-08-31.
Projected margin · MAE
9.82pts
1,960 held-out games
- Our calibrated stack
- 9.82
- Closing line (market)
- 9.82
- Best public model (DRatings)
- 10.10
Matches the closing line — the number nothing on the board beats. Matching it is the honest headline.
P(home cover) · calibration
0.6pp ECE
1,917 games with a posted spread
- Our log loss
- 0.6930
- Public consensus
- 0.7011
Probabilities mean what they say — and they beat the public consensus.
P(home win) · log loss
0.6080
1,954 games straight up
- Our calibrated stack
- 0.6080
- Devigged moneyline (market)
- 0.6083
- Public consensus
- 0.6358
Ties the devigged moneyline market; clearly ahead of consensus.
Margin scoreboard — 2019–2025
Mean absolute error of the projected final margin, on the games every model on the board predicted. Lower is better; nothing beats the closing line by a meaningful amount — including us.
| # | Model | Kind | Games | MAE | RMSE | ATS |
|---|---|---|---|---|---|---|
| 1 | BB final stack (ours, market-anchored) | Ours | 1,813 | 9.82 | 12.71 | 50.0% |
| 2 | Closing line (market) | Market | 1,813 | 9.84 | 12.73 | — |
| 3 | Computer Adjusted Line (market) | Market | 1,813 | 9.85 | 12.74 | 50.2% |
| 4 | Betting Benchmarks deep_honest (ours) | Ours | 1,813 | 9.87 | 12.75 | 52.3% |
| 5 | Midweek line (market) | Market | 1,813 | 9.88 | 12.77 | 48.0% |
| 6 | Opening line (market) | Market | 1,813 | 10.00 | 12.93 | 51.6% |
| 7 | DRatings | Public | 1,797 | 10.10 | 13.02 | 50.1% |
| 8 | Dokter Entropy | Public | 1,812 | 10.12 | 13.04 | 51.0% |
Market rows are betting lines — they are the market number, so they see information no forecaster does. Our final stack is market-anchored (closing line plus a disciplined step toward the model), so tracking the line this closely is by construction; the hair's-width MAE difference versus the closing line is sample wobble, not an edge. Rows without margin numbers publish probabilities only in this export. ATS is win rate against the closing spread — note how tightly the whole field clusters around 50%.
Do the probabilities mean what they say?
A reliability diagram answers the only question that matters for a probability: when we say 70%, does it happen 70% of the time? Each dot is one bin of games (2019–2025, all held-out); a calibrated forecast sits on the diagonal. Dot area is the number of games in the bin — the small edge bins are noisy by nature.
P(home win) — predicted vs realized
And P(home cover)? Every one of its 1,917 games lands in a single bin: predicted 48.6%, realized 49.1%. That's not a chart, it's a sentence — and it's the whole point: against a near-efficient closing spread, an honest cover probability is a coin flip, and ours says so. The +EV board's record lives earlier in the week, graded against opening lines — that edge decays as the market moves, which is exactly what these two numbers together predict.
The honest story
The projection is market-anchored. Our final stack starts from the closing line and takes a small, disciplined step toward the model's own number. That is why its margin error tracks the line so closely — by construction. On one matched subsample our MAE comes out a couple hundredths below the line's; that is sample wobble on a market-anchored forecast, not evidence of an edge, and we will not claim otherwise.
Calibration is the real product. Public models publish margins; almost none publish probabilities you can take at face value. Ours are calibrated on held-out data: cover probabilities with an expected calibration error under a percentage point, and win probabilities whose log loss ties the devigged moneyline market and clearly beats the public consensus.
Matches, not beats. The closing line is the strongest publicly observable forecast of an NFL game. Matching it — while adding honest, well-calibrated uncertainty around it — is the ceiling any transparent model should claim. Nothing on this page is betting advice; it is a measurement of forecast quality.
Methodology
Forecasts. Our model projects a final margin for every NFL game. Because a point estimate hides most of what matters, we treat each projection as the center of a distribution of outcomes, calibrated to the historical variance of NFL results (final margins deviate from the spread by roughly two touchdowns on average). The weekly forecast shows the full distribution: win probabilities, cover probabilities, and likely margin ranges.
Benchmarks. Every tracked model is scored the same way: one flat unit on each pick against the closing spread at standard −110 juice, pushes excluded. That makes 52.4% the break-even win rate, and profit in units the honest bottom line. We track 🏆 benchmark models (ours), 💰 paid services, and 👤 public models — the full field is here, scored identically back to 1999.
Data. Historical model picks are sourced from public archives such as The Prediction Tracker and the models' own published records. Lines are closing spreads. Benchmarks and forecasts are refreshed while the export pipeline runs; every page states the vintage of the data it shows.
Submit your model
Think you can beat the field? Submit your model and we'll score it exactly like everyone else — same lines, same juice, same leaderboard. Use the form below, or open it in a new tab.