About · Accuracy & Methodology

Honest benchmarks for a noisy market

We benchmark the performance of betting models — ours included — to separate real signal from marketing. Everything on this page is measured the hard way: on held-out seasons the model never trained on, against the closing line.

How accurate are we?

Honest answer: our projected margin matches the closing line — it does not beat it, and neither does anything else we track. What we add is calibration: probabilities that mean what they say, tie the devigged moneyline market, and beat the public consensus. Held-out evaluation, generated 2026-08-31.

held-out evaluation window

Projected margin · MAE

9.82pts

1,960 held-out games

Our calibrated stack
9.82
Closing line (market)
9.82
Best public model (DRatings)
10.10

Matches the closing line — the number nothing on the board beats. Matching it is the honest headline.

P(home cover) · calibration

0.6pp ECE

1,917 games with a posted spread

Our log loss
0.6930
Public consensus
0.7011

Probabilities mean what they say — and they beat the public consensus.

P(home win) · log loss

0.6080

1,954 games straight up

Our calibrated stack
0.6080
Devigged moneyline (market)
0.6083
Public consensus
0.6358

Ties the devigged moneyline market; clearly ahead of consensus.

Margin scoreboard — 2019–2025

Mean absolute error of the projected final margin, on the games every model on the board predicted. Lower is better; nothing beats the closing line by a meaningful amount — including us.

#ModelKindGamesMAERMSEATS
1BB final stack (ours, market-anchored)Ours1,8139.8212.7150.0%
2Closing line (market)Market1,8139.8412.73
3Computer Adjusted Line (market)Market1,8139.8512.7450.2%
4Betting Benchmarks deep_honest (ours)Ours1,8139.8712.7552.3%
5Midweek line (market)Market1,8139.8812.7748.0%
6Opening line (market)Market1,81310.0012.9351.6%
7DRatingsPublic1,79710.1013.0250.1%
8Dokter EntropyPublic1,81210.1213.0451.0%

Market rows are betting lines — they are the market number, so they see information no forecaster does. Our final stack is market-anchored (closing line plus a disciplined step toward the model), so tracking the line this closely is by construction; the hair's-width MAE difference versus the closing line is sample wobble, not an edge. Rows without margin numbers publish probabilities only in this export. ATS is win rate against the closing spread — note how tightly the whole field clusters around 50%.

Do the probabilities mean what they say?

A reliability diagram answers the only question that matters for a probability: when we say 70%, does it happen 70% of the time? Each dot is one bin of games (2019–2025, all held-out); a calibrated forecast sits on the diagonal. Dot area is the number of games in the bin — the small edge bins are noisy by nature.

P(home win) — predicted vs realized

0%0%25%25%50%50%75%75%100%100%PERFECT CALIBRATIONPredicted 8.7% → realized 25.0% · 12 games in [0.0, 0.1)12Predicted 15.8% → realized 16.4% · 67 games in [0.1, 0.2)67Predicted 25.8% → realized 21.6% · 204 games in [0.2, 0.3)204Predicted 36.3% → realized 38.8% · 356 games in [0.3, 0.4)356Predicted 42.8% → realized 50.4% · 129 games in [0.4, 0.5)129Predicted 57.0% → realized 52.7% · 294 games in [0.5, 0.6)294Predicted 63.8% → realized 60.3% · 395 games in [0.6, 0.7)395Predicted 74.2% → realized 72.5% · 284 games in [0.7, 0.8)284Predicted 84.6% → realized 87.5% · 176 games in [0.8, 0.9)176Predicted 92.7% → realized 94.6% · 37 games in [0.9, 1.0)37Predicted win probabilityRealized win rate

And P(home cover)? Every one of its 1,917 games lands in a single bin: predicted 48.6%, realized 49.1%. That's not a chart, it's a sentence — and it's the whole point: against a near-efficient closing spread, an honest cover probability is a coin flip, and ours says so. The +EV board's record lives earlier in the week, graded against opening lines — that edge decays as the market moves, which is exactly what these two numbers together predict.

The honest story

The projection is market-anchored. Our final stack starts from the closing line and takes a small, disciplined step toward the model's own number. That is why its margin error tracks the line so closely — by construction. On one matched subsample our MAE comes out a couple hundredths below the line's; that is sample wobble on a market-anchored forecast, not evidence of an edge, and we will not claim otherwise.

Calibration is the real product. Public models publish margins; almost none publish probabilities you can take at face value. Ours are calibrated on held-out data: cover probabilities with an expected calibration error under a percentage point, and win probabilities whose log loss ties the devigged moneyline market and clearly beats the public consensus.

Matches, not beats. The closing line is the strongest publicly observable forecast of an NFL game. Matching it — while adding honest, well-calibrated uncertainty around it — is the ceiling any transparent model should claim. Nothing on this page is betting advice; it is a measurement of forecast quality.

Methodology

Forecasts. Our model projects a final margin for every NFL game. Because a point estimate hides most of what matters, we treat each projection as the center of a distribution of outcomes, calibrated to the historical variance of NFL results (final margins deviate from the spread by roughly two touchdowns on average). The weekly forecast shows the full distribution: win probabilities, cover probabilities, and likely margin ranges.

Benchmarks. Every tracked model is scored the same way: one flat unit on each pick against the closing spread at standard −110 juice, pushes excluded. That makes 52.4% the break-even win rate, and profit in units the honest bottom line. We track 🏆 benchmark models (ours), 💰 paid services, and 👤 public models — the full field is here, scored identically back to 1999.

Data. Historical model picks are sourced from public archives such as The Prediction Tracker and the models' own published records. Lines are closing spreads. Benchmarks and forecasts are refreshed while the export pipeline runs; every page states the vintage of the data it shows.

Submit your model

Think you can beat the field? Submit your model and we'll score it exactly like everyone else — same lines, same juice, same leaderboard. Use the form below, or open it in a new tab.