Transparent performance across 1358 NBA regular-season games — every one of them out-of-sample (played after the last game the model was trained or calibrated on, i.e. since 2025-03-29).
Probabilistic forecast quality — lower is better
0.00 = perfect probabilistic forecasting (model is right with 100% confidence every time)
0.25 = random guessing (50/50 every game)
0.2156 = our model — measurably better than random, in line with published NBA-model benchmarks
For each predicted-probability bin, here is the model's confidence vs the actual win rate of games in that bin. Closer numbers = better calibrated.
| Predicted bin | Model says | Actually won | Games |
|---|---|---|---|
| 00-10% | 3.4% | 23.0% | 61 |
| 10-20% | 15.3% | 26.7% | 15 |
| 20-30% | 22.0% | 35.7% | 14 |
| 30-40% | 31.6% | 26.1% | 138 |
| 40-50% | 42.3% | 38.4% | 307 |
| 50-60% | 55.5% | 61.9% | 431 |
| 70-80% | 73.6% | 70.9% | 165 |
| 80-90% | 84.6% | 79.9% | 219 |
| 90-100% | 100.0% | 100.0% | 8 |
Bins with fewer than ~50 games are noisier and should be read with caution. Green = within 5pp of perfect calibration, yellow = within 10pp, orange = larger gap.
Each prediction is driven by these seven inputs. The bars below are XGBoost gain importances — how much each feature contributes to the model's decisions. See How It Works for what each one means.
| Date | Matchup | Score | Pick | Conf | Result |
|---|---|---|---|---|---|
| 04-12T00:00:00 | BKN @ TOR | 101-136 | TOR | 86.2% | ✓ |
| 04-12T00:00:00 | PHX @ OKC | 135-103 | OKC | 86.2% | ✗ |
| 04-12T00:00:00 | DEN @ SAS | 128-118 | SAS | 73.3% | ✗ |
| 04-12T00:00:00 | SAC @ POR | 110-122 | POR | 83.3% | ✓ |
| 04-12T00:00:00 | MIL @ PHI | 106-126 | PHI | 57.9% | ✓ |
| 04-12T00:00:00 | DET @ IND | 133-121 | DET | 100.0% | ✓ |
| 04-12T00:00:00 | CHA @ NYK | 110-96 | NYK | 51.5% | ✗ |
| 04-12T00:00:00 | NOP @ MIN | 126-132 | MIN | 73.3% | ✓ |
| 04-12T00:00:00 | ATL @ MIA | 117-143 | ATL | 69.3% | ✗ |
| 04-12T00:00:00 | UTA @ LAL | 107-131 | LAL | 86.2% | ✓ |
| 04-12T00:00:00 | GSW @ LAC | 110-115 | LAC | 71.1% | ✓ |
| 04-12T00:00:00 | MEM @ HOU | 101-132 | HOU | 86.2% | ✓ |
| 04-12T00:00:00 | CHI @ DAL | 128-149 | DAL | 51.5% | ✓ |
| 04-12T00:00:00 | WAS @ CLE | 117-130 | CLE | 86.2% | ✓ |
| 04-12T00:00:00 | ORL @ BOS | 108-113 | BOS | 80.2% | ✓ |
| 04-10T00:00:00 | DET @ CHA | 118-100 | CHA | 57.9% | ✗ |
| 04-10T00:00:00 | MIA @ WAS | 140-117 | MIA | 57.7% | ✓ |
| 04-10T00:00:00 | MEM @ UTA | 101-147 | UTA | 51.5% | ✓ |
| 04-10T00:00:00 | DAL @ SAS | 120-139 | SAS | 90.0% | ✓ |
| 04-10T00:00:00 | GSW @ SAC | 118-124 | SAC | 57.9% | ✓ |
Our model is an XGBoost classifier trained with seven features per game: ELO rating gap, rolling 20-/10-/5-game point differentials, rest-day gap, and back-to-back flags for each team. The train/calibration/test split is chronological — the model never trains on a game played after one it's evaluated on.
Raw model probabilities are calibrated using isotonic regression fit on a later, held-out 20% of games — that's why our published confidence numbers (e.g. "75%") match observed win rates in the table above.
Predictions are evaluated on a straight-up winner basis (did the model pick the winning team?). Only regular-season games where both teams had at least 5 games of completed history are included.
For the full walkthrough of features, training pipeline, and limitations, see How It Works.
The same model behind the 68.1% accuracy above — delivered to your inbox each morning. Free.