How the predictions work
How each weekly projection is produced, who gets one, and how I check whether the projections have been any good.
Summary
- A projection is built mostly from what the player has done recently. That mattered far more than anything else I tried adding.
- Every projection has a range, because one game is too noisy for a single number to mean much.
- Projections are written down before kickoff and never changed, so the accuracy record reflects real forecasts.
- A change to the model only goes in if it beats the current one on past seasons. Changes that fail are published too.
The analysis behind all of this is in the research paper.
The weekly cycle
The lock is the step that matters. Once a week's projections are written to disk they can't be rewritten, including by me. Without that, a later run could regenerate last week's numbers with a better model and I could publish the improved score as if it had been forecast in advance. The Report card would then be measuring nothing. I run the lock by hand for now, so the record starts once the first week has been locked and played.
How a projection is produced
A ridge regression and gradient boosting, averaged.
The inputs are recent scoring, how much of the team's work the player gets, the opponent, whether the game is at home, the betting line, and the roof. Weather is in the model too, but I haven't connected a forecast yet, so for games that haven't been played it uses typical values.
The two models are a ridge regression, which is linear and stable, and gradient boosting (LightGBM), which can pick up combinations of inputs. I average them. More elaborate approaches didn't help. On held-out weeks the ridge reached an R-squared of 0.280, boosting 0.288 and the average 0.289, against 0.232 for a plain recent average.
Prediction intervals
Quantile models with a split-conformal correction.
A projection of 24.3 points doesn't mean the player will score 24.3. The model's range for him is 10.7 to 34.0, and about four times in five the real score should land inside it.
I check that instead of assuming it. Two extra models estimate the 10th and 90th percentiles, and a conformal correction widens them by however much they turned out to be too narrow on weeks they hadn't seen. In the backtest the real score landed inside the published range 79.7% of the time, against a target of 80%. The Report card shows the same figure for live weeks.
Which players are projected
Training, calibration and publication use different groups of players.
Players with a small role don't get a projection. I compared the model with the simplest alternative, a player's own recent average, and below about four points a game the model did worse than that average. Publishing those projections would have made the site worse.
Players ruled Out or Doubtful are left off instead of shown at zero, because a zero reads as “he will play badly” when I mean “he isn't playing”. A player listed Questionable keeps his projection for if he plays, with the chance he does next to it. That chance comes from how often Questionable players played between 2021 and 2025: about 63% overall and 35% for quarterbacks.
The promotion gate
How a change is adopted, and why published accuracy can't regress.
Each week the pipeline reports where the model missed. I have Claude write one new feature from that report, as code, and the harness replays five past seasons with and without it, on weeks neither version was trained on.
The change is adopted only if it is clearly better and the ranges still cover about 80%. A tie keeps the current model, so published accuracy can hold or improve but not slip.
Changes that fail go on the Report card with their numbers. Most of them fail. The first feature I tried, a ratio of recent form to season average, improved 4,231 of 8,464 predictions, which is a coin flip, so it was rejected.
The achievable ceiling
Measured before the engine was built.
An oracle that knew each player's true season-long average in advance, which is more than any model can know, would still miss by about five fantasy points a game. A single game is dominated by things nobody can forecast: a tipped pass, a goal-line call, a fumble. My model is under 7% away from that bound, which is all the room there is for improvement.
Limitations
It aims at the likely outcome, not the extremes.
The biggest misses are almost always players who scored far more than expected. A 34-point game from someone averaging 8 isn't the most likely result, and the range is where that possibility shows up.
It only sees what is in the data.
A coach's comments, a scheme change nobody has recorded yet or a player's personal situation never reach it. It works from box scores, schedules, injury reports and betting lines.
Late scratches aren't counted against it.
A player ruled out shortly before kickoff isn't graded against the points model. That is an availability question, and I track it separately.
It is weakest early in the season.
Projections lean on recent form, and in week 1 there isn't much of it.
It only projects fantasy points.
I haven't built projections for yards or touchdowns. The research paper explains why they are harder to predict than volume.
Check it yourself
The Report card
Weekly accuracy against simple baselines, how often the ranges held, and every change I've tried, including the ones that failed.
See the recordThe research paper
What predicts a game and what doesn't, how much of it can be predicted at all, and the results that ruled out my first design.
Read the paperThis week's projections
Every projection with its range and, where relevant, the chance the player takes the field.
See the projections