Skip to content
Player Performance and Predictions

How the predictions work

How each weekly projection is produced, who gets one, and how I check whether the projections have been any good.

Summary

  1. A projection is built mostly from what the player has done recently. That mattered far more than anything else I tried adding.
  2. Every projection has a range, because one game is too noisy for a single number to mean much.
  3. Projections are written down before kickoff and never changed, so the accuracy record reflects real forecasts.
  4. A change to the model only goes in if it beats the current one on past seasons. Changes that fail are published too.

The analysis behind all of this is in the research paper.

The weekly cycle

The weekly prediction cycleFour steps in a row, then a loop back. The model makes projections, they are locked before the first kickoff, the games are played, and every projection is graded against what happened. The results go on the Report card, and what I learn feeds the next week.Projectevery eligible playerLockbefore first kickoffGames happennothing can changeGradeagainst the real scoreReport cardwhat I learn feeds the next week
Once projections are locked they are never changed, so the Report card measures forecasts, not hindsight.

The lock is the step that matters. Once a week's projections are written to disk they can't be rewritten, including by me. Without that, a later run could regenerate last week's numbers with a better model and I could publish the improved score as if it had been forecast in advance. The Report card would then be measuring nothing. I run the lock by hand for now, so the record starts once the first week has been locked and played.

How a projection is produced

A ridge regression and gradient boosting, averaged.

How one projection is builtA player's recent form, usage, opponent and game context feed three models. A linear model and a tree model are averaged to give the projection. A pair of range models are widened to match how often they missed, which gives the 80% range. Separately, the injury report gives the chance the player takes the field.Inputsform, usage, opponent,betting line, roofLinear modelsteady, simpleTree modelfinds interactionsRange modelslow end and high endAverage of boththe projection24.3 pointsthe number you seeWiden the rangechecked on unseen weeks10.7 to 34.0the 80% rangeInjury reportpractice statusseparate modelChance of playingQuestionable players only
Two simple models, averaged. Nothing more elaborate helped in testing.

The inputs are recent scoring, how much of the team's work the player gets, the opponent, whether the game is at home, the betting line, and the roof. Weather is in the model too, but I haven't connected a forecast yet, so for games that haven't been played it uses typical values.

The two models are a ridge regression, which is linear and stable, and gradient boosting (LightGBM), which can pick up combinations of inputs. I average them. More elaborate approaches didn't help. On held-out weeks the ridge reached an R-squared of 0.280, boosting 0.288 and the average 0.289, against 0.232 for a plain recent average.

Prediction intervals

Quantile models with a split-conformal correction.

Why every projection carries a rangeA single number of 24.3 points looks precise. The 80% range for the same player runs from about 11 to 34 points, meaning four times out of five the real score lands somewhere in that band.What a single number implies“24.3 points”The model's range10.734.04 games out of 5 land in here
The same projection shown two ways. The single number is the middle of the range.

A projection of 24.3 points doesn't mean the player will score 24.3. The model's range for him is 10.7 to 34.0, and about four times in five the real score should land inside it.

I check that instead of assuming it. Two extra models estimate the 10th and 90th percentiles, and a conformal correction widens them by however much they turned out to be too narrow on weeks they hadn't seen. In the backtest the real score landed inside the published range 79.7% of the time, against a target of 80%. The Report card shows the same figure for live weeks.

Which players are projected

Training, calibration and publication use different groups of players.

Which players are used for whatThe model learns from every player with at least three previous games. It sets the width of its ranges using only the players good enough to be published. And it publishes only players averaging at least four fantasy points over their last five games.Learn from: everyone with 3+ games12,411 player-gamesShow, and judge on: a real role4+ points over the last five games8,464 player-gamesrange widths are set here tooWhy not just use everyone?Below that line the model was measurablyworse than a player's own recent average.A deep-bench player who scores near zeroevery week needs no projection, and addingone only made the numbers noisier.Learning from them still helps, so the model does.
Three groups of players doing three different jobs. Each split was tested.

Players with a small role don't get a projection. I compared the model with the simplest alternative, a player's own recent average, and below about four points a game the model did worse than that average. Publishing those projections would have made the site worse.

Players ruled Out or Doubtful are left off instead of shown at zero, because a zero reads as “he will play badly” when I mean “he isn't playing”. A player listed Questionable keeps his projection for if he plays, with the chance he does next to it. That chance comes from how often Questionable players played between 2021 and 2025: about 63% overall and 35% for quarterbacks.

The promotion gate

How a change is adopted, and why published accuracy can't regress.

How a change to the model is adoptedWhere the current model missed is summarised. One new idea is written as code and replayed against five past seasons alongside the current model. A decision follows: if the change is clearly better and its ranges still hold, it replaces the model. If not, it is recorded as a dead end. Both outcomes go on the Report card.Where it missedlast week, summarisedOne new ideawritten as real codeReplay 5 seasonsold model vs newClearly better?and do the ranges hold?noyesRecorded as a dead endIt replaces the modelboth outcomes go on the Report card
A change is adopted only if it beats the current model on weeks neither was trained on. A tie keeps the current model.

Each week the pipeline reports where the model missed. I have Claude write one new feature from that report, as code, and the harness replays five past seasons with and without it, on weeks neither version was trained on.

The change is adopted only if it is clearly better and the ranges still cover about 80%. A tie keeps the current model, so published accuracy can hold or improve but not slip.

Changes that fail go on the Report card with their numbers. Most of them fail. The first feature I tried, a ratio of recent form to season average, improved 4,231 of 8,464 predictions, which is a coin flip, so it was rejected.

The achievable ceiling

Measured before the engine was built.

How much room for improvement existsA simple recent average is off by about 5.4 fantasy points a game. My model is off by about 5.27. A perfect knower of each player's true average would still be off by about 5.07. The gap between my model and that limit is under seven percent.A simple recent averageoff by 5.44 pointsThis modeloff by 5.27 pointsKnowing the true averageoff by 5.07 pointsLower is better. The whole remaining gap between this model and perfect knowledge is 0.20 points.
Even a model that knew each player's true season-long average in advance would only be about 7% better than mine. Most of a single game can't be predicted.

An oracle that knew each player's true season-long average in advance, which is more than any model can know, would still miss by about five fantasy points a game. A single game is dominated by things nobody can forecast: a tipped pass, a goal-line call, a fumble. My model is under 7% away from that bound, which is all the room there is for improvement.

Limitations

It aims at the likely outcome, not the extremes.

The biggest misses are almost always players who scored far more than expected. A 34-point game from someone averaging 8 isn't the most likely result, and the range is where that possibility shows up.

It only sees what is in the data.

A coach's comments, a scheme change nobody has recorded yet or a player's personal situation never reach it. It works from box scores, schedules, injury reports and betting lines.

Late scratches aren't counted against it.

A player ruled out shortly before kickoff isn't graded against the points model. That is an availability question, and I track it separately.

It is weakest early in the season.

Projections lean on recent form, and in week 1 there isn't much of it.

It only projects fantasy points.

I haven't built projections for yards or touchdowns. The research paper explains why they are harder to predict than volume.

Check it yourself

The Report card

Weekly accuracy against simple baselines, how often the ranges held, and every change I've tried, including the ones that failed.

See the record

The research paper

What predicts a game and what doesn't, how much of it can be predicted at all, and the results that ruled out my first design.

Read the paper

This week's projections

Every projection with its range and, where relevant, the chance the player takes the field.

See the projections
Esc

Loading players…