Methodology · injury-risk-v2
What the injury-risk number can—and cannot—say
It is a preseason statistical estimate of season-long football availability. It is not a medical prediction and does not claim to know more than a team’s medical staff.
Preseason season-level holdout
Fit on 2021, 2022, 2023 only; tested on 2024. 9,646 held-out player-team games from 29,890 training opportunities. Lower Brier is better; higher AUC is better.
| Method | AUC | Brier |
|---|---|---|
| Base per-game rate (training) | 0.5000 | 0.2015 |
| Preseason season estimate | 0.6863 | 0.1819 |
The preseason estimate is compared with the training base-rate predictor on the held-out outcomes and beat that baseline, so the product is labelled an injury-risk estimate. The build automatically relabels the product as availability context if that ceases to be true.
Calibration
Player-seasons are grouped by predicted preseason per-game rate. The largest bucket-level gap is 12.5 percentage points.
| Predicted bucket | Player-seasons | Predicted per-game rate | Realized miss rate |
|---|---|---|---|
| 7–11% | 68 | 8.8% | 16.2% |
| 11–13% | 68 | 12.0% | 13.1% |
| 13–17% | 67 | 14.8% | 17.2% |
| 17–20% | 68 | 18.3% | 19.8% |
| 20–26% | 67 | 23.0% | 20.9% |
| 26–32% | 68 | 28.3% | 29.5% |
| 32–33% | 9 | 32.1% | 19.7% |
| 33–33% | 126 | 32.5% | 35.1% |
| 33–47% | 65 | 39.0% | 46.1% |
| 47–92% | 71 | 62.0% | 59.2% |
Evidence boundary
The availability outcome starts with weekly rostered regular-season team-game opportunities, preserves explicitly injury/reserve-coded weeks, right-censors audited transaction/non-injury departure states, and marks participation only from a canonical weekly-stat row or a positive canonicalized football snap. Every scheduled eligible team game must be covered by the snap feed; unknown roster-status pairs and ambiguous team weeks fail the run, while rare unprovable identities (crosswalk-absent fringe players, mid-season position conversions, contested identity mappings) are quarantined whole within fail-closed budgets and never silently labeled. Season-level covariates use only prior seasons: outcome history and recent availability from earlier rostered opportunities, trailing snap share and age from prior feature observations. No game-week injury designation or weekly roster status is a predictor.
Expected games missed is an explicit per-game-rate × remaining-games extrapolation, not a diagnosis.