FPLRogue model accuracy and release decisions
We publish the tests behind model changes, including regressions and candidates we hold back. Lower error is useful, but a release must also improve the decisions FPL managers actually make.
xMins V11 release evaluation
V11 was compared with the production-policy control using only information available before each historical Gameweek. It passed all 55 release gates and improved every headline error measure below.
| Metric | Control | V11 | Change |
|---|---|---|---|
| xMins MAE | 21.451 min | 20.035 min | −1.416 min |
| xMins RMSE | 30.053 min | 28.810 min | −1.243 min |
| Start Brier score | 0.12915 | 0.12305 | −0.00610 |
| Calibration error (ECE) | 0.07159 | 0.06315 | −0.00844 |
| Predicted ≥60, played 0 | 3,716 | 3,025 | −691 |
| Within 15 minutes | 52.49% | 62.46% | +9.97 pp |
xPts V2 candidate evaluation
The candidate reduced point error and improved rank correlation, but its top-10 hit rate fell. Because one decision-quality gate failed, the candidate was not promoted.
| Metric | Legacy | V2 candidate | Change |
|---|---|---|---|
| xPts MAE | 3.1473 pts | 2.7170 pts | −0.4303 pts |
| xPts RMSE | 3.7520 pts | 3.5180 pts | −0.2340 pts |
| Prediction bias | +1.4572 pts | +0.5551 pts | −0.9021 pts |
| Mean Spearman rank | 0.1711 | 0.2129 | +0.0418 |
| Top-10 hit rate | 27.03% | 25.95% | −1.08 pp |
Model decision log
- xMins V11 moved to current-Gameweek shadow review
Passed 55 of 55 release gates across error, calibration, position, season, Gameweek-band and bootstrap checks. No automatic production promotion was made by this test.
- xPts V2 remained on hold
RMSE and overall ranking improved, but the top-10 hit-rate gate regressed. The public table above retains that negative result.
What these results do not prove
The xMins replay does not contain historical injury flags, suspension flags, predicted-lineup snapshots or today’s manual overrides. Some older seasons use 60 minutes as a start proxy. The xPts test covers one held-out season and cannot guarantee future performance. Actual FPL points remain noisy, so these metrics should be read as model-comparison evidence rather than promises.
For reuse or reporting, cite the evaluation date, sample, holdout period and release outcome—not just the lowest error number.


