Model evaluation
GW2 Evaluation: Inside the Data
The model is only 2 weeks into the season, so nothing is for certain. But there were some extremely encouraging signs. Not “the model is proven after one week,” but there is real evidence that:
- The overall scoring levels were well calibrated.
- Higher-ranked players materially outperformed lower-ranked players.
- FPL haul probabilities were excellent in aggregate.
- Sorare’s upper tail performed better than forecast.
- DEFCON is the clearest area to monitor.
FPL
Clean sample: 207 recorded starters who played 60+ minutes.
| Metric | Result |
|---|---|
| Average projected: | 3.74 |
| Average actual: | 3.67 |
| Bias: | −0.07 |
| MAE | 2.39 |
| Median absolute error | 2.00 |
| Pearson correlation | 0.37 |
| Spearman correlation | 0.28 |
| Within ±2 points | 50.2% |
| Within ±3 points | 72.5% |
This is especially encouraging. For both FPL and Sorare, we aren’t averaging averages, we’re stripping down a player’s actions and rebuilding from zero. We model based on what the opponent should allow, and what actions players should get, in that specific match environment, area of the pitch, and more. Then we sim each match thousands of times: and all this comes out to a projected average that is extremely close to mean scores across the board for that GW2. Sorare’s Average was also extremely close (50.83 Projected vs 50.47 Actual). More on that further down…
The aggregate probability calibration was particularly good:
| Metric | Result |
|---|---|
| Expected 10+ hauls: | 13.28 |
| Actual 10+ hauls: | 14 |
| Expected 15+ hauls: | 1.33 |
| Actual 15+ hauls: | 1 |
Remember, we can’t predict exactly who is going to break out and haul for every match, we can only present the likelihood we believe a player may do so, but overall we basically matched the expected number of hauls across the gameweek.
Ranking separation was strong:
| Group | Result |
|---|---|
| Top-25 clean survivors (players who played over 60mins): | 8.08 average points, with 5/12 hauling. |
| Everyone outside that top 25: | 3.39 average, with 9/195 hauling. |
That is a 41.7% versus 4.6% haul rate, although the top group is obviously still small.
The four clean survivors from the original top 10 averaged 12.75, with three hauling.
Genuine non-obvious calls included:
- Morgan Gibbs-White: rank 5 (4th on probables list), 6.18 projected → 13
- Rayan Cherki: rank 15 (7th on probables list), 5.45 → 14
- João Pedro: rank 16, 5.43 → 9
- Gabriel: rank 41, 4.77 → 8
- Dominic Calvert-Lewin: rank 47, 4.74 → 8
- Nordi Mukiele: rank 51, 4.72 → 9
- Kevin Schade: rank 97, 4.33 → 10
Useful cooler stances included Wirtz, Gakpo, Marmoush and Rice, none of whom hauled.
The honest major misses were Virgil van Dijk, Havertz and Okafor. Doku and Watkins were DNPs; Tzolis and Reece James were partial-minute cases, so none are counted as clean projection errors. But Tzolis cannot be claimed as a hit, with 0 points before being pulled at half time. We think that Tzolis’ projections are high because of his history in Belgium, but all foreign leagues are weighted down for relevance and confidence. No other new Premier League player has such outlier projections, so the model sees something in him. But we’ll be keeping an eye out over the next few weeks, especially as he plays more, and the model learns more about him and his profile in English football.
Pascal Groß, Tarkowski, Calafiori and Neco Williams produced good FPL results, but they were not highly ranked by the model. We cannot cleanly say “FME called it” with these players. We didn’t.
Sorare
Clean sample: 203 matched starters who played 60+ minutes.
| Metric | Result |
|---|---|
| Average projected | 50.83 |
| Average actual | 50.47 |
| Bias | −0.36 |
| MAE | 14.53 |
| Median absolute error | 12.39 |
| Pearson correlation | 0.31 |
| Spearman correlation | 0.25 |
| Within ±15 | 59.1% |
| Within ±20 | 73.9% |
The top-end ranking performance was excellent:
| Group | Result |
|---|---|
| Top 10: | six clean survivors; 5/6 scored 70+, averaging 78.43 |
| Top 20: | 7/12 scored 70+ |
| Top 30: | 9/17 scored 70+ |
| Outside the top 10, | 28/197 scored 70+: 14.2% |
So the clean full-population comparison is:
Sorare top-10: 83% hit 70+. Everyone else: 14%.
Strong calls beyond the obvious names included Marc Guéhi, Elliot Anderson, Rúben Dias, João Pedro, Gibbs-White and Xhaka. Truffert was an especially good upside hit: 57.31 projected, 79.33 P90, 94.10 actual.
Honest misses included:
- Boscagli: rank 11, 63.70 → 23.56
- Garner: rank 9, 64.55 → 39.50
- McGinn: rank 26, 60.78 → 26.20
- Havertz: rank 32, 60.23 → 39.50
- Pedro Porro: rank 44, 58.53 → 23.34
Sorare’s upper tail was hotter than predicted:
| Metric | Result |
|---|---|
| Expected 60+ scores: | 53.29; actual: 68 |
| Expected 70+ scores: | 26.26; actual: 32 |
| P90 exceedance: | 26/201 = 12.9%, versus nominal 10% |
That suggests the Sorare simulation tail may be slightly conservative. It is a monitoring observation, after one week of data, no kneejerk tweaking the model, but we’re keeping an eye on it.
The watch item
DEFCON expected 38.37 clean-sample hits but produced 29. That is the one obvious underperformance relative to probability. We’re going to track this across subsequent GWs before changing anything.
Overall, for a model that’s only a couple of weeks old, we’re really encouraged by the start. More to come very soon…