Fantasy Matchup Edge

← Back to Blog

Model evaluation

GW2 Evaluation: Inside the Data

The model is only 2 weeks into the season, so nothing is for certain. But there were some extremely encouraging signs. Not “the model is proven after one week,” but there is real evidence that:

FPL

Clean sample: 207 recorded starters who played 60+ minutes.

MetricResult
Average projected:3.74
Average actual:3.67
Bias:−0.07
MAE2.39
Median absolute error2.00
Pearson correlation0.37
Spearman correlation0.28
Within ±2 points50.2%
Within ±3 points72.5%

This is especially encouraging. For both FPL and Sorare, we aren’t averaging averages, we’re stripping down a player’s actions and rebuilding from zero. We model based on what the opponent should allow, and what actions players should get, in that specific match environment, area of the pitch, and more. Then we sim each match thousands of times: and all this comes out to a projected average that is extremely close to mean scores across the board for that GW2. Sorare’s Average was also extremely close (50.83 Projected vs 50.47 Actual). More on that further down…

The aggregate probability calibration was particularly good:

MetricResult
Expected 10+ hauls:13.28
Actual 10+ hauls:14
Expected 15+ hauls:1.33
Actual 15+ hauls:1

Remember, we can’t predict exactly who is going to break out and haul for every match, we can only present the likelihood we believe a player may do so, but overall we basically matched the expected number of hauls across the gameweek.

Ranking separation was strong:

GroupResult
Top-25 clean survivors (players who played over 60mins):8.08 average points, with 5/12 hauling.
Everyone outside that top 25:3.39 average, with 9/195 hauling.

That is a 41.7% versus 4.6% haul rate, although the top group is obviously still small.

The four clean survivors from the original top 10 averaged 12.75, with three hauling.

Genuine non-obvious calls included:

Useful cooler stances included Wirtz, Gakpo, Marmoush and Rice, none of whom hauled.

The honest major misses were Virgil van Dijk, Havertz and Okafor. Doku and Watkins were DNPs; Tzolis and Reece James were partial-minute cases, so none are counted as clean projection errors. But Tzolis cannot be claimed as a hit, with 0 points before being pulled at half time. We think that Tzolis’ projections are high because of his history in Belgium, but all foreign leagues are weighted down for relevance and confidence. No other new Premier League player has such outlier projections, so the model sees something in him. But we’ll be keeping an eye out over the next few weeks, especially as he plays more, and the model learns more about him and his profile in English football.

Pascal Groß, Tarkowski, Calafiori and Neco Williams produced good FPL results, but they were not highly ranked by the model. We cannot cleanly say “FME called it” with these players. We didn’t.

Sorare

Clean sample: 203 matched starters who played 60+ minutes.

MetricResult
Average projected50.83
Average actual50.47
Bias−0.36
MAE14.53
Median absolute error12.39
Pearson correlation0.31
Spearman correlation0.25
Within ±1559.1%
Within ±2073.9%

The top-end ranking performance was excellent:

GroupResult
Top 10:six clean survivors; 5/6 scored 70+, averaging 78.43
Top 20:7/12 scored 70+
Top 30:9/17 scored 70+
Outside the top 10,28/197 scored 70+: 14.2%

So the clean full-population comparison is:

Sorare top-10: 83% hit 70+. Everyone else: 14%.

Strong calls beyond the obvious names included Marc Guéhi, Elliot Anderson, Rúben Dias, João Pedro, Gibbs-White and Xhaka. Truffert was an especially good upside hit: 57.31 projected, 79.33 P90, 94.10 actual.

Honest misses included:

Sorare’s upper tail was hotter than predicted:

MetricResult
Expected 60+ scores:53.29; actual: 68
Expected 70+ scores:26.26; actual: 32
P90 exceedance:26/201 = 12.9%, versus nominal 10%

That suggests the Sorare simulation tail may be slightly conservative. It is a monitoring observation, after one week of data, no kneejerk tweaking the model, but we’re keeping an eye on it.

The watch item

DEFCON expected 38.37 clean-sample hits but produced 29. That is the one obvious underperformance relative to probability. We’re going to track this across subsequent GWs before changing anything.

Overall, for a model that’s only a couple of weeks old, we’re really encouraged by the start. More to come very soon…