Model evaluation
GW3 Review: The Model Faded Some of FPL’s Most Popular Picks
Gameweek 3 gave us another proper opportunity to put Fantasy Matchup Edge under the microscope. We locked the pre-game projections, compared them with what actually happened, and looked beyond a handful of successful picks to ask whether the model was calibrated, whether its rankings separated stronger from weaker opportunities, and whether it spotted places where FPL ownership looked more optimistic than the matchup data.
- Overall FPL and Sorare scoring levels were very closely calibrated.
- The simulation model came very close to the actual number of FPL hauls.
- Higher-rated FPL players materially outscored the lower-rated half.
- The exact location of the biggest hauls remained noisy.
- Several highly-owned players were ranked very poorly by FME — and then returned poorly in GW3.
One important reminder before the numbers: FME projections are IF STARTS. They are not generic minutes forecasts and they are not multiplied by start probability. For the cleanest accuracy tests below, we therefore focus on players who logged at least 60 minutes. DNPs and short cameos still matter to fantasy managers, but they are not clean tests of an IF-STARTS performance projection.
The headline numbers
| Metric | FPL | Sorare |
|---|---|---|
| Clean sample | 212 | 214 |
| Average projection | 3.77 | 50.51 |
| Average actual | 3.73 | 49.89 |
| Bias (actual − projected) | −0.04 | −0.62 |
| Mean absolute error | 2.21 | 14.17 |
| RMSE | 2.85 | 17.20 |
| Spearman rank correlation | 0.264 | 0.144 |
The aggregate calibration was extremely close. Across the clean FPL sample, the model projected an average of 3.77 points and the players actually scored 3.73. On Sorare, the model projected 50.51 and the actual average was 49.89.
That does not mean every individual projection was close — football scoring is far too noisy for that. The individual errors and relatively modest rank correlations make that clear. But it is a strong sign that the model estimated the overall scoring environment of the gameweek well.
The probability model got the number of hauls remarkably close
One of the main reasons we simulate each match rather than publishing only a projected mean is that fantasy decisions are about distributions. A 4.5-point projection with serious haul upside is not necessarily the same asset as another 4.5-point projection with a narrower range of outcomes.
FPL haul calibration
| Outcome | Model expected | Actual |
|---|---|---|
| 10+ point scores | 14.10 | 13 |
| 15+ point scores | 1.44 | 1 |
That is very encouraging at the slate level. The model did not know exactly which players would supply those hauls — and several of the highest individual haul probabilities disappointed — but it was extremely close on how many big scores the gameweek would produce.
For example, Christos Tzolis entered the week with a 30.2% chance of 10+, Erling Haaland 28.2%, Bruno Fernandes 28.1% and Morgan Gibbs-White 27.0%. Their actual FPL returns were 5, 9, 2 and 3. This was a useful reminder that a probability is not a promise: the aggregate distribution can be right even when the exact identity of the winners is noisy.
Sorare threshold calibration
On the public ranking population that played 60+ minutes, the Sorare threshold probabilities were also close to the realised totals:
| Outcome | Model expected | Actual |
|---|---|---|
| 60+ scores | 56.2 | 59 |
| 70+ scores | 27.9 | 31 |
Did higher-ranked players actually score more?
Broadly, yes — and the separation was much clearer on FPL than Sorare.
| FPL group | Avg projection | Avg actual | 10+ hit rate |
|---|---|---|---|
| Higher-projected half | 4.42 | 4.41 | 8.5% |
| Lower-projected half | 3.12 | 3.06 | 3.8% |
The higher-rated half scored around 44% more FPL points on average and was more than twice as likely to produce a 10+ return.
| Sorare group | Avg projection | Avg actual | 60+ rate |
|---|---|---|---|
| Higher-projected half | 56.05 | 52.45 | 29.9% |
| Lower-projected half | 44.96 | 47.33 | 26.2% |
The Sorare split was weaker. The very top still contained useful scores — the ten highest projected clean players averaged 60.5 actual Sorare points — but the overall rank ordering needs a larger sample before we draw strong conclusions.
Some of the calls that worked
FPL
| Player | FME projection | Actual | 10+ chance |
|---|---|---|---|
| Erling Haaland | 7.17 | 9 | 28.2% |
| Luka Vušković | 5.08 | 12 | 13.0% |
| Harvey Barnes | 4.68 | 12 | 13.4% |
| Joško Gvardiol | 5.10 | 8 | 10.4% |
| Kai Havertz | 5.50 | 8 | 17.0% |
| Virgil van Dijk | 5.47 | 6 | 16.6% |
Vušković is a particularly useful example of what we are trying to surface: a strong matchup projection that was not simply an obvious premium fantasy name. Harvey Barnes was another useful mid-price call, returning 12 from a 4.68 projection.
Sorare
| Player | Projection | Actual |
|---|---|---|
| Christos Tzolis | 63.9 | 82.5 |
| Virgil van Dijk | 64.2 | 89.6 |
| Kai Havertz | 59.4 | 82.0 |
| Cody Gakpo | 59.4 | 91.7 |
| Marc Guéhi | 63.9 | 71.2 |
| Erling Haaland | 65.8 | 72.2 |
And the misses
There is no value in reviewing a model if we only show the winners.
| Player | Projection | Actual |
|---|---|---|
| Bruno Fernandes | 6.45 | 2 |
| João Pedro | 5.39 | 1 |
| Morgan Gibbs-White | 6.42 | 3 |
| Dominic Calvert-Lewin | 4.57 | 1 |
| Cole Palmer | 4.51 | 1 |
Bruno was the clearest high-end miss across both platforms. He was our third-highest FPL projection and the highest Sorare projection at 84.7, but finished with 2 FPL points and 35.4 on Sorare.
At the other end, the model badly underestimated some surprise hauls. Tyrick Mitchell was projected at only 2.92 FPL points and scored 15; Jayden Bogle was projected at 2.94 and scored 14. Those are misses, not results to explain away after the fact. The question over a larger sample is whether outcomes like these occur at approximately the frequency the simulation distributions imply.
The interesting bit: popular players FME actively didn’t like
This may eventually be one of the most useful parts of Fantasy Matchup Edge. Fantasy analysis naturally focuses on finding players to buy, but avoiding a heavily-owned asset in a weak matchup can be just as valuable.
Using the ownership snapshot supplied for this review, several popular assets entered GW3 with surprisingly poor FME projections — and many then returned very little. The ownership figures below are a review snapshot, not a claim that they exactly reproduce the GW3 deadline ownership.
| Player | Ownership | FME public rank | Projection | GW3 pts |
|---|---|---|---|---|
| Riccardo Calafiori | 47.4% | 233 | 3.53 | 2 |
| Yoane Wissa | 16.9% | 293 | 3.24 | 1 |
| Pascal Groß | 16.2% | 285 | 3.26 | 1 |
| Harry Maguire | 12.6% | 237 | 3.52 | 2 |
| Senne Lammens | 12.5% | 256 | 3.44 | 2 |
| Anthony Elanga | 10.9% | 286 | 3.26 | 1 |
| Issa Diop | 14.9% | 359 | 2.74 | 3 |
| Milan van Ewijk | 11.2% | 387 | 1.71 | 3 |
Wissa is probably the clearest example
Yoane Wissa was 16.9% owned overall in the supplied snapshot and appeared in plenty of template-looking squads. FME was nowhere near as enthusiastic.
- FME public-board rank: 293rd
- Full FPL model rank: 383rd of 538 mapped players
- Forward rank: 41st of 58
- FPL projection: 3.24
- 10+ haul probability: 3.2%
- GW3 result: 1 point in 90 minutes
That is exactly the sort of disagreement with the fantasy meta that we want the model to expose. It also was not simply a case of there being nothing attractive around his budget.
| Player | Position | Pre-GW3 price | FME projection | GW3 actual |
|---|---|---|---|---|
| Christos Tzolis | MID | £6.5m | 6.79 | 5 |
| Joško Gvardiol | DEF | £5.6m | 5.10 | 8 |
| Mikkel Damsgaard | MID | £5.5m | 5.08 | 8 |
| Marc Guéhi | DEF | £6.0m | 4.98 | 8 |
| Harvey Barnes | MID | £6.0m | 4.68 | 12 |
| Yoane Wissa | FWD | about £6m | 3.24 | 1 |
These are not all direct positional swaps — FPL squad construction still matters — but they make the value point clearly. An affordable, popular player is not automatically good value. FME saw substantially stronger expected returns elsewhere in the same broad price bracket.
Don’t just tell me who projects well. Tell me which popular players the matchup model thinks I shouldn’t be following.
A positional pattern to monitor
There was also a position-level pattern worth keeping an eye on rather than reacting to after one week.
| Position | Avg actual minus projection |
|---|---|
| Defenders | +0.51 |
| Goalkeepers | +0.32 |
| Midfielders | −0.40 |
| Forwards | −1.12 |
Sorare showed a similar directional pattern: forwards came in around 3.5 points below projection on average while goalkeepers finished roughly 4.8 above. One gameweek is far too small a sample to recalibrate anything from this, but if it persists it may point to a position-specific scoring or decisive-event adjustment worth investigating.
What did we learn from GW3?
The biggest positive is not that a particular player hit. It is that the distribution of the gameweek was broadly where the model said it should be. Average FPL and Sorare scoring were almost perfectly centred, expected FPL haul counts were close to the actual totals, and Sorare threshold probabilities were similarly encouraging.
The ranking signal was useful, especially in FPL, and the model produced an interesting group of popular-player fades where ownership appeared considerably stronger than the matchup projection.
But GW3 also showed where we need more evidence. The highest individual haul probabilities did not produce the actual hauls, Sorare rank ordering was relatively weak despite good overall calibration, and individual misses such as Bruno Fernandes, Tyrick Mitchell and Jayden Bogle were large.
We should not change a simulation model because of one surprising football match. We should track whether those misses become patterns. That is what these weekly reviews are for.
Onto GW4
The GW4 model has now been refreshed with another full week of current-season evidence, updated team and opponent profiles, fresh market environments and another 100,000 simulations.
The goal is not to produce a perfect list every week. It is to build a model that consistently tells us which matchups create the most fantasy opportunity, which players are best equipped to exploit them, where the market or fantasy ownership appears too optimistic, and where less obvious value exists elsewhere.
GW3 was another useful step in that direction.