Model evaluation
GW4 Evaluation: Inside the Data
Gameweek 4 gave us another useful test of Fantasy Matchup Edge. We are not interested in pretending that a projection model should predict every individual score exactly. Football is far too noisy for that. The more useful questions are whether the overall scoring level is calibrated, whether higher-ranked players outperform the wider pool, whether the probability forecasts behave sensibly, and whether the model can find players — or avoid players — that more conventional approaches might miss.
For the clean evaluation, we only include players who actually played 60+ minutes. FME projections are built on an IF STARTS basis, so bench cameos and DNPs are not treated as failed football predictions.
The headline from GW4:
- Overall FPL scoring was again very close to the model average.
- The top-ranked FPL players comfortably outscored the wider player pool.
- The expected number of 10+ FPL hauls was almost exact.
- Noah Okafor was one of the strongest calls of the week.
- Our very cold stance on Yoane Wissa paid off again.
- Sorare average scoring was also extremely well calibrated.
- Sorare’s upper tail was hotter than forecast — something we are continuing to watch.
- There were also some clear misses. We’ll show those too.
FPL
Clean sample: 197 players who played 60+ minutes.
| Metric | GW4 result |
|---|---|
| Average projected | 3.72 |
| Average actual | 3.99 |
| Bias | +0.27 |
| Mean absolute error | 2.56 |
| Median absolute error | 2.21 |
| Pearson correlation | 0.18 |
| Spearman correlation | 0.20 |
| Within ±2 points | 46.2% |
| Within ±3 points | 72.1% |
The first thing we look at is the population as a whole. Across almost 200 players, FME's average predicted points figure was 3.72. The actual average was 3.99.
That is a difference of just 0.27 points per player.
Individual FPL scores will always be noisy — one goal, penalty save, red card or late bonus swing can completely change a player's final total — and the single-week correlations reflect that. But at the population level, the scoring environment was captured very closely.
That matters because FME is not simply averaging historic FPL scores. The model is rebuilding expected scoring from the actions a player is likely to produce in that specific matchup, then running those actions through the FPL scoring system.
The highest-ranked players did better
Calibration is useful, but a ranking model also needs to separate the players we should actually care about.
It did.
- The eight players from FME's original top 10 averaged 6.50 FPL points.
- Everyone outside that original top 10 averaged 3.88.
- The 14 clean players from the original top 25 averaged 5.79.
- Everyone outside that top 25 averaged 3.85.
- The 24 clean players from the original top 50 averaged 5.38.
- Everyone outside that top 50 averaged 3.80.
So while GW4 was not a week where every high-ranked player exploded, there was still clear ranking separation. The top of the board produced materially more points than the broader population.
The 10+ haul probabilities were almost dead-on
This was arguably the nicest aggregate result of the week.
| Metric | Result |
|---|---|
| Expected 10+ FPL hauls | 12.52 |
| Actual 10+ FPL hauls | 13 |
| Expected 15+ hauls | 1.28 |
| Actual 15+ hauls | 3 |
The important distinction here is that probability models are not supposed to identify the exact 13 players who will haul. If Player A has a 20% haul probability and Player B has a 5% probability, both can still haul — or neither can.
What we want over the whole population is for those probabilities to add up to something resembling reality. In GW4, our 10+ probabilities did exactly that.
Okafor: one of the clearest calls of GW4
Noah Okafor
- FME rank: 2nd
- Predicted points: 6.40
- P90 ceiling: 12
- 10+ probability: 27.0%
- Actual: 9 points
Okafor scored and returned nine FPL points.
And this is exactly the kind of call we want FME to make. He was not simply sitting near the top because Leeds had been assigned a favourable green box on a fixture table. FME liked the way his own attacking profile matched what the Newcastle fixture could produce for him.
Why FPL fixture difficulty only tells part of the story
Traditional FPL fixture difficulty charts are useful. They tell you whether a team has a broadly attractive or unattractive matchup. But every player on that team does not score points in the same way.
A striker dependent on central shots, a wide attacker creating chances from the half-space, a progressive midfielder accumulating defensive contributions and a full-back generating crosses may all respond very differently to exactly the same opponent.
That is where a team-level fixture rating can miss something important.
GW4 gave us an almost perfect example. Newcastle were playing Leeds. A conventional fixture-first view could easily lead you towards the Newcastle forward and away from the Leeds attacker.
FME did effectively the opposite.
| Player | FME view | Actual |
|---|---|---|
| Noah Okafor | Rank 2 — 6.40 projected — 27.0% chance of 10+ | 9 |
| Yoane Wissa | Rank 407 — 3.06 projected — 2.7% chance of 10+ | 2 |
That is a huge difference in the model's view of two attackers involved in the same fixture. And Okafor delivered. Wissa blanked.
Wissa: the trap worked again
Wissa deserves a section of his own because we specifically highlighted him as one of our Traps of the Week.
- FME rank: 407
- Predicted points: 3.06
- P90 ceiling: 8
- 10+ probability: 2.7%
- Actual: 2
And he remains a genuinely popular FPL asset: the Fantasy Football Pundit ownership table had Wissa at 18.1% ownership on 16 September 2026.
This is one of the uses of the model that interests us most. Finding a good player is useful. Finding a popular player whose matchup does not fit his usual scoring routes can be just as useful.
The point is not that Wissa is a bad footballer or that he will continue blanking. It is that ownership and general reputation do not automatically make an individual gameweek a good matchup. GW4 was another week where FME wanted substantially less Wissa than the market.
Other strong FPL calls
There were several more good results near the top of the rankings.
- João Pedro: rank 3 — 6.08 projected → 12 actual
- Erling Haaland: rank 4 — 5.83 → 9
- Morgan Rogers: rank 5 — 5.79 → 8
- Morgan Gibbs-White: rank 6 — 5.78 → 8
- Virgil van Dijk: rank 9 — 5.52 → 8
- Anan Khalaili: rank 10 — 5.31 → 7
- Gabriel Magalhães: rank 15 — 5.12 → 9
- Maxim De Cuyper: rank 32 — 5.05 → 11
There is a nice mix there. Some obvious premium assets should rank highly. A good model should not avoid obvious players merely to look clever. But Okafor, Khalaili and De Cuyper are much more interesting demonstrations of the matchup layer finding something beyond the standard template.
Useful cooler stances
Wissa was the clearest outright trap call, but he was not the only popular attacker the model was relatively cool on.
- Bryan Mbeumo: rank 169 — 3.91 projected → 2
- Alexander Isak: rank 58 — 4.65 projected → 2
Neither returned. That does not mean every player outside the top 50 should be sold. It means there was a substantial gap between the players FME believed had the strongest GW4 matchup and some much more widely discussed fantasy names.
That is precisely why we prefer player-level matchup analysis to simply following ownership, recent points or a fixture colour.
And the misses
There were plenty. We are going to keep publishing these because a locked projection board is only useful if we are willing to show the bad calls as well as the good ones.
- Christos Tzolis: rank 1 — 6.49 projected → 3
- Jack Rudoni: rank 7 — 5.64 → 2
- Reece James: rank 8 — 5.63 → 1
- Kai Havertz: rank 11 — 5.24 → 1
- Bruno Fernandes: rank 14 — 5.14 → 2
Those are genuine misses.
And there were also several large hauls that FME did not identify particularly well.
- Pascal Groß: 17 points from model rank 368
- Jayden Bogle: 15 from rank 322
- Leif Davis: 14 from rank 462
- Joško Gvardiol: 11 from rank 292
- Vitalii Mykolenko: 11 from rank 458
We cannot call those FME successes. They weren't.
That is also why we care more about population-level calibration and ranking separation than cherry-picking screenshots of whichever player happened to score.
DEFCON was much closer this week
GW4's clean non-goalkeeper sample produced:
| Metric | Result |
|---|---|
| Expected DEFCON hits | 36.63 |
| Actual DEFCON hits | 33 |
That's reasonably close. DEFCON had been one of the clearest things we wanted to keep monitoring after the earlier evaluation, so this week's result is encouraging rather than something that currently demands a model change.
Sorare
Clean sample: 205 players who played 60+ minutes.
| Metric | GW4 result |
|---|---|
| Average projected | 50.42 |
| Average actual | 51.01 |
| Bias | +0.59 |
| MAE | 15.13 |
| Median absolute error | 13.10 |
| Pearson correlation | 0.27 |
| Spearman correlation | 0.28 |
| Within ±15 | 54.6% |
| Within ±20 | 70.2% |
Once again, the overall Sorare scoring environment landed extremely close. Average projection: 50.42. Average actual: 51.01. A difference of just 0.59 points across 205 players.
The higher Sorare ranks separated again
Among the original rankings, the 10 clean survivors from the original top 20 averaged 67.71. Four scored 70+ and two scored 80+.
The 16 clean survivors from the original top 30 averaged 66.99, with seven hitting 70+.
Across the whole clean sample, the actual average was only 51.01.
So the model was again successfully concentrating stronger Sorare scores towards the upper part of the rankings.
Some excellent Sorare calls
- James Garner: rank 9 — 62.62 projected → 91.90
- Gabriel Magalhães: rank 6 — 63.12 → 90.26
- Matheus Nunes: rank 13 — 61.61 → 85.30
- Bukayo Saka: rank 10 — 62.28 → 79.00
- Virgil van Dijk: rank 3 — 64.28 → 74.90
- João Pedro: rank 15 — 61.37 → 75.30
Those are exactly the kinds of players we want the upper rankings to contain.
Sorare's upper tail was hotter than the model again
This is probably the clearest modelling watch item from GW4.
| Threshold | Expected | Actual |
|---|---|---|
| 50+ | 96.90 | 96 |
| 60+ | 53.33 | 68 |
| 70+ | 26.14 | 38 |
| 80+ | 11.80 | 15 |
The 50+ calibration was almost perfect. But once we move into the upper tail, reality was noticeably hotter than forecast.
There is another useful check here. A P90 score is, by definition, intended to be exceeded roughly 10% of the time. In GW4, 34 of 205 players exceeded their P90. That's 16.6%.
So there is now further evidence that the Sorare simulation may be a little conservative in the upper tail. That does not mean we immediately tweak the model after one gameweek. Football results themselves can run hot. But it is something worth continuing to track.
Sorare misses
Again, the top of the rankings was not universally successful.
- Bruno Fernandes: rank 1 — 78.09 projected → 53.80
- Reece James: rank 2 — 69.19 → 59.48
- Christos Tzolis: rank 11 — 61.70 → 42.60
- Ross Barkley: rank 17 — 59.92 → 34.00
- Kai Havertz: rank 20 — 58.79 → 28.50
And some of the biggest scores of the week came from players the model did not rank especially highly.
- Pascal Groß: 100 from rank 30
- Jeremy Jacquet: 100 from rank 79
- Bart Verbruggen: 91.7 from rank 142
Those are not predictions we are going to retrofit into successes after the event.
What do we take from GW4?
The most encouraging result is probably not any individual player. It is the broader shape of the forecasts.
Across 197 FPL players, average predicted points were only 0.27 away from average actual points. Across 205 Sorare players, average projected score was only 0.59 away.
The FPL model expected 12.52 10+ hauls and got 13. The highest-ranked FPL players materially outscored the wider pool. And there were some genuinely useful player-level calls.
Okafor at number two was excellent.
Wissa at number 407 — and on our Traps of the Week list — was arguably just as interesting.
That Okafor/Wissa contrast is also a good illustration of what we want Fantasy Matchup Edge to do differently. A normal FPL fixture difficulty view starts with the team.
FME starts with the player: how he scores, what actions his opponent tends to allow, where those actions occur, the likely match environment and what that combination means for his fantasy scoring distribution.
Sometimes the two approaches will give you the same answer. Sometimes they won't.
GW4 gave us a pretty good example of why that difference can matter.
There were also clear misses, and the Sorare upper tail remains something we want to monitor. Four gameweeks into a season, none of this proves a model is “solved”.
But the process is beginning to give us exactly what we wanted from it: a locked set of pre-match projections, measurable probabilities, transparent results afterwards, and matchup-specific calls we can actually test.
And then we do it again next week.
Want to see how FME turns player action profiles, opponent tendencies and match environment into projections? Read how our model works →