Fantasy Matchup Edge

← Back to Blog

Model evaluation

GW4 Evaluation: Inside the Data

Gameweek 4 gave us another useful test of Fantasy Matchup Edge. We are not interested in pretending that a projection model should predict every individual score exactly. Football is far too noisy for that. The more useful questions are whether the overall scoring level is calibrated, whether higher-ranked players outperform the wider pool, whether the probability forecasts behave sensibly, and whether the model can find players — or avoid players — that more conventional approaches might miss.

For the clean evaluation, we only include players who actually played 60+ minutes. FME projections are built on an IF STARTS basis, so bench cameos and DNPs are not treated as failed football predictions.

The headline from GW4:

FPL

Clean sample: 197 players who played 60+ minutes.

MetricGW4 result
Average projected3.72
Average actual3.99
Bias+0.27
Mean absolute error2.56
Median absolute error2.21
Pearson correlation0.18
Spearman correlation0.20
Within ±2 points46.2%
Within ±3 points72.1%

The first thing we look at is the population as a whole. Across almost 200 players, FME's average predicted points figure was 3.72. The actual average was 3.99.

That is a difference of just 0.27 points per player.

Individual FPL scores will always be noisy — one goal, penalty save, red card or late bonus swing can completely change a player's final total — and the single-week correlations reflect that. But at the population level, the scoring environment was captured very closely.

That matters because FME is not simply averaging historic FPL scores. The model is rebuilding expected scoring from the actions a player is likely to produce in that specific matchup, then running those actions through the FPL scoring system.

The highest-ranked players did better

Calibration is useful, but a ranking model also needs to separate the players we should actually care about.

It did.

So while GW4 was not a week where every high-ranked player exploded, there was still clear ranking separation. The top of the board produced materially more points than the broader population.

The 10+ haul probabilities were almost dead-on

This was arguably the nicest aggregate result of the week.

MetricResult
Expected 10+ FPL hauls12.52
Actual 10+ FPL hauls13
Expected 15+ hauls1.28
Actual 15+ hauls3

The important distinction here is that probability models are not supposed to identify the exact 13 players who will haul. If Player A has a 20% haul probability and Player B has a 5% probability, both can still haul — or neither can.

What we want over the whole population is for those probabilities to add up to something resembling reality. In GW4, our 10+ probabilities did exactly that.

Okafor: one of the clearest calls of GW4

Noah Okafor

Okafor scored and returned nine FPL points.

And this is exactly the kind of call we want FME to make. He was not simply sitting near the top because Leeds had been assigned a favourable green box on a fixture table. FME liked the way his own attacking profile matched what the Newcastle fixture could produce for him.

Why FPL fixture difficulty only tells part of the story

Traditional FPL fixture difficulty charts are useful. They tell you whether a team has a broadly attractive or unattractive matchup. But every player on that team does not score points in the same way.

A striker dependent on central shots, a wide attacker creating chances from the half-space, a progressive midfielder accumulating defensive contributions and a full-back generating crosses may all respond very differently to exactly the same opponent.

That is where a team-level fixture rating can miss something important.

GW4 gave us an almost perfect example. Newcastle were playing Leeds. A conventional fixture-first view could easily lead you towards the Newcastle forward and away from the Leeds attacker.

FME did effectively the opposite.

PlayerFME viewActual
Noah OkaforRank 2 — 6.40 projected — 27.0% chance of 10+9
Yoane WissaRank 407 — 3.06 projected — 2.7% chance of 10+2

That is a huge difference in the model's view of two attackers involved in the same fixture. And Okafor delivered. Wissa blanked.

Wissa: the trap worked again

Wissa deserves a section of his own because we specifically highlighted him as one of our Traps of the Week.

And he remains a genuinely popular FPL asset: the Fantasy Football Pundit ownership table had Wissa at 18.1% ownership on 16 September 2026.

This is one of the uses of the model that interests us most. Finding a good player is useful. Finding a popular player whose matchup does not fit his usual scoring routes can be just as useful.

The point is not that Wissa is a bad footballer or that he will continue blanking. It is that ownership and general reputation do not automatically make an individual gameweek a good matchup. GW4 was another week where FME wanted substantially less Wissa than the market.

Other strong FPL calls

There were several more good results near the top of the rankings.

There is a nice mix there. Some obvious premium assets should rank highly. A good model should not avoid obvious players merely to look clever. But Okafor, Khalaili and De Cuyper are much more interesting demonstrations of the matchup layer finding something beyond the standard template.

Useful cooler stances

Wissa was the clearest outright trap call, but he was not the only popular attacker the model was relatively cool on.

Neither returned. That does not mean every player outside the top 50 should be sold. It means there was a substantial gap between the players FME believed had the strongest GW4 matchup and some much more widely discussed fantasy names.

That is precisely why we prefer player-level matchup analysis to simply following ownership, recent points or a fixture colour.

And the misses

There were plenty. We are going to keep publishing these because a locked projection board is only useful if we are willing to show the bad calls as well as the good ones.

Those are genuine misses.

And there were also several large hauls that FME did not identify particularly well.

We cannot call those FME successes. They weren't.

That is also why we care more about population-level calibration and ranking separation than cherry-picking screenshots of whichever player happened to score.

DEFCON was much closer this week

GW4's clean non-goalkeeper sample produced:

MetricResult
Expected DEFCON hits36.63
Actual DEFCON hits33

That's reasonably close. DEFCON had been one of the clearest things we wanted to keep monitoring after the earlier evaluation, so this week's result is encouraging rather than something that currently demands a model change.

Sorare

Clean sample: 205 players who played 60+ minutes.

MetricGW4 result
Average projected50.42
Average actual51.01
Bias+0.59
MAE15.13
Median absolute error13.10
Pearson correlation0.27
Spearman correlation0.28
Within ±1554.6%
Within ±2070.2%

Once again, the overall Sorare scoring environment landed extremely close. Average projection: 50.42. Average actual: 51.01. A difference of just 0.59 points across 205 players.

The higher Sorare ranks separated again

Among the original rankings, the 10 clean survivors from the original top 20 averaged 67.71. Four scored 70+ and two scored 80+.

The 16 clean survivors from the original top 30 averaged 66.99, with seven hitting 70+.

Across the whole clean sample, the actual average was only 51.01.

So the model was again successfully concentrating stronger Sorare scores towards the upper part of the rankings.

Some excellent Sorare calls

Those are exactly the kinds of players we want the upper rankings to contain.

Sorare's upper tail was hotter than the model again

This is probably the clearest modelling watch item from GW4.

ThresholdExpectedActual
50+96.9096
60+53.3368
70+26.1438
80+11.8015

The 50+ calibration was almost perfect. But once we move into the upper tail, reality was noticeably hotter than forecast.

There is another useful check here. A P90 score is, by definition, intended to be exceeded roughly 10% of the time. In GW4, 34 of 205 players exceeded their P90. That's 16.6%.

So there is now further evidence that the Sorare simulation may be a little conservative in the upper tail. That does not mean we immediately tweak the model after one gameweek. Football results themselves can run hot. But it is something worth continuing to track.

Sorare misses

Again, the top of the rankings was not universally successful.

And some of the biggest scores of the week came from players the model did not rank especially highly.

Those are not predictions we are going to retrofit into successes after the event.

What do we take from GW4?

The most encouraging result is probably not any individual player. It is the broader shape of the forecasts.

Across 197 FPL players, average predicted points were only 0.27 away from average actual points. Across 205 Sorare players, average projected score was only 0.59 away.

The FPL model expected 12.52 10+ hauls and got 13. The highest-ranked FPL players materially outscored the wider pool. And there were some genuinely useful player-level calls.

Okafor at number two was excellent.

Wissa at number 407 — and on our Traps of the Week list — was arguably just as interesting.

That Okafor/Wissa contrast is also a good illustration of what we want Fantasy Matchup Edge to do differently. A normal FPL fixture difficulty view starts with the team.

FME starts with the player: how he scores, what actions his opponent tends to allow, where those actions occur, the likely match environment and what that combination means for his fantasy scoring distribution.

Sometimes the two approaches will give you the same answer. Sometimes they won't.

GW4 gave us a pretty good example of why that difference can matter.

There were also clear misses, and the Sorare upper tail remains something we want to monitor. Four gameweeks into a season, none of this proves a model is “solved”.

But the process is beginning to give us exactly what we wanted from it: a locked set of pre-match projections, measurable probabilities, transparent results afterwards, and matchup-specific calls we can actually test.

And then we do it again next week.

Want to see how FME turns player action profiles, opponent tendencies and match environment into projections? Read how our model works →