Model evaluation
GW5 Projections vs Actuals: Close on the Big Picture
Gameweek 5 gave Fantasy Matchup Edge another useful reality check. The headline is mostly encouraging: overall FPL and Sorare scoring again landed reasonably close to the model, FPL 10+ haul probability was well calibrated, and Sorare's 80+ forecast was almost exact.
But this was not a perfect week. Ranking separation was much weaker than we want, and both platforms produced more P90 exceedances than a perfectly calibrated 90th percentile would imply. Those are exactly the kinds of things we want these reviews to expose.
As always, the clean evaluation respects the model contract. FME projections are IF STARTS. DNPs, substitute appearances and starters who played fewer than 60 minutes remain visible in our review data, but they are not counted as clean projection misses.
The GW5 headline:
- FPL average projected: 3.72. Average actual: 4.08.
- Sorare average projected: 50.80. Average actual: 52.28.
- FPL 10+ calibration was close: 6.5% predicted, 7.2% observed.
- Sorare 80+ was almost exact: 6.0% predicted, 6.1% observed.
- Ranking correlation was weak in GW5, and we are watching that closely.
- Both platforms ran a little hotter than their P90 distributions suggested.
- And the model remained cold on Yoane Wissa. He blanked again.
First: the overall scoring level was close again
The most important starting point for any set of fpl player projections is whether the scoring environment as a whole is sensible. If a model projects everyone too high or too low, individual rankings can look clever while the underlying scale is wrong.
| Metric | FPL | Sorare |
|---|---|---|
| Clean sample | 209 | 213 |
| Average projected | 3.72 | 50.80 |
| Average actual | 4.08 | 52.28 |
| Bias (actual − projected) | +0.35 | +1.48 |
| MAE | 2.45 | 14.18 |
| Spearman correlation | 0.104 | 0.095 |
| P90 exceedance | 14.8% | 16.9% |
FPL actual scoring came in 0.35 points per player above the model. Sorare came in 1.48 points higher. Neither is perfectly centred, but both are close enough that the broad scoring environment was captured reasonably well across more than 200 clean starters on each platform.
That is encouraging because FME is not simply taking historic fantasy averages and nudging them up or down. The model builds player-level action expectations from the matchup, role, opponent tendencies and match environment, then converts those into platform scoring distributions.
Put another way: our fpl predicted points are intended to be the output of the underlying football model, not a cosmetic ranking score. GW5 was another week where the population-level total stayed in the right neighbourhood.
The probability calibration had some very good signs
A projection model should not be judged by whether it names every single player who hauls. That is not how probability works. What we want is for a group of players carrying, say, a 10% chance of a haul to hit that outcome at roughly the right frequency over time.
FPL thresholds
| Threshold | Predicted rate | Observed rate |
|---|---|---|
| 10+ points | 6.5% | 7.2% |
| 15+ points | 0.7% | 1.4% |
| 20+ points | 0.1% | 0.0% |
The 10+ number is the one we care most about here because there is enough event volume for it to be useful in a single gameweek. A 6.5% model rate against 7.2% observed is a good result.
The 15+ line ran hotter than forecast, but we are talking about a very small number of actual events. Twenty-point hauls are rarer still. We record them, but we are not going to make model changes because one or two huge individual fantasy scores happened in a particular week.
Sorare thresholds
| Threshold | Predicted rate | Observed rate |
|---|---|---|
| 60+ | 26.6% | 34.3% |
| 70+ | 13.2% | 16.9% |
| 80+ | 6.0% | 6.1% |
Sorare is more mixed. The 80+ tail was almost absurdly close — 6.0% predicted against 6.1% observed — while 60+ and 70+ scores happened more often than forecast.
That fits something we have already been monitoring: the model can be very good at the centre and far upper tail while still being a little conservative across the broader high-score region.
The P90 watch item is still there
A P90 is intended to represent a score that should be exceeded roughly 10% of the time over a sufficiently large sample.
In GW5, 14.8% of clean FPL players exceeded their P90. For Sorare it was 16.9%.
One gameweek can absolutely run hot. So can five. We are not going to widen every distribution because of a handful of weeks. But if this persists across a much larger sample, it would suggest that the simulated outcome distributions are still a little too tight or conservative in the tails.
For now: log it, accumulate it, do not knee-jerk the model.
The main thing we did not like: ranking separation
This is the part of GW5 that deserves some realism.
FPL's clean-sample Spearman correlation was 0.104. Sorare's was 0.095. That is weak.
It does not mean the forecasts were useless — the overall scoring level and several probability calibrations were good — but a fantasy product ultimately needs to do more than estimate the population average. It needs to put the stronger opportunities towards the top of the board.
We have seen better ranking separation in earlier weeks, so we are not treating GW5 as a reason to rewrite the model. But ranking quality is now one of the metrics we most want to accumulate longitudinally. If weak separation becomes a pattern rather than a noisy week, that is something we investigate properly.
Hits and misses: what actually matters from GW5
The strongest hits this week were not a cherry-picked screenshot of one player who happened to score. They were structural:
- FPL average scoring stayed close: 3.72 projected versus 4.08 actual.
- FPL 10+ probability calibrated well: 6.5% predicted versus 7.2% observed.
- Sorare 80+ probability was almost exact: 6.0% versus 6.1%.
- The model's recurring Wissa fade produced another useful outcome.
And the misses/watch items were equally clear:
- Ranking separation was weak on both FPL and Sorare this week.
- Sorare 60+ and 70+ outcomes ran hotter than the probabilities suggested.
- P90 was exceeded too often on both platforms for a perfectly calibrated long-run distribution.
That is the point of locking the projections before kickoff and reviewing them afterwards. We do not need to retrofit a story around whichever players scored. We can look at the model as a forecasting system.
The meta matters too: Wissa, again
Now, we aren't trying to start a hate campaign about Yoann Wissa, but....
...this is becoming a bit of a recurring feature.
Wissa remains a popular FPL asset. FME has repeatedly been much colder on him than the market, and he blanked again in GW5.
In GW3, when he was still heavily owned, FME had Wissa at just 3.24 projected points with a 3.2% chance of 10+. He returned one point.
In GW4, ownership was reported at 18.1%. FME ranked him 407th, projected 3.06 points, gave him a P90 of 8 and only a 2.7% chance of 10+. He returned two points.
And in GW5, despite remaining highly owned, the model continued to fade him. He blanked again.
The important point is not that Wissa is a bad footballer. He isn't. It is not even that managers should permanently avoid him. They shouldn't.
The useful part is the disagreement between market popularity and matchup fit. Ownership tells us what the fantasy meta believes. FME is trying to answer a different question: does this specific match suit the actions and scoring routes this player normally relies on?
That is also why a simple fpl fixture difficulty colour can only tell part of the story. A fixture can look attractive for a team without being equally attractive for every player in that team. Player role, action profile and opponent tendencies can point somewhere very different from the crowd.
So no, this is not the official Fantasy Matchup Edge anti-Wissa campaign.
But we're not apologising for the fade either.
What are we changing?
Nothing because of GW5 alone.
That is probably the most important discipline in this whole process. A model can always be made to look better in hindsight if you change it after every strange result. That is not validation. That is chasing noise.
What we are doing instead:
- continuing to accumulate clean starter-only FPL and Sorare evaluation samples;
- tracking bias and MAE across gameweeks, not just one-week snapshots;
- tracking rank correlation and top-of-board separation as a core product-health metric;
- tracking 10+/15+/20+ FPL probability calibration;
- tracking Sorare 60+/70+/80+ calibration;
- tracking P90 exceedance rates to see whether the simulation tails remain too conservative.
The encouraging part is that GW5 did not give us a model that was wildly wrong about the scoring environment. Quite the opposite. The broad averages and several important probability forecasts were close.
The realistic part is that ranking was noisy and weak this week, and the upper-tail monitoring signal is still there.
Five gameweeks into the season, that is exactly what we want from these reviews: enough positive evidence to keep building, enough honesty to know what still needs proving, and a locked record that stops us from pretending every result was obvious after the event.
Overall: another encouraging week for calibration, another useful set of watch items — and another Wissa blank.
Want to see how FME turns player action profiles, opponent tendencies and match environment into projections? Read how our model works →