FantasyOMatic Defense Ratings vs Industry aFPA: A 24-Season Backtest
Five companies publish a matchup number for the same defense and no two mean the same thing. Here is what 62,134 out-of-sample predictions say about each of them, about ours, and about the one habit that costs more than any table.
Five different companies publish a "matchup" number for the same defense, and no two of them mean the same thing. One is a level, one is a delta, one is a percentage, and two of them disagree about which direction is good. We took all five, rebuilt them from the same play data, and ran them against 24 seasons — 62,134 predictions, each one made using only the weeks that came before it.
Every one of them carries real signal. Every one of them is bigger than the effect it measures. Ours included.
Five numbers, five meanings
These are the constructions in circulation, as published:
| Source | What it is | Window | Direction |
|---|---|---|---|
| FantasyPros, FantasyData | Fantasy Points Allowed | season to date | high = generous |
| FantasyPros, numberFire | Adjusted FPA | season to date | positive = favorable |
| 4for4 | Schedule-adjusted aFPA | rolling 10 weeks | low = tougher |
| Establish The Run | DvP, a percentage | last 8 games, recency-weighted | +10% = inflates scoring 10% |
| FTN | Fantasy Points Against, DVOA-adjusted | rolling | negative = better defense |
Establish The Run's is multiplicative, which is a real disagreement rather than a formatting choice. A defense that inflates scoring 8% gives a 20-point player four times what it gives a 5-point player. The additive tables charge both the same.
Every number is bigger than the effect it measures
There is a way to test this that does not depend on anyone's opinion. Take the number a source printed before the game. Take what actually happened. Ask how much of the printed swing showed up in the result.
If a table says "+3.0 against quarterbacks" and quarterbacks really do gain about three points, that number is worth face value. If they gain one point, it is worth a third of what it says.
| Share of the printed number that shows up | Overstated by | |
| Every published table | 0.17 – 0.39 | 2.6× – 6× |
| FantasyOMatic · QB | 0.39 | 2.5× |
| FantasyOMatic · RB | 0.45 | 2.2× |
| FantasyOMatic · WR | 0.28 | 3.6× |
| FantasyOMatic · TE | 0.15 | 6.7× |
At quarterback and running back our number is closer to honest than anything published — 2.5× and 2.2× against their 3.4× to 3.8× and 2.6× to 2.7×. At receiver it is a tie. At tight end we are the worst of the six. That is the real state of it, and tight end is the position we are actively rebuilding because of it.
The reason ours is smaller is structural. The published tables take each opposing player's points, subtract that player's own season average, and average the result. The trouble is that the player's average was itself set by the defenses he happened to face, so the yardstick is contaminated by the very thing being measured. Plant a known six-point defense in synthetic data and that method returns 6.48. It finds the effect, and it finds too much of it.
Our model solves for every player and every defense at once, in one system, under a penalty that stops a defense with four games behind it from claiming a six-point effect. On that same planted six-point defense it returns 2.71 — deliberately short, because in September it does not yet know enough to say six.
The mistake that costs more than any table
Here is the finding with the most money in it, and it has nothing to do with whose arithmetic is better.
The industry ships projections and matchup tables as separate products. A subscriber reads the projection, then looks up the matchup, then adjusts. That counts the opponent twice — because the projection already had it.
Accuracy against the player's own season average, weeks 6 onward:
| QB | RB | WR | TE | |
| Our projection, matchup already inside | +1.22% | −0.83% | +0.47% | −0.24% |
| The same projection, plus a published table on top | −8.67% | −6.01% | −3.88% | −6.42% |
| A raw average plus that table, no model at all | −3.28% | −1.13% | −1.43% | −2.38% |
The matchup is already in the This Week number. On any player page, the projection breakdown shows it as its own row, so you can see exactly how many points the opponent is worth before you decide anything. Do not add anyone else's on top.
Does adjusting for schedule even help?
This is the claim the whole category rests on, so we split it by era rather than averaging 24 seasons into one comfortable figure.
How much better an opponent-adjusted ranking predicts the rest of the season than raw points allowed:
| Position | 2002 – 2012 | 2013 – 2025 |
| WR | +38.6% | +54.4% |
| TE | noise | +16.4% |
| RB | +20.6% | +3.2% |
| QB | +5.0% | −1.7% |
There is a matching finding at tight end that we would rather state than bury: before 2013, no method here — ours, theirs, or raw points allowed — detects anything at tight end at all. The 24-season tight end result is a modern-era finding wearing a long label.
Where we disagree with the public table
When our ranking of a defense differs sharply from the raw public number, one of them is wrong. Over 24 seasons, at wide receiver, ours has been the right one about 57% of the time — with a confidence interval running from 50% to 64%.
That interval clears a coin flip, and it is deliberately printed with the number. It is a wide-receiver result. We tested the same thing at quarterback, running back and tight end and found nothing that separates from chance.
What we checked and did not find
A study is only worth the claims it retires.
- "FPA doesn't account for opponent strength" is only true of the raw tables. 4for4 and Establish The Run already adjust. It is a fair criticism of the number on a generic fantasy site, not of the whole category.
- Ranking defenses is close to a commodity. After putting every method on a common scale, we do not rank defenses better than the adjusted tables do at any position. What differs is how big the printed number is, and at tight end they are ahead of us.
- Adjusting the player for the defenses he has faced is worth about 0.2 points. We built it, measured it, and it changes essentially no lineup decisions. It stays in the model because it is correct, not because it is an edge.
- No matchup number wins you close calls on its own. On pairs of players projected within a point of each other, every construction — ours included — lands within about a point of the others.
What to do with this
- Pick one source and learn its sign convention. Mixing a 4for4 level with a FantasyPros delta is how a good matchup becomes a bad start.
- Do not stack a matchup table on a projection that already has one. It is the single most expensive habit measured here.
- Trust the receiver matchup most. It is the position where opponent adjustment clearly predicts the rest of the season, and it is where our disagreements with the public table have been right.
- Treat any September matchup number as provisional — ours and everyone's. Four games is not enough for a six-point claim, and a number that admits it is more useful than one that does not.
- Discount what you read elsewhere by about two-thirds. A published +6 has behaved like a +2.
The Matchup Matrix (All-Pro) ranks every remaining opponent for every player on your roster using the opponent effect described above, in fantasy points rather than a rank. The free matchups page and the projection breakdown on any player page show the same number without the season-long view.
Talk this over with AI Coach
The AI Coach (Hall of Famer) can pull the opponent effect for any player on your roster and tell you how much of your projection it accounts for — which is the only version of this question that decides a lineup.
AI Coach is available on the Hall of Famer season pass.
The fine print
Every number here comes from a walk-forward test: at each week, the model saw only the weeks before it. The industry constructions were rebuilt from the same play data rather than scraped, so the comparison is of methods, not of data quality, and each was given a slightly better version of its own method than the one it ships. Standard errors are clustered on the defense-week, because players facing the same defense in the same week are not independent observations. Half-PPR scoring. The study is reproducible and gets re-run after this season; anything that stops holding up comes out.
A matchup number you can defend at face value beats a bigger one you have to discount in your head.
Don't share this with anyone in your league.
Share with the people you want to see win.
Discussion
No comments yet. Be the first to share your take.
