DraftSOS publishes its full methodology and its complete out-of-sample accuracy record on this page, because a schedule model is only worth trusting if its errors are measured openly. The SOS Edge composite weights full-season opponent ease at 70 percent and playoff weeks 15-17 at 30 percent, and every input is a curated three-leg consensus of defensive ranks re-fetched on 2026-08-23 from Sportsnaut, FOX Sports, and a PFF anchor, with position splits from PFF, FTN Data, ESPN, and FantasyData. Validation uses leave-one-season-out cross-validation against nflverse actuals spanning the 2018 through 2025 seasons across quarterbacks, running backs, wide receivers, and tight ends, and the exact published formula, the per-position error tables, and the per-field refresh cadence documented below let any reader reproduce every number we publish before relying on it during a real fantasy draft this season.
How we model strength of schedule
DraftSOS turns a full season of matchups into floor, median, and ceiling outcomes — so your draft and lineup calls rest on the whole range of what each schedule can produce, position by position.
What we can prove
Loading validation figures…
The model
The free SOS grid is a deterministic composite — no randomness. It is built from two dated inputs: the published 2026 NFL schedule and curated defensive rankings (1–32, 1 = best) hand-blended from five preseason publications, refreshed pre-camp and at T-2 weeks. For each team-position cell the SOS Edge is 70% mean full-season schedule ease plus 30% mean ease across playoff weeks 15–17. The floor and ceiling ranks shown for a matchup are the minimum and maximum of the opponent's five curated defensive ranks (overall plus the position-specific splits) — the curated model's own profile spread for that defense, not sampled outcomes. Monte Carlo simulation is used one level up: the app board's championship simulator draws thousands of season paths from player-level distributions to estimate title odds. That is a separate, disclosed layer — not the free grid.
How we validate it
We measure the model out-of-sample. Using leave-one-season-out cross-validation, we hold an entire NFL season aside, predict it from the remaining seasons, and compare the prediction to what actually happened. We track three measures across positions: rank correlation (how closely the predicted order matches the actual order), mean absolute error (average distance between predicted and actual points per game), and directional accuracy (how often we call the above-below-median direction correctly). Holding a full season out keeps the test honest and free of hindsight fitting.
What “70.8% directional accuracy” actually means
For each held-out season, every player-season we evaluate gets an above/below-median call from the model and an above/below-median outcome. The call follows the rule in our validation code verbatim: a prediction is above if it is strictly greater than the median of that season’s predictions, and the outcome is above if actual points per game is strictly greater than the median of that season’s actuals
(predicted > median(predicted) vs actual > median(actual), per fold — quant-engine.run_leave_one_season_out_cv). A call is correct when the model and the season land on the same side of the median. Directional accuracy is the share of calls that are correct — across the 2018–2025 out-of-sample player-seasons, 70.8% (90% CI 68.8–72.7%, cluster-robust, n=2,400).
The baseline it has to beat: “always call the easier defense” — predict above-median whenever a player’s season schedule was easier than the league-median ease (measured ex-post, from actual points allowed by that season’s opponents), else below. That strategy hits 53.8% on the 1,486 player-seasons with a team-schedule mapping. On that identical subset, the model’s directional accuracy is 76.0% — the schedule is not the whole story, and the gap is the part of the skill the model adds on top of “just read the schedule.”
Honest caveats. The pooled figure is driven by the skill positions — within-position out-of-sample ρ: QB 0.64, RB 0.72, WR 0.71, TE 0.76, versus LB −0.10, DB −0.07, K 0.00 — so for IDP and kicker lines the model’s edge over a coin flip is, frankly, thin, and the pooled number overstates those. The baseline’s “ease” uses ex-post points allowed (it knows the defense it’s judging); the 2,400-unit sample caps each held-out season at 300 players. Season-by-season baseline performance itself swings 32–75%, so the lift is not uniform across eras.
Published results
Season-by-season out-of-sample results (2018–2025) are in the table above: rank correlation, MAE, and directional accuracy per held-out season. The same deterministic composite described above powers every ranking in the free SOS grid — you can put it to work today.
Notes to the analysis
What the numbers above do and do not cover, in eight items:
- Inputs & vintages (Data SLA)
- Curated defensive ranks (3-leg consensus re-fetched 2026-08-23: Sportsnaut + FOX + PFF anchor; refreshed at T-2 weeks) · position splits (PFF/FTN/ESPN/FantasyData, 2026-05-24 vintage) · validation history: nflverse 2018–2025 · ADP: PPG-based proxy, 300 players, refreshed daily. Per-field refresh schedule + owners: Data SLA.
- Consensus rule
- def_rank is a hand-blended rank average, rounded to 1–32 (1 = best). 2026-08-23 re-fetch: Sportsnaut defense rankings (updated 2026-05-03) + FOX power rankings (2026-07-28) equal-weight, PFF (2026-07-21) as #1 anchor (LAR); Athlon and Game Haus were unreachable at re-fetch (documented). Blending is manual and logged; artifact hashes ship in methodology v2.
- Clustered precision
- The headline directional accuracy carries a cluster-robust 90% CI of ±2.0 pp (bootstrap over the 2,400 player-seasons) — wider than the naive ±1.5 pp a plain binomial would suggest.
- Position specificity & cross-position ρ
- Weekly cells carry per-position defensive ranks (pass/run/WR/TE splits), not one generic number. Cross-position Spearman ρ across the 32 teams (2026-08-23 config): pass–run 0.91, pass–WR 0.98, pass–TE 0.94, run–WR 0.93, run–TE 0.94, WR–TE 0.90 — the splits are highly (near-)redundant; the overall def_rank carries essentially all the team-level signal, position splits refine it. F16-full (cardinal specificity) is a W6 item.
- Board objective & boosts (draft companion)
- Board score = (VOR + need floor) × (1 + gap boost + scarcity + bye + division) × depth × injury; VOR = projected PPG − replacement level (full-season ADP-anchored floor). When the live pool thins below required depth the floor holds at its pre-draft value and the payload flags
pool-thinned(no silent worst-remaining substitution). Active boosts are named in each recommendation's reason string. - Sleeper list
- Hand-curated (10 players) with written rationale; each edge is projected % vs an ADP-implied baseline (heuristic). No out-of-sample backtest of the sleeper rule yet — treat as scouted, not proved.
- Kicker is modeled
- sos-k.json is derived from the defense file by inverting points-allowed around the league mean with 0.6× shrinkage — badged “(modeled)” in the UI, not measured kicker data.
- Refresh SLA & limitations
- Projections refresh daily; schedule at release; defensive consensus at camp + T-2 weeks; ADP weekly. Limitations: the free Top 50 spans 17 distinct teams (each team's easiest cell only), the composite measures the easy side of a matchup (opponent defense), and projections do not ingest injuries.
See the model in action.
The same schedule model powers the free SOS grid. Try it now — instant, free access.
Our validation uses mean absolute error, rank correlation, and out-of-sample cross-validation, with methodology adapted from open-source fantasy analytics research. Open source credits
Data sources: nflverse (open NFL data, 2018–2025 regular seasons). Methodology last reviewed: 2026-08-23. Validation dataset: 8 NFL seasons, 2,400 out-of-sample player-season evaluations (300 per held-out season). Publisher: Vision Tech Solutions LLC. How this powers the SOS Edge →
Defensive ranks: Athlon, Game Haus, Sportsnaut, FOX Sports, PFF — 2026-05-24 consensus (as of file
sos-config.json · _last_updated), refreshed per the T-2-weeks SLA; blending is a manual rank average (consensus rule: average of the five outlet ranks, rounded to 1–32).Position splits: PFF, FTN Data, ESPN, FantasyData — 2026-05-24 (
sos-position-config.json).Validation: nflverse 2018–2025 weekly stats, leave-one-season-out CV; computed output committed with the scripts (monorepo,
da_directional.py / da_from_units.py).Per-artifact SHA-256 hashes: pending (methodology v2).
Cite this methodology
The DraftSOS SOS Edge methodology is published openly for citation and replication. Suggested formats:
APA: DraftSOS. (2026). How DraftSOS models strength of schedule. Vision Tech Solutions LLC. https://draftsos.com/accuracy
MLA: "How DraftSOS Models Strength of Schedule." DraftSOS, Vision Tech Solutions LLC, 29 Jul. 2026, draftsos.com/accuracy.