1. What DFS Benchmark is
DFS Benchmark grades DFS tools against what actually happened, under rules published before any data is collected. It covers projections, ownership projections, optimizers and AI assistants. It answers one question: which tools are actually the strongest?
Who runs it. DFS Benchmark is built by the founders of The DFS Kitchen. Because of that:
- The DFS Kitchen appears on DFS Benchmark as a house entry: graded under the same rules as every tool and labeled "House" everywhere it appears.
- Data DFS Benchmark collects from other companies is never shared with The DFS Kitchen.
- The disclosure below appears in the site footer, on the About page and beside every house result.
- DFS Benchmark never calls itself "independent" or "unbiased." It earns trust through the rules below.
Disclosure. DFS Benchmark is owned and operated by the founders of The DFS Kitchen, a daily fantasy sports content and tools company. The DFS Kitchen appears on DFS Benchmark as a labeled house entry, graded under the same rules as every other tool. Data DFS Benchmark collects from other companies is never shared with The DFS Kitchen.
2. Principles
Every tool is graded the same way, and the rules can't move once data exists.
- Same rules for every tool. Default settings, the same capture minute, the same grading.
- Rules before data. This rulebook is published before the first capture. Changes get a new version number and never apply to past slates.
- Locked at lock. Every captured file is timestamped and fingerprinted before games start.
- Ranges, not hype. Every number is shown with its margin of error. A tool only leads a row when the gap is bigger than the noise.
- Consent first. Companies are graded by invitation. A company that declines is listed as declined.
- No pay for placement. Participation is free, and money never touches results.
3. Who's tested
v1 grades four groups: invited DFS tools, a house entry, Claude as the AI Division, and two simple baselines every tool should beat.
Tools Division (graded by invitation, and only with consent)
- SaberSim, Stokastic, RotoGrinders and FantasyLabs.
- Any other DFS tool can apply under the same rules.
AI Division
- Claude gets the same fixed prompt at the capture minute, in the Claude app with web search on and the slate's DraftKings salary file attached.
- The model version is recorded every slate.
- The prompt is published here and doesn't change mid-season.
Baselines
- Recent-form baseline: a player's average DraftKings points over his last 10 games (NBA) or last 6 games (NFL), counting last season when needed.
- Salary baseline: his salary in thousands × the slate's average points per $1K. Used when a player has too few games for recent form.
House entry (owned by DFS Benchmark's founders)
- The DFS Kitchen is graded under exactly the same rules and labeled "House" everywhere it appears.
- Its files are captured and fingerprinted before any other tool's files are downloaded that slate, so it never sees a competitor's numbers first.
- Data from other companies never reaches The DFS Kitchen.
- Where the Kitchen doesn't offer something, its cell reads "—", like any other tool.
4. Scope for v1
v1 grades DraftKings classic main slates: NFL as internal dry runs, then every NBA main slate from opening night.
| Area | v1 plan |
|---|---|
| Site | DraftKings |
| Contest format | Classic, main slate |
| NFL | Sunday main slate, internal dry runs from Sep 27, 2026 (not published) |
| NBA | Every main slate from opening night, Oct 20, 2026 |
| Later | MLB (2027 season), Showdown, FanDuel |
Two reference contests are picked each slate:
- Reference GPP: the main slate's largest-field GPP. Used for actual ownership, field finish and dupes.
- Reference double-up: the main slate's largest double-up. Used for the cash line.
5. Before lock: capture rules
Every source is captured at the same minute before lock, and nothing that happens after that minute counts.
- Capture minute
- Fingerprints posted
- Games play
- Actual results in
- Grading
- Results published
Each slate runs this cycle once. Sections 5 and 6 set the rules for each step.
Capture minute
- NBA: 30 minutes before the first game on the main slate.
- NFL: 12:00 PM ET Sunday, after inactives for the 1 PM games are posted.
What's captured from each tool
- The full projections file: DraftKings points, plus ceiling and floor if offered.
- Ownership projections.
- The tool's #1 lineup on default settings.
From the AI Division: the lineup Claude returns to the fixed prompt, plus its full answer.
How it's captured
- By hand, through each tool's own export buttons, or sent directly by the company.
- No scraping, bots or automated collection.
- Each capture logs the tool's plan name and version.
Locking
- Every file gets a SHA-256 fingerprint at capture. Fingerprints are posted publicly before lock.
- Each company can check its own files against the posted fingerprints.
- Baseline and AI Division files are published in full after grading.
- Late swap isn't graded. The capture is final.
6. After games: grading rules
Every source is graded against official DraftKings results on the same player pool.
Actual results
- Points: DraftKings scoring, computed from official box scores.
- Ownership: actual ownership in the reference GPP, from its standings file.
- Cash line: the lowest score that cashed in the reference double-up.
Player pool
- Players are matched by DraftKings player ID.
- Graded players are those every projection source projected, with a consensus of at least 10 DK points (NBA) or 5 (NFL). Consensus is the median across sources.
- The floor keeps deep bench players from padding anyone's accuracy.
Scratches
- Ruled out before the capture minute: counts as 0 points for any source still projecting him. That's a real miss.
- Ruled out after the capture minute: removed for every source. Nobody could have known.
Same-engine lineups
- DFS Benchmark's own solver builds each source's highest-projected legal lineup from its projections. Only the salary cap and positions apply.
- One engine for every source means lineup differences come from the projections.
- The solver's code is public.
7. The v1 benchmarks
Eight headline rows rank the tools. Record rows show everything else with ranges but never rank.
The headline number is Beat-the-Baseline: how much less error a source makes than the recent-form baseline.
Beat-the-Baseline = 1 − source MAE ÷ baseline MAE
Headline rows
| # | Benchmark | What it measures | How it's scored | Better |
|---|---|---|---|---|
| 1 | Beat-the-Baseline | Projection skill over a naive forecast | 1 − source error ÷ baseline error, same player pool | Higher |
| 2 | Points Accuracy | Average points missed per player | Mean absolute error in DK points, by position group | Lower |
| 3 | Rank Accuracy | Whether players were ordered correctly | Rank correlation of projected vs actual, within each position | Higher |
| 4 | Value Finder | Whether it found the slate's best values | Of the 10 best actual values (DK points per $1K of salary), how many were in its top 20 projected values | Higher |
| 5 | Ownership Accuracy | Whether it knew the crowd | Mean absolute error in ownership %, reference GPP | Lower |
| 6 | Chalk Finder | Whether it spotted the most-owned players | Overlap of its top 5 projected ownership with the actual top 5 | Higher |
| 7 | % of Perfect | Same-engine lineup quality | Same-engine lineup points ÷ best possible lineup's points | Higher |
| 8 | Cash-Line Rate | Whether the lineup would have cashed | Share of slates the same-engine lineup beat the cash line | Higher |
Record rows (shown with ranges, never ranked)
- Own #1 lineup: each tool's default #1 lineup, with its finish percentile in the reference GPP, dupes (identical lineups in the field) and donuts (roster spots scoring 1 DK point or less).
- Optimal Overlap: how many of the perfect lineup's players a lineup had, such as 4 of 8 (NBA) or 4 of 9 (NFL). Tracked for both the same-engine lineup and the tool's own #1.
- Top-1% rate and ROI: reported, never headlined. No season has enough entries to prove a GPP ROI edge.
AI Division rows
- Valid lineup rate: lineups that are legal (salary cap, positions, players on the slate).
- Stale info rate: lineups that include a player already ruled out at the capture minute.
- AI lineups are also graded on % of Perfect and Cash-Line Rate using the lineup the model returned, plus the record rows.
8. How results are reported
A tool only leads a row when its lead is bigger than the margin of error.
- Ranges on everything. Each number carries a 95% range from resampling slates. A tool leads only when its range clears the next tool's; otherwise the row reads "statistically tied."
- Minimum sample. No row is ranked before 20 NBA slates or 8 NFL slates.
- Labels. "—" means the tool doesn't offer it. "Declined" means the company declined. "Pending" means awaiting consent.
- Price context. Each tool's cheapest plan that covers what was graded, plus its cost per slate. The accuracy-vs-cost chart plots Beat-the-Baseline against cost per slate.
- Cadence. A results log posts after each slate is graded. The leaderboard updates weekly.
- Corrections. Grading errors are fixed publicly in a changelog. Any company can request a review of a published result.
9. Participation and data use
Companies take part by consent, and their data is used only to grade them.
How a company takes part
- Say yes by email. Participation is free.
- Pick how capture works: send your files at the capture minute, or authorize DFS Benchmark to download them from its paid subscription.
What DFS Benchmark does with your data
- Uses it only to grade, and stores it privately.
- Never republishes, sells or shares it, including with The DFS Kitchen.
- Publishes graded numbers only, never your projections or ownership.
Your rights
- Check your files against the public fingerprints.
- Request a review of any published result.
- Withdraw at any time. Grading stops from the next slate; published results stay on the record, marked "withdrew."
What DFS Benchmark doesn't do
- No paid placement and no sponsored rankings. Any sponsorship is disclosed and kept away from results.
Badges and season awards
- A badge needs a lead whose range clears the next tool's, after the minimum sample in section 8. No clear lead, no badge.
- Each badge names the row, sport and season, such as "#1 Ownership Accuracy, NBA 2026–27," and links back to DFS Benchmark.
- Season awards per sport: Most Accurate Projections, Best Ownership, Best Value Finder, Best Value for Money (Accuracy vs Cost) and Most Improved (form guide).
- The house entry can earn badges and awards under the same rules; they carry its House label.
- Badges and awards can't be bought.
Changelog
- Oct 2, 2026: Rulebook v1 draft published.