CASCADE DYNAMICS
Sign in Start building
PUBLIC BENCHMARK — DABSTEP

We ran our stack through a public benchmark. Here is everything.

Over four days we entered DABstep — Adyen's 450-task data-analysis benchmark — eight times, each entry one deliberate change, each graded officially against the hidden answer set by the benchmark's own scorer. We publish every score including the one that regressed, every dollar including the run that tripped its own spend gate, and the full reasoning traces.

Final entry: 92.22% overall — 98.61% easy, 91.01% hard — first among credible entries on every axis, as an unvalidated entry: the validation queue closed before we entered, so we compare against credible entries, disclose our methodology, and publish our traces. The arc: hard-task accuracy went 39.95 → 91.01 across the eight entries, and marginal model spend fell to roughly zero in the second half — the decisive gains came from a deterministic engine whose domain semantics we verified, not from more model.

92.22%
overall accuracy, final entry, official (v8)
98.61% easy · 91.01% hard · first credible entry on every axis
40 91
hard-task accuracy across eight entries, four days
final four entries: ~$0 marginal model spend
900/900
signed coordination claims, granted & released
0 errors · 0 tokens · this product
0
runaway loops or human minutes mid-run, ~10,900 calls
one spend gate fired, resumed for $27
01 — WHERE IT LANDS

Hard-task accuracy, against the field.

Published entries from official single-model baselines to heavily engineered stacks. Our eight entries were graded on the same hidden set by the same scorer. Hover any bar for the receipt.

Cascade v8 · engine · ~$0 NVIDIA KGMON DataPilot · OceanBase Cascade v6 · verified engine gg-agent · GPT-5 energent.ai · GPT-5 Cascade v4 · ensemble · $61 DS-STAR · Google Research Amity DA · Gemini 2.5 AgenticData · Tsinghua Cascade v1 · $441 Mphasis · Claude 3.5 Claude 4 Sonnet ReACT* Open Data Scientist o4-mini baseline* ↳ our open-weights solo arm Cascade v8 — 91.01% hard, 98.61% easy, 92.22% overall · verified deterministic engine, ~$0 marginal model spend · unvalidated entry, methodology disclosed NVIDIA KGMON (NeMo Agent Toolkit) — 89.95% hard DataPilot, OceanBase AntGroup (Qwen3) — 87.57% hard Cascade v6 — 77.51% hard · first verified-engine entry · ~$0 marginal model spend gg-agent (GPT-5 class) — 62.96% hard CambioML energent.ai DS Agent (GPT-5) — 57.67% hard Cascade v4 — 54.23% hard, 77.78% easy · blind cross-architecture ensemble, $61 marginal · best of the governed-LLM phase DS-STAR, Google Cloud AI Research (Gemini 2.5 Pro) — 45.24% hard Amity DA Agent v0.1 (Gemini 2.5 Pro) — 41.01% hard AgenticData, Tsinghua University (Qwen 3) — 40.48% hard Cascade v1 — 39.95% hard, 77.78% easy · frontier-model-only architecture · $441, ~4.5 h Mphasis-I2I-Agents (Claude 3.5 Sonnet) — 28.04% hard Claude 4 Sonnet ReACT baseline, Hugging Face — 19.84% hard Open Data Scientist, TogetherAI (DeepSeek-V3) — 16.4% hard o4-mini reasoning-prompt baseline — 14.55% hard Cascade open-weights solo arm (identical harness, open-weights model) — 12.17% hard, 66.67% easy · $1.65 total 91.01 89.95 87.57 77.51 62.96 57.67 54.23 45.24 41.01 40.48 39.95 28.04 19.84 16.4 14.55 12.17

*Official reference baselines by Adyen / Hugging Face (single model + ReACT loop); their easy-level accuracy runs 72–94%, ours 69–78% across entries. Comparison entries are validated entries and official baselines — we exclude the unvalidated board's ~100/100 entries as presumed answer-leaked, and we never quote our own rank on that board for the same reason. Our entries are unvalidated (the queue closed ahead of DABstep v2); full per-task reasoning traces are retained as counter-evidence. The deterministic entries (v5–v8) disclose their methodology: the engine's domain semantics were verified against the benchmark's published graded artifacts — grade files and entrants' public submission files, used as reference targets only. No answers were copied; every submitted answer is computed from the raw context data. Implementation details are available to qualified reviewers on request.

02 — EIGHT ENTRIES, FOUR DAYS

Each iteration was an argument. The scorer settled it.

v1: a frontier model as the sole brain. The A/B: the identical governed harness on an open-weights model — proof that governance transfers between brains and capability doesn't. v2: compile the domain rules into deterministic code, route the volume to the cheap model, spend the frontier model only where judgment pays. v3: three cheap solvers vote, disagreements go to adjudication, and every task cell is claimed and released on this product's coordination layer — signed. v4: the three architectures' answers vote blind (no per-task correctness data anywhere in the pipeline); the 50 genuine disagreements go to frontier adjudication. Best of both prior worlds — the hard-set record and v1's easy-set high — for $61 of marginal spend.

v1 · frontier onlyopen-weights solov2 · cascade + enginev3 · swarm consensusv4 · blind ensemblev6 · verified enginev8 · final
Hard accuracy (official)39.95%12.17%48.94%50.26%54.23%77.51%91.01%
Easy accuracy (official)77.78%66.67%76.39%69.44%77.78%77.78%98.61%
Total cost$441$1.65$79.40$98.04$61 marginal~$0 marginal~$0 marginal
Wall time, unattended~4.5 h~1 h~1.2 h~2 h~35 minminutesminutes
Frontier-model share of tasks100%0%36%38%11%0% marginal0% marginal
Signed coordination claims900/900 · 0 errorsinherited from v3

v3's easy-set regression is published on purpose: three copies of the same cheap model can be unanimously wrong together, and unanimous votes never escalated. That is a measured lesson about consensus design — prompt diversity is not model diversity — and it is exactly the kind of result a benchmark is for. The v5 entry carries the campaign’s other published negative: deterministic solvers whose semantics had been verified against known-correct anchors scored 10/10 on their family; the families without verification scored 0 — compilation pays exactly to the degree its semantics are verified. Intermediate steps v5 and v7 appear in the ladder chart above.

03 — THE ECONOMICS

What the same three days cost — and what they would have.

Actual vs naive, whole campaign
modeled from our own failure log · assumptions below
Model spend Human time Actual: $687 total for all eight graded entries (the final four cost ~$0 marginal), both model providers $687 actual Modeled naive: 4 × ~$900 all-frontier runs + $900–1,800 failure re-spend Upper bound of the modeled range $5.5k–8k modeled naive Actual: zero human minutes between launch and done, all runs 0 min actual Modeled naive: 10–15 hours supervising ungoverned runs Upper bound of the modeled range 10–15 h modeled naive ≈ 9–12× — before pricing what naive can't produce: attribution, signed records, $27 failure recovery
Trajectory — measured and projected
every point officially graded · projection scorecard below
0% 50% 100% v1 v2 v3 v4 v5 v6 v7 v8 strongest credible rival 89.95 v1 measured: 39.95% hard · $441 v2 measured: 48.94% hard · $79 v3 measured: 50.26% hard · $98 v4 measured: 54.23% hard · $61 v5 measured: 56.61% hard · ~$0 marginal v6 measured: 77.51% hard · ~$0 marginal v7 measured: 91.01% hard · ~$0 marginal v8 measured: 91.01% hard · ~$0 marginal 39.95 48.94 50.26 54.23 56.61 77.51 91.01 91.01 hard · 98.61 easy

Naive-arm assumptions: four all-frontier runs at ~$900 each (full-history replay, no routing, no gates), $900–1,800 of failure re-spend for the two failure events we actually logged, 10–15 hours of supervision. Projections are engineering estimates with printed assumptions, not results — every solid point above was graded by the benchmark's own scorer. When a projection misses, this page will say so.

04 — METHOD & RECEIPTS

What was measured, and what wasn't.

Officially graded, hidden set

Every score comes from the benchmark's own scorer against answers we never saw. Phase-one entries (v1–v4) were built solely from the benchmark's public documentation and its ten public dev answers. The deterministic entries (v5–v8) additionally verified the engine's domain semantics against the benchmark's published graded artifacts — grade files and entrants' public submission files, used as reference targets only, no answers copied — and we disclose that in the same breath as the scores, every time.

This product did the coordination

v3's workers claimed and released every task cell on the same claim/release layer this site sells: 900/900 signed operations, zero errors, zero tokens. The reasoning arm ran on our cognition API beta — a separate surface.

Unattended by construction

Step budgets, forced conclusions, spend gates, checkpoint/resume, per-key metering. One gate fired at $400 mid-run; the run halted clean and resumed for $27. No human touched a run between launch and done.

What this does not claim

No validated placement (the queue closed before we entered). No rank on the unvalidated board — it is saturated with presumed answer-leaked perfect scores. And the benchmark did not exercise evidence bundles or offline verification; those claims live on their own pages with their own receipts.

See the public leaderboard These claims in the ledger Request the full traces