every row is runnable
The config leaderboard
An orchestration config — models, judges, gates, escalation rules — is a model. Point ANTHROPIC_BASE_URL at it and Claude Code runs it like any other. Here they compete against stock Opus and Fable on measured pass rate, latency, cost, and blind-judged quality. Recipes may stay sealed; scores may not.
Daily Driver
Minutes-scale coding tasks driven through Claude Code itself — the client is the product surface.
| Rank | Config | Pass | Wall p50 | TTFT | Judged | $/task | Verification |
|---|---|---|---|---|---|---|---|
#11–4 | waymark-maxopen≈ ties Opus + Fable Speechify Research · n=48 runs | 95.8% | 24.4s | 3.6s | 9.17 | $0.622 | platform verified |
#21–6 | Claude Code + Opus 4.8 Anthropic (stock) · n=48 runs · reference row | 93.8% | 25.2s | 2.2s | 9.08 | $0.143 | platform verified |
#31–3 | Claude Code + Fable 5 Anthropic (stock) · n=48 runs · reference row | 93.8% | 36.2s | 4.9s | 9.00 | $0.300 | platform verified |
#43–6 | waymark-fastopen≈ ties Opus + Fable Speechify Research · n=48 runs | 89.6% | 7.0s | 1.0s | 7.00 | $0.119 | platform verified |
#53–6 | waymark-moa-v2open≈ ties Opus + Fablenegative result: chair-swap experiment Speechify Research · n=48 runs | 87.5% | 9.0s | 1.0s | 6.38 | $0.125 | platform verified |
#63–6 | waymark-moaopen≈ ties Opus + Fable Speechify Research · n=48 runs | 85.4% | 8.9s | 1.3s | 8.92 | $0.083 | platform verified |
daily_driver v1 (15 tasks, 4 reps, Claude Code client) · generated 2026-07-03 · exp/daily-driver-eval @ de1b2be · raw results · methodology · shaded rows are pinned stock baselines · rank ranges are 95% bootstrap CIs — overlapping ranges mean the data cannot separate those entries.
Codebase Q&A
Repo-comprehension questions over real codebases — 124-task SWE-Atlas suite, official opus rubric judge.
| Rank | Config | Rubric agg | Must-have | Turns | $/task | Verification |
|---|---|---|---|---|---|---|
| #1 | waymark-agent (GLM batch)open Speechify Research · n=124 tasks | 0.818 | 42% | 62.6 | $0.710 | platform verified |
| #2 | waymark-agent + verify levers (batch_verify)openworst-9 slice · paired vs control 0.395 Speechify Research · n=9 tasks | 0.699 | — | — | — | platform verified |
SWE-Atlas Codebase-QnA v1 (124 tasks, official opus rubric judge) · generated 2026-07-06 · methodology · stock Opus/Fable baselines not yet run on this suite · slice-scoped entries report only the metrics measured on their slice.
ML Research
Synthetic, leak-proof regression tasks scored on a held-out split — feature-engineering skill, not memorized Kaggle leaderboards. Everyone clears the tuned-GBM reference; the ranking is the signal.
| Rank | Config | Mean score | Nonlinear | Interaction | Noisy | Verification |
|---|---|---|---|---|---|---|
| #1 | Fable 5 Speechify Research · n=3 tasks | 0.956 | 0.939 | 0.984 | 0.943 | platform verified |
| #2 | gpt-5.5 Speechify Research · n=3 tasks | 0.951 | 0.938 | 0.975 | 0.940 | platform verified |
| #3 | GLM-5.2 Speechify Research · n=3 tasks | 0.940 | 0.920 | 0.972 | 0.928 | platform verified |
ML-Researcher v1 (synthetic leak-proof tasks, held-out scored) · generated 2026-07-06 · methodology · tuned-GBM reference 0.83–0.93 (all entries clear it) · All three beat the tuned-GBM reference (0.83–0.93) via feature engineering; the ranking is the signal, not a pass/fail tier. Tasks are synthetic (leak-proof), not Kaggle.
Human-in-the-Loop
Does the model know when it's blocked on human-only knowledge, ask the right expert, and integrate the answer? Necessity A/B: score with an expert roster minus without. The weaker the base model, the more the roster helps.
| Rank | Config | Roster off | Roster on | Necessity gap | Right expert | Verification |
|---|---|---|---|---|---|---|
| #1 | GLM-5.2weakest base, biggest lift Speechify Research · n=3 tasks · roster A/B | 0.33 | 0.83 | +0.50 | 69%9/13 | platform verified |
| #2 | Fable 5 Speechify Research · n=3 tasks · roster A/B | 0.25 | 0.50 | +0.25 | 78%7/9 | platform verified |
| #3 | gpt-5.5 Speechify Research · n=3 tasks · roster A/B | 0.17 | 0.33 | +0.17 | 83%10/12 | platform verified |
Human-mix-AI v0 (expert roster, necessity A/B) · generated 2026-07-06 · methodology · Measures whether an AI knows when it's blocked on human-only knowledge, asks the right expert, and integrates the answer. Mean necessity gap +0.31 across arms; 76% of asks routed to the correct blocker expert. Tasks grounded in real cases where human insight corrected AI.
Overnight Solver
Hidden-test autonomous solving (DeepSWE class). A config that wins Daily Driver is allowed to lose here — and the board will say so.
0 verified entries — methodology published, suite ready.
Verified, not vibes
Sealed and open configs alike are scored by the platform on a held-out task split — Kaggle rules. Nobody sees your recipe; everybody believes your number. Open configs also get community-reproduction badges.
Sealed recipes welcome
Disclose which models you chain — never how. Prompts, judges, gates and budgets stay yours. A sealed config can still be run by anyone, as a hosted model name, without revealing a byte of the recipe.
Bring your own eval
The harness ships with the format. Point it at your repo and your task types, and the leaderboard that matters becomes the one scored on your own workload.
Submissions: open configs via PR · sealed configs via the scoring vault. Held-out split rotates quarterly.