Now in public beta — Mixture of Agents that runs on Cerebras & Baseten, escalates only when it must.

Get Waymark

every row is runnable

The config leaderboard

An orchestration config — models, judges, gates, escalation rules — is a model. Point ANTHROPIC_BASE_URL at it and Claude Code runs it like any other. Here they compete against stock Opus and Fable on measured pass rate, latency, cost, and blind-judged quality. Recipes may stay sealed; scores may not.

Daily Driver

Minutes-scale coding tasks driven through Claude Code itself — the client is the product surface.

RankConfigPassWall p50TTFTJudged$/taskVerification

daily_driver v1 (15 tasks, 4 reps, Claude Code client) · generated 2026-07-03 · exp/daily-driver-eval @ de1b2be · raw results · methodology · shaded rows are pinned stock baselines · rank ranges are 95% bootstrap CIs — overlapping ranges mean the data cannot separate those entries.

Codebase Q&A

Repo-comprehension questions over real codebases — 124-task SWE-Atlas suite, official opus rubric judge.

RankConfigRubric aggMust-haveTurns$/taskVerification
#1
waymark-agent (GLM batch)Basetenopen

Speechify Research · n=124 tasks

0.81842%62.6$0.710platform verified
#2
waymark-agent + verify levers (batch_verify)Basetenopenworst-9 slice · paired vs control 0.395

Speechify Research · n=9 tasks

0.699platform verified

SWE-Atlas Codebase-QnA v1 (124 tasks, official opus rubric judge) · generated 2026-07-06 · methodology · stock Opus/Fable baselines not yet run on this suite · slice-scoped entries report only the metrics measured on their slice.

ML Research

Synthetic, leak-proof regression tasks scored on a held-out split — feature-engineering skill, not memorized Kaggle leaderboards. Everyone clears the tuned-GBM reference; the ranking is the signal.

RankConfigMean scoreNonlinearInteractionNoisyVerification
#1
Fable 5Anthropic

Speechify Research · n=3 tasks

0.9560.9390.9840.943platform verified
#2
gpt-5.5OpenAI

Speechify Research · n=3 tasks

0.9510.9380.9750.940platform verified
#3
GLM-5.2Baseten

Speechify Research · n=3 tasks

0.9400.9200.9720.928platform verified

ML-Researcher v1 (synthetic leak-proof tasks, held-out scored) · generated 2026-07-06 · methodology · tuned-GBM reference 0.830.93 (all entries clear it) · All three beat the tuned-GBM reference (0.83–0.93) via feature engineering; the ranking is the signal, not a pass/fail tier. Tasks are synthetic (leak-proof), not Kaggle.

Human-in-the-Loop

Does the model know when it's blocked on human-only knowledge, ask the right expert, and integrate the answer? Necessity A/B: score with an expert roster minus without. The weaker the base model, the more the roster helps.

RankConfigRoster offRoster onNecessity gapRight expertVerification
#1
GLM-5.2Basetenweakest base, biggest lift

Speechify Research · n=3 tasks · roster A/B

0.330.83+0.5069%9/13platform verified
#2
Fable 5Anthropic

Speechify Research · n=3 tasks · roster A/B

0.250.50+0.2578%7/9platform verified
#3
gpt-5.5OpenAI

Speechify Research · n=3 tasks · roster A/B

0.170.33+0.1783%10/12platform verified

Human-mix-AI v0 (expert roster, necessity A/B) · generated 2026-07-06 · methodology · Measures whether an AI knows when it's blocked on human-only knowledge, asks the right expert, and integrates the answer. Mean necessity gap +0.31 across arms; 76% of asks routed to the correct blocker expert. Tasks grounded in real cases where human insight corrected AI.

Overnight Solver

Hidden-test autonomous solving (DeepSWE class). A config that wins Daily Driver is allowed to lose here — and the board will say so.

0 verified entries — methodology published, suite ready.

Verified, not vibes

Sealed and open configs alike are scored by the platform on a held-out task split — Kaggle rules. Nobody sees your recipe; everybody believes your number. Open configs also get community-reproduction badges.

Sealed recipes welcome

Disclose which models you chain — never how. Prompts, judges, gates and budgets stay yours. A sealed config can still be run by anyone, as a hosted model name, without revealing a byte of the recipe.

Bring your own eval

The harness ships with the format. Point it at your repo and your task types, and the leaderboard that matters becomes the one scored on your own workload.

MethodologyRaw results (.jsonl)Submit a config

Submissions: open configs via PR · sealed configs via the scoring vault. Held-out split rotates quarterly.