Running model comparison
When a notable model comes out, we run Claude Code with it on fixed autoresearch tasks and publish the trajectories here. Pick the resource you care about, the task and the models; the plot and the standings follow. Why we measure it this way is covered in the blog post.
Prices are USD per 1M input / output tokens on each provider's first-party API as of 2026-08-02; Inkling is priced at its 256K-context tier, the one these runs used, and Sonnet 5 at its introductory rate ($2.00 / $10.00, standard $3.00 / $15.00 from 2026-09-01). Cost per run and the cost curves price every token from the run transcripts at these rates: cache reads at each provider's cached-input rate, cache writes at Anthropic's 1-hour-TTL write rate where applicable.
How the benchmark works
- Same harness. Every model runs inside Claude Code (versions 2.1.215–2.1.220 across cohorts) with identical prompts and configs, so the model is the only thing that changes between runs.
- 7 runs per model per task. A single run is mostly noise. We report mean ± std and keep every raw trajectory public. Inkling runs at 3.
- Fixed budgets. Runs get about an hour of wall-clock on KernelBench and several hours on Shakespeare; the standings also score everyone at fixed cutoffs so unequal run lengths can't tilt the comparison.
- Costs are in current API pricing. Token counts come from the run transcripts and are priced at each provider's public API rates, including cache reads and writes.
- New releases get added. When a model comes out, it goes into the queue and the standings update. The task set is growing too.
2026-08-02 · Sonnet 5 and Haiku 4.5 added as baselines
2026-08-01 · run count raised from 3 to 7 per model · cost curves and per-run token and cost columns added
2026-07-26 · Opus 5 added
Follow the standings
We ping the Discord when a new model is added to this page. Raw trajectories and every individual run: research.autolab.ai.