RESEARCH · RUNNING BENCHMARK · UPDATED 2026-08-02

Running model comparison

When a notable model comes out, we run Claude Code with it on fixed autoresearch tasks and publish the trajectories here. Pick the resource you care about, the task and the models; the plot and the standings follow. Why we measure it this way is covered in the blog post.

I care about
Task
Line = mean of best-so-far, band = ±1 std, over 7 runs per model (3 for Inkling). Hover the plot for exact values, hover a model name anywhere for its card, and use the corner button to enlarge.

Prices are USD per 1M input / output tokens on each provider's first-party API as of 2026-08-02; Inkling is priced at its 256K-context tier, the one these runs used, and Sonnet 5 at its introductory rate ($2.00 / $10.00, standard $3.00 / $15.00 from 2026-09-01). Cost per run and the cost curves price every token from the run transcripts at these rates: cache reads at each provider's cached-input rate, cache writes at Anthropic's 1-hour-TTL write rate where applicable.

PROTOCOL

How the benchmark works

changelog
2026-08-02 · Sonnet 5 and Haiku 4.5 added as baselines
2026-08-01 · run count raised from 3 to 7 per model · cost curves and per-run token and cost columns added
2026-07-26 · Opus 5 added

Follow the standings

We ping the Discord when a new model is added to this page. Raw trajectories and every individual run: research.autolab.ai.