Agentic coding is still the hard part
Across 28 models in LiveBench's 2026-06-25 release, the best score on real repository tasks is 56.1 — while the best score on classic coding questions is 83.6. The average model gives up 30.9 points when it has to work in a repository instead of answering a prompt.
Cost does not track capability either. GLM 5.2 reaches 51.9 — 93% of the best score — at $0.225 per successful task. That is roughly 4× cheaper than Claude Opus 4.8, which pays $0.986 for 56.1. One of those two costs is taken from the nearest published effort variant, so treat the multiple as indicative — see the method note below.
Live · refreshed every 3 hours
What Cognition is actually shipping
Every changelog entry and post we ingest, by product and theme. Unlike the benchmark below, this updates continuously — it comes from our own feed of Cognition's published sources, not a third-party snapshot.
- This month
- 23 -48%
- Last 90 days
- 114
- Tracked since
- 2024-03
Changelog entries per month, by product
Last 18 months · 544 entries tracked in total
Days since last ship
As of 2026-08-27
What they've been working on
Theme appearances across the last 180 days. Entries carry more than one theme, so these are appearances rather than shares.
Derived from Devin Central's own ingest of Cognition's published changelogs and blog (docs.devin.ai). Counts reflect published entries, which are a proxy for shipping activity — not a measure of engineering effort or code volume.
Snapshot · LiveBench 2026-06-25
How the underlying models score
The agentic gap, ranked
Points each model loses moving from classic coding questions to real repository work. Every model on the board degrades; the question is by how much.
Best agentic
56.1
real repo tasks
Best coding
83.6
classic benchmarks
Mean gap
30.9 pts
across 28 models
Models
28
28 with cost
Score against cost
The dashed line is the frontier — nothing is both cheaper and better. Cheap models sit surprisingly close to frontier ones. Hover any point.
| Model | Reasoning | Coding | Agentic ↓ | Mathematics | Data Analysis | Language | IF | $/task | $/Mtok in | out |
|---|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 4.8frontierAnthropic · xhigh effort | 89.7 | 79.3 | 56.1 | 95.3 | 78.3 | 81.4 | 72.4 | $0.986* | $5.00 | $25.00 |
| GPT 5.4frontierOpenAI · xhigh | 88.1 | 77.5 | 53.8 | 94.1 | 79.3 | 82.6 | 70.2 | $0.387 | $2.50 | $15.00 |
| GPT 5.5OpenAI · xhigh | 89.7 | 82.1 | 52.1 | 95.9 | 81.6 | 87.4 | 70.7 | $0.436 | $5.00 | $30.00 |
| GLM 5.2frontierZhipu AI | 78.6 | 79.7 | 51.9 | 89.8 | 73.7 | 76.2 | 62.3 | $0.225 | $1.40 | $4.40 |
| Claude Sonnet 5Anthropic · xhigh effort | 88.7 | 80.7 | 51.1 | 92.9 | 71.7 | 75.0 | 63.9 | $0.513 | $3.00 | $15.00 |
| Claude Fable 5Anthropic · xhigh effort | 87.7 | 82.5 | 50.7 | 95.7 | 78.7 | 89.5 | 72.0 | $1.48* | $10.00 | $50.00 |
| Claude Opus 4.7Anthropic · xhigh effort | 87.2 | 82.1 | 50.7 | 92.9 | 78.3 | 77.9 | 66.7 | $0.528 | $5.00 | $25.00 |
| GPT 5.2OpenAI · high | 83.2 | 76.1 | 50.3 | 93.2 | 78.2 | 79.8 | 61.8 | $0.234 | $1.75 | $14.00 |
| GPT 5.2 CodexfrontierOpenAI | 77.7 | 83.6 | 49.4 | 88.8 | 78.2 | 73.7 | 66.4 | $0.187 | $1.75 | $14.00 |
| Claude Opus 4.6Anthropic · thinking auto high effort | 88.7 | 78.2 | 49.0 | 89.3 | 69.9 | 83.3 | 63.3 | $0.404 | $5.00 | $25.00 |
| Gemini 3.5 FlashGoogle · high | 82.0 | 78.2 | 49.0 | 88.2 | 64.9 | 84.6 | 75.6 | $0.249 | $1.50 | $9.00 |
| Kimi K2.6frontierMoonshot AI · thinking | 79.4 | 78.6 | 46.9 | 84.3 | 65.1 | 75.1 | 64.4 | $0.169 | $0.950 | $4.00 |
| GPT 5.4 NanofrontierOpenAI · xhigh | 81.1 | 70.8 | 46.8 | 91.0 | 67.6 | 62.5 | 67.2 | $0.091 | $0.200 | $1.25 |
| GPT 5.5OpenAI · high | 89.7 | 80.0 | 46.8 | 95.2 | 80.4 | 87.8 | 71.4 | — | $5.00 | $30.00 |
| Grok Build 0.1frontierxAI | 76.4 | 65.4 | 45.8 | 78.4 | 70.8 | 72.5 | 65.2 | $0.024 | $1.00 | $2.00 |
| Kimi K2.7 CodeMoonshot AI | 82.8 | 74.0 | 45.7 | 79.6 | 62.7 | 77.9 | 56.3 | $0.100 | $0.950 | $4.00 |
| Gemini 3.1 Pro PreviewGoogle · high | 84.0 | 76.5 | 45.4 | 91.0 | 78.5 | 85.4 | 79.1 | $0.285 | $2.00 | $12.00 |
| Qwen3.7Alibaba · max | 83.3 | 74.2 | 43.6 | 85.2 | 71.8 | 79.7 | 74.0 | $0.182 | $2.50 | $7.50 |
| Claude Sonnet 4.6Anthropic · thinking auto medium effort | 84.8 | 79.3 | 42.6 | 87.0 | 77.9 | 76.1 | 63.2 | $0.306 | $3.00 | $15.00 |
| DeepSeek V4 ProDeepSeek | 82.7 | 70.0 | 42.6 | 90.7 | 74.5 | 78.1 | 62.4 | $0.050 | $0.435 | $0.870 |
| GPT 5.4 MiniOpenAI · xhigh | 71.3 | 71.6 | 41.7 | 78.5 | 70.8 | 71.0 | 59.8 | $0.334 | $0.750 | $4.50 |
| Qwen3.6 PlusAlibaba | 75.8 | 78.2 | 41.4 | 83.7 | 69.9 | 75.0 | 58.3 | $0.227 | $0.500 | $3.00 |
| GPT 5.5OpenAI | 87.3 | 78.6 | 41.1 | 69.8 | 77.0 | 85.6 | 65.7 | — | $5.00 | $30.00 |
| MiniMax M3MiniMax | 74.5 | 68.2 | 40.7 | 76.9 | 76.2 | 76.8 | 57.5 | $0.060 | $0.300 | $1.20 |
| Claude Opus 4.5Anthropic · thinking 64k high effort | 80.1 | 79.7 | 39.7 | 90.4 | 74.4 | 81.3 | 62.5 | $0.610 | $5.00 | $25.00 |
| Qwen3.6.27bAlibaba | 70.3 | 71.8 | 39.3 | 79.9 | 70.4 | 63.3 | 53.2 | $0.202 | $0.600 | $3.60 |
| DeepSeek V4 FlashfrontierDeepSeek | 70.6 | 69.2 | 37.6 | 79.6 | 68.0 | 70.1 | 63.1 | $0.016 | $0.140 | $0.280 |
| Grok 4.3xAI | 70.8 | 69.9 | 18.5 | 84.3 | 55.8 | 73.6 | 62.8 | $0.061 | $1.25 | $2.50 |
* 2 model(s) have cost taken from the nearest effort variant, because the two upstream tables label effort differently. Treat those figures as indicative.
Agentic coding leaderboard
Mean of the JavaScript, TypeScript and Python repository tasks, with cost per successful task on the right.
- Claude Opus 4.856.1$0.986
- GPT 5.453.8$0.387
- GPT 5.552.1$0.436
- GLM 5.251.9$0.225
- Claude Sonnet 551.1$0.513
- Claude Fable 550.7$1.48
- Claude Opus 4.750.7$0.528
- GPT 5.250.3$0.234
- GPT 5.2 Codex49.4$0.187
- Claude Opus 4.649.0$0.404
- Gemini 3.5 Flash49.0$0.249
- Kimi K2.646.9$0.169
- GPT 5.4 Nano46.8$0.091
- GPT 5.546.8—
- Grok Build 0.145.8$0.024
- Kimi K2.7 Code45.7$0.100
- Gemini 3.1 Pro Preview45.4$0.285
- Qwen3.743.6$0.182
- Claude Sonnet 4.642.6$0.306
- DeepSeek V4 Pro42.6$0.050
- GPT 5.4 Mini41.7$0.334
- Qwen3.6 Plus41.4$0.227
- GPT 5.541.1—
- MiniMax M340.7$0.060
- Claude Opus 4.539.7$0.610
- Qwen3.6.27b39.3$0.202
- DeepSeek V4 Flash37.6$0.016
- Grok 4.318.5$0.061
Vendor-published · FrontierCode 1.1 Extended · updated 2026-08-07
What Cognition's own benchmark says about the newer models
Several models missing from the LiveBench snapshot above — including Opus 5, GPT-5.6 Sol, Kimi K3, and Grok 4.5 — do appear on Cognition's FrontierCode 1.1 Extended chart, alongside Devin Fusion itself. These are the vendor's figures on the vendor's benchmark: useful precisely because they are current, and to be read with exactly that skepticism. They are quoted here, never mixed into the independent figures above.
| Model | Lab | FrontierCode score | Avg cost / task |
|---|---|---|---|
| Fable 5 (xhigh) | Anthropic | 64.9 | $10.53 |
| Opus 5 (medium) | Anthropic | 63.6 | $3.51 |
| Devin Fusion | Cognition | 63.1 | $1.35 |
| GPT-5.6 Sol (high) | OpenAI | 58.7 | $3.41 |
| Kimi K3 | Moonshot | 58.2 | $3.12 |
| Grok 4.5 (high) | xAI | 56.6 | $1.09 |
Source: Devin Fusion post (chart data updated 2026-08-07); method in the FrontierCode 1.1 post. Cognition benchmarks and sells Devin Fusion — this table is the vendor grading itself. Fable 5 was measured before its access suspension (June 12, US directive) and is not currently available. Scores here are not comparable to the LiveBench numbers above; they are different benchmarks with different tasks and graders.
Method and caveats
- This is a monthly snapshot, not a live feed. LiveBench publishes new questions on roughly a monthly cadence and replaces about a sixth of them each time. The release shown is 2026-06-25; we last synced it on 2026-08-24.
- Agentic Coding is the mean of LiveBench's JavaScript, TypeScript and Python tasks, which run an agent harness against issues from real GitHub repositories. Coding is the mean of its code generation and code completion tasks. Category composition comes from LiveBench, not from us.
- Token prices come from the per-model configs in LiveBench's main repository, where the project declares them — 28 of 28 models. That is the same publisher as the scores, in an actively maintained Apache-2.0 repository, so the pricing column does not depend on any preview endpoint.
- Cost per successful task is different: it is measured during evaluation rather than declared, and is only published in LiveBench's redesigned leaderboard repository. 26 of 28 models have it; the rest show “—” rather than being dropped. If that source moves, this page keeps scores and prices and loses only this column. 2 model(s) are marked with an asterisk because the score and cost tables label effort levels differently, so the nearest variant was used.
- Scores are not directly comparable across releases. Questions change, so a model moving between releases may reflect a new question set rather than a new model.
- Not scored in this release (20). LiveBench prices a model when it begins evaluating it, so its cost table runs ahead of its scores. These shipped models are priced upstream but absent from every figure here: Claude Opus 5, DeepSeek V4 Flash 0731, DeepSeek V4 Flash Vision Exp, DeepSeek V4 Pro 0813, Gemini 3.5 Flash Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, GPT 5.6 Luna, GPT 5.6 Sol, GPT 5.6 Terra, Grok 4.5, Grok 4.6, Inkling, Kimi K3, Muse Spark 1.1, Muse Spark 1.2, Ox Alpha, Qwen3.8, Qwen3.8.27b, Smaug Agentic.
- LiveBench publishes irregularly. Across 11 releases the median gap is 36 days, but the mean is 73 and the longest is 179 — and publication has trailed the release label by up to 185 days. We sync the newest release that exists; when it ages past our thresholds this page says so above, and past six months we retire it.
Source and attribution
All figures are derived from LiveBench (source), released under the Apache-2.0 licence. We publish derived views — category rollups and the cost/score frontier — and link back rather than mirroring their leaderboard.
Colin White et al. “LiveBench: A Challenging, Contamination-Free LLM Benchmark.” ICLR 2025. arXiv:2406.19314