Skip to content
Models
PLAN Free $0/mo PLAN Pro $20/mo PLAN Max $200/mo PLAN Teams $80/mo + $40/seat Claude Opus 4.8 agentic 56.1 · $5/$25 MTok · $0.9858/task GPT 5.4 agentic 53.8 · $2.5/$15 MTok · $0.3874/task GPT 5.5 agentic 52.1 · $5/$30 MTok · $0.4356/task GLM 5.2 agentic 51.9 · $1.4/$4.4 MTok · $0.2246/task Claude Sonnet 5 agentic 51.1 · $3/$15 MTok · $0.5134/task Claude Opus 4.7 agentic 50.7 · $5/$25 MTok · $0.5282/task Claude Fable 5 agentic 50.7 · $10/$50 MTok · $1.4777/task GPT 5.2 agentic 50.3 · $1.75/$14 MTok · $0.2336/task Claude Opus 4.8 $5 / $25 MTok Claude Fable 5 $10 / $50 MTok Devin overage At model API pricing Devin Cloud Updates Queued Messages and Adds Undo for Folder Moves Devin CLI v3000.6.2: Enhanced Path Resolution The partnership we missed: Fiserv is deploying Devin against core bank… SpaceX reportedly tried to buy Cognition. Scott Wu says it is not for … Devin CLI v3000.5.20: Enhanced Session Management Devin Cloud Updates Session Management Devin Desktop: Continuity — 2026-08-21 Devin Cloud Updates PLAN Free $0/mo PLAN Pro $20/mo PLAN Max $200/mo PLAN Teams $80/mo + $40/seat Claude Opus 4.8 agentic 56.1 · $5/$25 MTok · $0.9858/task GPT 5.4 agentic 53.8 · $2.5/$15 MTok · $0.3874/task GPT 5.5 agentic 52.1 · $5/$30 MTok · $0.4356/task GLM 5.2 agentic 51.9 · $1.4/$4.4 MTok · $0.2246/task Claude Sonnet 5 agentic 51.1 · $3/$15 MTok · $0.5134/task Claude Opus 4.7 agentic 50.7 · $5/$25 MTok · $0.5282/task Claude Fable 5 agentic 50.7 · $10/$50 MTok · $1.4777/task GPT 5.2 agentic 50.3 · $1.75/$14 MTok · $0.2336/task Claude Opus 4.8 $5 / $25 MTok Claude Fable 5 $10 / $50 MTok Devin overage At model API pricing Devin Cloud Updates Queued Messages and Adds Undo for Folder Moves Devin CLI v3000.6.2: Enhanced Path Resolution The partnership we missed: Fiserv is deploying Devin against core bank… SpaceX reportedly tried to buy Cognition. Scott Wu says it is not for … Devin CLI v3000.5.20: Enhanced Session Management Devin Cloud Updates Session Management Devin Desktop: Continuity — 2026-08-21 Devin Cloud Updates
LiveBench 2026-06-25

Agentic coding is still the hard part

Across 28 models in LiveBench's 2026-06-25 release, the best score on real repository tasks is 56.1 — while the best score on classic coding questions is 83.6. The average model gives up 30.9 points when it has to work in a repository instead of answering a prompt.

Cost does not track capability either. GLM 5.2 reaches 51.9 — 93% of the best score — at $0.225 per successful task. That is roughly 4× cheaper than Claude Opus 4.8, which pays $0.986 for 56.1. One of those two costs is taken from the nearest published effort variant, so treat the multiple as indicative — see the method note below.

Live · refreshed every 3 hours

What Cognition is actually shipping

Every changelog entry and post we ingest, by product and theme. Unlike the benchmark below, this updates continuously — it comes from our own feed of Cognition's published sources, not a third-party snapshot.

This month
23 -48%
Last 90 days
114
Tracked since
2024-03

Changelog entries per month, by product

Last 18 months · 544 entries tracked in total

Mar 25Apr 25May 25Jun 25Jul 25Aug 25Sep 25Oct 25Nov 25Dec 25Jan 26Feb 26Mar 26Apr 26May 26Jun 26Jul 26Aug 260102030405060changelog entriesDesktopPluginsCloudCLICognitionproduct

Days since last ship

As of 2026-08-27

05101520253035days since last changelog entryCloudCLIDesktopPluginsCognition

What they've been working on

Theme appearances across the last 180 days. Entries carry more than one theme, so these are appearances rather than shares.

020406080100120140160180theme appearances, last 180 daysbugfixdesktopcloud-agentuxclicode-reviewmcpsecurityintegrationsmodelsperformancerelease

Derived from Devin Central's own ingest of Cognition's published changelogs and blog (docs.devin.ai). Counts reflect published entries, which are a proxy for shipping activity — not a measure of engineering effort or code volume.

Snapshot · LiveBench 2026-06-25

How the underlying models score

The agentic gap, ranked

Points each model loses moving from classic coding questions to real repository work. Every model on the board degrades; the question is by how much.

0102030405060708090100110points lost on real repo workGPT 5.5Grok 4.3Claude Opus 4.5Qwen3.6 PlusClaude Sonnet 4.6GPT 5.2 CodexQwen3.6.27bClaude Fable 5Kimi K2.6DeepSeek V4 FlashClaude Opus 4.7Gemini 3.1 Pro PreviewQwen3.7GPT 5.4 MiniClaude Sonnet 5Claude Opus 4.6Gemini 3.5 FlashKimi K2.7 CodeGLM 5.2MiniMax M3DeepSeek V4 ProGPT 5.2GPT 5.4 NanoGPT 5.4Claude Opus 4.8Grok Build 0.1

Best agentic

56.1

real repo tasks

Best coding

83.6

classic benchmarks

Mean gap

30.9 pts

across 28 models

Models

28

28 with cost

Score against cost

The dashed line is the frontier — nothing is both cheaper and better. Cheap models sit surprisingly close to frontier ones. Hover any point.

204060$0.020$0.050$0.100$0.200$0.500$1.00cost per successful task (log)
Lab28 models
ModelReasoningCodingAgenticMathematicsData AnalysisLanguageIF$/task$/Mtok inout
Claude Opus 4.8frontierAnthropic · xhigh effort89.779.356.195.378.381.472.4$0.986*$5.00$25.00
GPT 5.4frontierOpenAI · xhigh88.177.553.894.179.382.670.2$0.387$2.50$15.00
GPT 5.5OpenAI · xhigh89.782.152.195.981.687.470.7$0.436$5.00$30.00
GLM 5.2frontierZhipu AI78.679.751.989.873.776.262.3$0.225$1.40$4.40
Claude Sonnet 5Anthropic · xhigh effort88.780.751.192.971.775.063.9$0.513$3.00$15.00
Claude Fable 5Anthropic · xhigh effort87.782.550.795.778.789.572.0$1.48*$10.00$50.00
Claude Opus 4.7Anthropic · xhigh effort87.282.150.792.978.377.966.7$0.528$5.00$25.00
GPT 5.2OpenAI · high83.276.150.393.278.279.861.8$0.234$1.75$14.00
GPT 5.2 CodexfrontierOpenAI77.783.649.488.878.273.766.4$0.187$1.75$14.00
Claude Opus 4.6Anthropic · thinking auto high effort88.778.249.089.369.983.363.3$0.404$5.00$25.00
Gemini 3.5 FlashGoogle · high82.078.249.088.264.984.675.6$0.249$1.50$9.00
Kimi K2.6frontierMoonshot AI · thinking79.478.646.984.365.175.164.4$0.169$0.950$4.00
GPT 5.4 NanofrontierOpenAI · xhigh81.170.846.891.067.662.567.2$0.091$0.200$1.25
GPT 5.5OpenAI · high89.780.046.895.280.487.871.4$5.00$30.00
Grok Build 0.1frontierxAI76.465.445.878.470.872.565.2$0.024$1.00$2.00
Kimi K2.7 CodeMoonshot AI82.874.045.779.662.777.956.3$0.100$0.950$4.00
Gemini 3.1 Pro PreviewGoogle · high84.076.545.491.078.585.479.1$0.285$2.00$12.00
Qwen3.7Alibaba · max83.374.243.685.271.879.774.0$0.182$2.50$7.50
Claude Sonnet 4.6Anthropic · thinking auto medium effort84.879.342.687.077.976.163.2$0.306$3.00$15.00
DeepSeek V4 ProDeepSeek82.770.042.690.774.578.162.4$0.050$0.435$0.870
GPT 5.4 MiniOpenAI · xhigh71.371.641.778.570.871.059.8$0.334$0.750$4.50
Qwen3.6 PlusAlibaba75.878.241.483.769.975.058.3$0.227$0.500$3.00
GPT 5.5OpenAI87.378.641.169.877.085.665.7$5.00$30.00
MiniMax M3MiniMax74.568.240.776.976.276.857.5$0.060$0.300$1.20
Claude Opus 4.5Anthropic · thinking 64k high effort80.179.739.790.474.481.362.5$0.610$5.00$25.00
Qwen3.6.27bAlibaba70.371.839.379.970.463.353.2$0.202$0.600$3.60
DeepSeek V4 FlashfrontierDeepSeek70.669.237.679.668.070.163.1$0.016$0.140$0.280
Grok 4.3xAI70.869.918.584.355.873.662.8$0.061$1.25$2.50

* 2 model(s) have cost taken from the nearest effort variant, because the two upstream tables label effort differently. Treat those figures as indicative.

Agentic coding leaderboard

Mean of the JavaScript, TypeScript and Python repository tasks, with cost per successful task on the right.

  • Claude Opus 4.856.1$0.986
  • GPT 5.453.8$0.387
  • GPT 5.552.1$0.436
  • GLM 5.251.9$0.225
  • Claude Sonnet 551.1$0.513
  • Claude Fable 550.7$1.48
  • Claude Opus 4.750.7$0.528
  • GPT 5.250.3$0.234
  • GPT 5.2 Codex49.4$0.187
  • Claude Opus 4.649.0$0.404
  • Gemini 3.5 Flash49.0$0.249
  • Kimi K2.646.9$0.169
  • GPT 5.4 Nano46.8$0.091
  • GPT 5.546.8
  • Grok Build 0.145.8$0.024
  • Kimi K2.7 Code45.7$0.100
  • Gemini 3.1 Pro Preview45.4$0.285
  • Qwen3.743.6$0.182
  • Claude Sonnet 4.642.6$0.306
  • DeepSeek V4 Pro42.6$0.050
  • GPT 5.4 Mini41.7$0.334
  • Qwen3.6 Plus41.4$0.227
  • GPT 5.541.1
  • MiniMax M340.7$0.060
  • Claude Opus 4.539.7$0.610
  • Qwen3.6.27b39.3$0.202
  • DeepSeek V4 Flash37.6$0.016
  • Grok 4.318.5$0.061

Vendor-published · FrontierCode 1.1 Extended · updated 2026-08-07

What Cognition's own benchmark says about the newer models

Several models missing from the LiveBench snapshot above — including Opus 5, GPT-5.6 Sol, Kimi K3, and Grok 4.5 — do appear on Cognition's FrontierCode 1.1 Extended chart, alongside Devin Fusion itself. These are the vendor's figures on the vendor's benchmark: useful precisely because they are current, and to be read with exactly that skepticism. They are quoted here, never mixed into the independent figures above.

Model Lab FrontierCode score Avg cost / task
Fable 5 (xhigh) Anthropic 64.9 $10.53
Opus 5 (medium) Anthropic 63.6 $3.51
Devin Fusion Cognition 63.1 $1.35
GPT-5.6 Sol (high) OpenAI 58.7 $3.41
Kimi K3 Moonshot 58.2 $3.12
Grok 4.5 (high) xAI 56.6 $1.09

Source: Devin Fusion post (chart data updated 2026-08-07); method in the FrontierCode 1.1 post. Cognition benchmarks and sells Devin Fusion — this table is the vendor grading itself. Fable 5 was measured before its access suspension (June 12, US directive) and is not currently available. Scores here are not comparable to the LiveBench numbers above; they are different benchmarks with different tasks and graders.

Method and caveats

  • This is a monthly snapshot, not a live feed. LiveBench publishes new questions on roughly a monthly cadence and replaces about a sixth of them each time. The release shown is 2026-06-25; we last synced it on 2026-08-24.
  • Agentic Coding is the mean of LiveBench's JavaScript, TypeScript and Python tasks, which run an agent harness against issues from real GitHub repositories. Coding is the mean of its code generation and code completion tasks. Category composition comes from LiveBench, not from us.
  • Token prices come from the per-model configs in LiveBench's main repository, where the project declares them — 28 of 28 models. That is the same publisher as the scores, in an actively maintained Apache-2.0 repository, so the pricing column does not depend on any preview endpoint.
  • Cost per successful task is different: it is measured during evaluation rather than declared, and is only published in LiveBench's redesigned leaderboard repository. 26 of 28 models have it; the rest show “—” rather than being dropped. If that source moves, this page keeps scores and prices and loses only this column. 2 model(s) are marked with an asterisk because the score and cost tables label effort levels differently, so the nearest variant was used.
  • Scores are not directly comparable across releases. Questions change, so a model moving between releases may reflect a new question set rather than a new model.
  • Not scored in this release (20). LiveBench prices a model when it begins evaluating it, so its cost table runs ahead of its scores. These shipped models are priced upstream but absent from every figure here: Claude Opus 5, DeepSeek V4 Flash 0731, DeepSeek V4 Flash Vision Exp, DeepSeek V4 Pro 0813, Gemini 3.5 Flash Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, GPT 5.6 Luna, GPT 5.6 Sol, GPT 5.6 Terra, Grok 4.5, Grok 4.6, Inkling, Kimi K3, Muse Spark 1.1, Muse Spark 1.2, Ox Alpha, Qwen3.8, Qwen3.8.27b, Smaug Agentic.
  • LiveBench publishes irregularly. Across 11 releases the median gap is 36 days, but the mean is 73 and the longest is 179 — and publication has trailed the release label by up to 185 days. We sync the newest release that exists; when it ages past our thresholds this page says so above, and past six months we retire it.

Source and attribution

All figures are derived from LiveBench (source), released under the Apache-2.0 licence. We publish derived views — category rollups and the cost/score frontier — and link back rather than mirroring their leaderboard.

Colin White et al. “LiveBench: A Challenging, Contamination-Free LLM Benchmark.” ICLR 2025. arXiv:2406.19314