When Cognition shipped Devin Fusion at the end of June, the pitch was blunt: “Engineering teams are lighting money on fire.” Six weeks later, the preview has produced its first wave of field commentary — and the interesting part is not whether people like it, but that practitioners and platform vendors alike have started treating Fusion’s architecture as the reference design for how agents should spend money.
What Cognition actually published
Fusion runs two parallel agents: a frontier model in the lead seat, and a cheaper “sidekick” model, each a full agent with its own tools and its own cached context. The lead delegates, monitors, and keeps the judgment calls — the plan, the interpretation of ambiguity, the final review. Cognition’s design note is explicit: “the main agent should take minimal actions, and only read what is absolutely necessary.”
The published numbers, all on Cognition’s own FrontierCode 1.1 benchmark:
- 35% lower cost at frontier-level scores at publication — revised up to 60% on FrontierCode 1.1 Extended data updated August 7, with the largest gains in implementation-focused sessions.
- On the August 7 chart, Fusion scores 63.1 at $1.35 per task, against Fable 5 (xhigh) at 64.9 for $10.53 and Opus 5 (medium) at 63.6 for $3.51.
- With Fable 5 in the lead seat, Fusion measured 41% cheaper than a pure Fable 5 harness while holding its score — measured before Fable access was suspended under a US government directive on June 12, and not yet re-run.
- In internal use, 88% of Cognition employees’ merged PRs were driven entirely by the automated Fusion router.
These are vendor figures on a vendor benchmark. Cognition publishes per-task worked examples — including a failure case, where delegating a judgment-heavy React feature dropped the score from 54 to 27 — which is more than most routing claims come with.
The field reception: routing as a design discipline
The sharpest early commentary is coming from Japan’s practitioner blogs, which have taken apart the launch post rather than summarize it.
One Qiita analysis argues Fusion’s real contribution is conceptual: it reframes the question from “which model should I pick?” to “when in a task does the expensive model need to think?” — routing as a mid-session design discipline, not a procurement decision. A second hands-on post walks the preview on app.devin.ai and echoes the published 41% figure while flagging the practical caveats: Fusion is preview-stage, and the cost win depends on your task mix behaving like FrontierCode’s — mechanical work delegates cleanly; judgment-heavy work should not.
The pattern also went mainstream this month: AI Engineer’s State of Model Routing panel (August 6) put Cognition alongside NVIDIA and OpenRouter, with a chapter on how Fusion performs against frontier models and a thesis the whole panel shared — model capabilities are jagged, so the winning move is a frontier model that plans and delegates rather than one model that does everything.
What’s still open
Three things to watch, in plain terms:
- The Fable asterisk. Fusion’s best number (41%) was measured with a model nobody can currently run — Fable 5 access remains suspended. Cognition says it will extend Fable to Fusion users “once access is restored.”
- Preview economics vs. production economics. The 60% figure is benchmark telemetry. Whether your bill drops depends on how much of your workload is the mechanical middle of tasks — the part sidekicks eat — versus judgment calls.
- The failure mode is documented but real. Cognition’s own worked example shows delegation destroying quality when “the judgment is the deliverable.” The router decides; the router can be wrong.
Method: All performance and cost figures are Cognition's, published on cognition.com and measured on FrontierCode 1.1 (fetched 2026-08-12). The Qiita posts are practitioner commentary — analysis and preview walkthroughs, not independent benchmark replication.
References
- Devin Fusion: Frontier Performance at 35% Lower Cost — architecture, worked examples, 35–60% and 41% claims, 88% internal stat, Fable suspension note.
- Making Fable Cheaper Than Opus — earlier per-run cost data ($1.86 vs $2.04) behind the lead-model thesis.
- FrontierCode 1.1 — the eval all of these numbers run on.
- Devin Fusionはモデル選定ではなくルーティング設計である — routing-as-design analysis (Japanese).
- Devinの「Fusion」は賢いモデル1体をやめた — preview walkthrough and caveats (Japanese).
- The State of Model Routing — NVIDIA, Cognition, OpenRouter — AI Engineer panel, August 6.
Editorial from Devin Central — a fan news desk, not Cognition. Devin is a trademark of Cognition.