AI Tools
Cognition SWE-2: what it costs and where it beats Fable 5.1
Cognition SWE-2 scoring 50.0% on FrontierCode 1.1 Main at $1.18 per rollout makes it the cheapest model above 45% on Cognition's own leaderboard, checked September 11, 2026: Fable 5.1 scores 50.9% at $3.28 and GPT-6 Astra scores 53.3% at $4.59. SWE-2 ships inside Devin, not as a standalone API.
What matters
- Cognition's leaderboard puts SWE-2 at 50.0% on FrontierCode 1.1 Main for $1.18 per rollout; Fable 5.1 needs $3.28 for 50.9% and GPT-6 Astra $4.59 for 53.3%.
- SWE-2 is post-trained from Kimi K3, a 2.8T-parameter third-party base model, so Cognition's contribution is training and harness work rather than new weights.
- Cognition's own benchmark table shows SWE-2 at 27.3% on Terminal-Bench 4 against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra.
- There is no public SWE-2 token price: you buy it through Devin, from the $20/mo Pro plan up to a $80/mo Teams minimum with $40 full seats.
- Cognition reports SWE-2 medium finishing FrontierCode 1.1 Main in 53 steps against SWE-1.7's 127, at 81% lower cost.
What is Cognition SWE-2?
SWE-2 is Cognition's newest software-engineering model, announced on September 10, 2026, and post-trained from Kimi K3, a 2.8-trillion-parameter base model. Cognition calls it its most advanced coding model, and the launch post claims it lands within one point of Anthropic's Fable 5.1 on FrontierCode while costing 64% less per run.
Availability on day one was narrow: Devin Desktop and the Devin CLI, with Devin Web and Fusion rolling out behind them. The Devin CLI docs treat swe as an alias that resolves to the latest SWE model, so existing CLI users pick up SWE-2 by switching model inside a session rather than editing config.
The lineage is the part most coverage skips. SWE-2 is a post-trained checkpoint of Kimi K3, a model Moonshot had already put through heavy agentic RL. That makes the interesting comparison the same base trained two ways: unmodified Kimi K3 sits 9th on Cognition's board at 44.2%, while Cognition's version of it climbs to 50.0% and 5th. Cognition says its RL adds 5 to 6 points across several benchmarks on top of the base model.
SWE-2 benchmarks: what Cognition published
The launch post reports four benchmark columns. Two of them flatter SWE-2 and two do not, which is why the whole table is worth reading rather than the headline percentage.
- FrontierCode 1.1 Main: SWE-2 50.0%, Fable 5.1 50.9%, GPT-6 Astra 53.3%, Grok 4.6 48.0%, GPT-5.6 Sol 47.5%, Kimi K3 44.2%, SWE-1.7 42.0%.
- DeepSWE 1.1: SWE-2 73.0%, ahead of Fable 5.1 (67.4%) and just behind GPT-6 Astra (74.1%).
- Terminal-Bench 2.1: SWE-2 92.8%, the highest number in Cognition's table, ahead of Fable 5.1 (91.4%) and GPT-6 Astra (89.9%).
- Terminal-Bench 4: SWE-2 27.3%, against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra.
FrontierCode is built by Cognition, and the methodology page says each task is written by open-source maintainers working more than 40 hours per task, then graded on whether a real maintainer would merge the pull request. Runs that consult solution-bearing sources are scored zero. That is a stricter frame than pass-rate benchmarks, and it is also a benchmark where the vendor both writes the tasks and reports the scores.
The cost column is the real headline
FrontierCode's board publishes a cost per rollout column next to score, and that is where SWE-2 separates from the field. Cognition's leaderboard defines cost as mean US-dollar spend per rollout, and the model name carries the reasoning effort that produced its best score.
The top of the board, checked on September 11, 2026: Fable 5 (xhigh) 53.5% at $13.09, Opus 5 (medium) 53.4% at $4.31, GPT-6 Astra (max) 53.3% at $4.59, Fable 5.1 (medium) 50.9% at $3.28, then SWE-2 (max) 50.0% at $1.18. GPT-5.6 Sol (max) scores 47.5% for $5.19, Grok 4.6 (high) 48.0% for $2.88, and SWE-1.7 42.0% for $1.97.
Two cross-checks fall out of those numbers. SWE-2's $1.18 against Fable 5.1's $3.28 is a 64% gap, which matches the launch post's claim exactly. Against GPT-6 Astra's $4.59 it is roughly a quarter of the cost, which also matches. The cost is not a token price, though: it is what a rollout cost that vendor to run on that harness at that effort level, so it moves with the agent wrapper, not just the model.
Where SWE-2 falls short
The honest caveat sits in Cognition's own table. On Terminal-Bench 4, SWE-2 scores 27.3% while Fable 5.1 scores 55.8% and GPT-6 Astra 57.9%. GPT-5.6 Sol, a cheaper and lower-scoring model overall, still beats SWE-2 there at 37.3%. If your work is long-horizon terminal and tool use, the launch post gives you a reason to stay put.
Three more limits matter for a buying decision:
- The benchmark is the vendor's. Cognition writes FrontierCode, grades it, and now sells the model topping its cost-efficiency story. The September 10 changelog shows the company also corrects competitor prices and entries on that board, which cuts both ways: useful maintenance, but the same hand controls the inputs and the output.
- No independent replication yet. SWE-2 launched the same day the leaderboard entry appeared, so there is no third-party harness run to compare against.
- Flag rate is blank for SWE-2. Rivals carry a flag rate for unfair internet use (Fable 5.1 at 0.0%, DeepSeek V4 Pro 0813 at 10.6%). SWE-2's cell shows a dash, so that check has no published result.
What SWE-2 costs you: Devin plan prices
SWE-2 is not sold per token. It runs inside Devin, so the buyer question is a Devin plan question, and Cognition's own docs spell the tiers out:
- Free: limited usage, Devin Review and DeepWiki, one member.
- Pro: $20/month for one person, with a daily and weekly quota shared across Devin sessions, the CLI and Devin Desktop, plus pay-as-you-go on-demand credits past the quota.
- Max: $200/month, a larger weekly quota with no daily cap.
- Teams: an $80/month minimum, up to 200 members, $40/month for each full seat (Pro-equivalent quota plus Devin Desktop) and free flex seats that draw from the team's shared on-demand credits.
Credits roll over month to month and do not expire. Usage is metered on the work Devin does, not on wall-clock time: a sleeping session costs nothing, Devin sleeps after 30 minutes idle by default, and Windows sessions consume about 9% more than the Linux equivalent. There are no concurrent session limits, so the documented cost control is splitting work across sessions rather than queuing it.
Two accounting notes. Cognition points to devin.ai/pricing as authoritative, and that page is behind a bot checkpoint, so the tier numbers above come from the product docs, including the $80 Teams minimum and the $40 full seat. And users on the old ACU-based Core plan were migrated to Free, which means the cheapest path to SWE-2 in the CLI is the $20/mo Pro plan.
Who should switch to SWE-2, and who should not
Switch if your measurable cost is dollars per merged change and you already run Devin through the CLI or Desktop. The $1.18 rollout figure and the 58% fewer turns against SWE-1.7 are the two numbers that translate into a smaller bill, and DeepSWE 1.1 is the benchmark closest to repo-scale refactoring work.
Stay if your agent spends its life in a terminal on long, multi-step tool chains: Terminal-Bench 4 is a 27.3% against 55.8% gap in Fable 5.1's favour, and that is the weakest cell in Cognition's own table. Wait if you want a second opinion before moving. A model that launches on its author's own leaderboard has not yet been reproduced elsewhere.
One structural point for anyone standardising on Cognition: the Devin CLI is model-agnostic and lists Anthropic, OpenAI, Google and open-source families (DeepSeek, Kimi, GLM) alongside `swe`, with short names like `opus`, `sonnet` and `gpt`. SWE-2 is an option on that menu, not the menu itself, and Cognition's own router, Adaptive, will pick between them per task.
How to test SWE-2 on a small budget
Start on the Free plan to see the interface, then take Pro for a month if the CLI is what you need, since Devin CLI and Devin Desktop access belong to the paid tier. Inside a session, /model swe switches to the current SWE family member, and Alt+T cycles the thinking level, which is where the effort levels that produced the $1.18 score live.
Price the trial against your own failures rather than a leaderboard: give SWE-2 the tasks your current agent gets wrong, watch the step count and the credit burn, and compare the invoice at month end. On-demand credits make that comparison possible without buying a seat upgrade first, because they roll over instead of expiring.
At a glance
| Model (best effort) | Score | Cost per rollout |
|---|---|---|
| Cognition SWE-2 (max) | 50.0% | $1.18 |
| Anthropic Fable 5.1 (medium) | 50.9% | $3.28 |
| OpenAI GPT-6 Astra (max) | 53.3% | $4.59 |
| Anthropic Opus 5 (medium) | 53.4% | $4.31 |
| Anthropic Fable 5 (xhigh) | 53.5% | $13.09 |
| OpenAI GPT-5.6 Sol (max) | 47.5% | $5.19 |
| xAI Grok 4.6 (high) | 48.0% | $2.88 |
| Moonshot Kimi K3 | 44.2% | $3.82 |
| Cognition SWE-1.7 | 42.0% | $1.97 |
FAQ
Is Cognition SWE-2 free?
No. SWE-2 runs inside Devin, and Devin's Free plan lists limited usage plus Devin Review and DeepWiki. Devin CLI and Devin Desktop access, where Cognition shipped SWE-2 first, are listed under the $20/mo Pro plan.
How much does SWE-2 cost per million tokens?
Cognition has published no per-token rate for SWE-2, and there is no standalone SWE-2 API. The only public cost figure is per rollout on Cognition's FrontierCode leaderboard: $1.18 at max effort, against $3.28 for Fable 5.1 and $4.59 for GPT-6 Astra.
Is SWE-2 better than Fable 5.1?
On Cognition's board they land within one point (50.0% against 50.9% on FrontierCode 1.1 Main) while SWE-2 costs 64% less per rollout. On Terminal-Bench 4, Fable 5.1 more than doubles SWE-2 (55.8% against 27.3%). Pick by workload, not by the single headline score.
Related reading
GPT-6 Astra vs Claude Fable 5.1 API pricing, Codex vs Claude Code, Claude Code weekly limits
Sources
- Cognition: Introducing SWE-2, Pushing the Pareto Frontier (September 10, 2026)
- FrontierCode 1.1 Main leaderboard (Cognition, updated September 10, 2026)
- Devin docs: self-serve plans (Free, Pro, Max, Teams)
- Devin docs: how usage is metered
- Devin CLI docs: models and thinking levels
- Hacker News discussion: Cognition launches new SWE-2 model