Somebody has already pasted the line into your team chat: within one point of Fable 5.1, 64% cheaper. Cognition shipped SWE-2 on 10 September 2026, a coding model post-trained from the 2.8 trillion parameter Kimi K3, and on FrontierCode 1.1 Main it scores 50.0% against 50.9% for Claude Fable 5.1. That's the good row. The other row is Terminal-Bench 4, where SWE-2 lands at 27.3% and Fable 5.1 at 55.8%. Same table, same launch post. Only one of them made the headline.
The short answer
SWE-2 is live in Devin Desktop and the Devin CLI as of 10 September 2026, with Devin Web and Fusion following. There's no public API and no per token price, and the weights stay closed, so the "64% cheaper than Fable 5.1" claim is Cognition's own cost accounting on its own benchmark, and you can't reproduce it. It tops the table on Terminal-Bench 2.1 at 92.8%, sits a point under Fable 5.1 on FrontierCode 1.1 Main, and collapses to 27.3% on Terminal-Bench 4, where Fable 5.1 posts 55.8% and GPT-6 Astra 57.9%. The post doesn't explain that gap.
What the table says, both halves of it
Start with the number Cognition wants you to read. FrontierCode 1.1 Main is Cognition's own benchmark, published in July 2026, and on it SWE-2 posts 50.0% at its best effort level. Fable 5.1 posts 50.9%, GPT-6 Astra 53.3%, GPT-5.6 Sol 47.5%. The Kimi K3 base sits at 44.2%, so the reinforcement learning bought 5.8 points, and the previous SWE-1.7 is down at 42.0%. DeepSWE 1.1 reads the same way: 73.0% for SWE-2, 74.1% for Astra. On Terminal-Bench 2.1 SWE-2 actually tops the table at 92.8%, ahead of Fable 5.1 at 91.4%. Three rows, one of them a win.
Then there's the fourth row. Terminal-Bench 4 is the newer, harder terminal benchmark, and it's the one OpenAI leaned on the week before to sell GPT-6 Astra at $10 and $50. On it SWE-2 scores 27.3%. Astra scores 57.9%, Fable 5.1 scores 55.8%, and even GPT-5.6 Sol, which Cognition beats on the headline benchmark, scores 37.3%. The Kimi K3 base is at 21.5%, so the RL lifted it by about six points here too, and SWE-2 still lands under half of the two frontier models. The post prints the row and says nothing about it. Not a footnote, not a caveat.
Honestly, a gap that size isn't noise, and I'd want it explained before I believed the rest. The charitable reading is that Terminal-Bench 4 rewards long, messy shell sessions that Cognition's RL environments don't cover yet. The less charitable one is that a model tuned hard against one lab's benchmark family looks exactly like this: great where it trained, ordinary where it didn't. I might be wrong about which applies. Cognition could settle it in a paragraph, and chose not to.
64% cheaper than what, and paid how
Now the money. The exact sentence is that SWE-2 reaches 50.0% on FrontierCode 1.1 Main "within one point of Fable 5.1 while being 64% cheaper", and that it "comes within a few points of GPT-6 Astra at a quarter of the cost". The footnote says costs assume list pricing for all models. Fine for Fable 5.1 and Astra, which have rate cards. SWE-2 doesn't. You can't buy it by the token. It lives inside Devin, which bills in its own compute units, and Cognition hasn't published a per token rate for it, any more than it did for SWE-1.7.
So the 64% is two public rate cards against one internal number nobody outside Cognition can see. That's not a lie, and I don't think it is one. It's just unfalsifiable from where you're sitting, which for a purchasing decision amounts to the same thing. The comparison you can check is internal: on the same benchmark, SWE-2 at medium effort scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less on average, and it makes its first real edit after a median of 18 steps against 48. Those we'd trust, because it's the same vendor measuring itself twice with the same ruler.
Fewer steps matters more than the percentage if you run agents at any volume. An agent that spends 48 turns reading before it touches a file is burning budget on exploration, and cutting that to 18 is the sort of change you notice in a day. Whether it stays at 18 on your repo is a different question. Cognition's chart breaks the steps down by phase, and exploration is the bar that shrank. That fits a model that has learned when to stop looking. It also fits one that has learned what FrontierCode tasks look like.
A Chinese open base, a closed product on top
SWE-2 is the second SWE model Cognition has built on a Moonshot base. SWE-1.7, in July, came from Kimi K2.7. This one comes from Kimi K3, the 2.8 trillion parameter model whose weights landed at 1.4 terabytes in late July. Cognition says this is the first time RL of this kind has been scaled to the multi trillion parameter regime, and the technical half of the post is genuinely about that: cost penalties tuned to the Pareto slope at each effort level, plus a draft model for speculative decoding that had to be retrained online because its acceptance rate kept sliding during RL. The three effort levels, medium, high and max, come out of one training run.
None of that reaches you as weights. Kimi K3 is open, SWE-2 isn't, and there's no licence discussion in the post at all. You get a closed product on an open base, which is a perfectly normal business model and worth naming plainly, because "built on an open model" tends to get shortened to "open" by the time it reaches a slide.
Cognition does something with the Chinese origin that we haven't seen other vendors do, and we'll give it credit for that. It reruns its propaganda and censorship test: 145 questions on politically sensitive topics in China, asked in English, Simplified Chinese and Traditional Chinese, each answer graded pass or fail by a single judge, GPT-5.6 Luna. SWE-2 passes 98.0% overall: 99.8% in English, 95.2% in Simplified Chinese, 99.1% in Traditional Chinese. It also reruns a test where identical coding tasks arrive with Western, Pakistani, Chinese, Tibetan or Falun Gong affiliated customer framings, and reports no statistically significant change in how vulnerable the code came out, for any of the six models tested. Good tests. One judge, from a rival lab, on 145 questions, is a thin basis for the word "trustworthy", though, and we'd rerun it before quoting it.
Our take, for what it's worth: if you already pay for Devin, SWE-2 is the model you'll get and it looks like a real step over SWE-1.7 on every row. If you don't, nothing here gives you a price to compare or a terminal benchmark you'd want to lead with. Wait for a number you can put in a spreadsheet.
Sources
Cognition, Introducing SWE-2: Pushing the Pareto Frontier, 10 September 2026 (the benchmark table, the 64% and quarter of the cost claims, the 18 versus 48 steps, the RL method, the propaganda and framing evaluations, the availability wording). Cognition, SWE-1.7: Frontier Intelligence at a Fraction of the Cost, July 2026 (the Kimi K2.7 base of the previous model). OfficeChai, Cognition Releases SWE-2, 10 September 2026 (independent confirmation of the release, the base model and the rollout surfaces). BenchLM, SWE-2 model page (confirms no published pricing and a 1M context listing, with partial benchmark coverage).
Frequently asked questions
What is SWE-2?
Cognition's coding model for Devin, released on 10 September 2026. It's post-trained with reinforcement learning from Kimi K3, Moonshot's 2.8 trillion parameter open weight model, and ships three effort levels, medium, high and max, trained in a single run. Cognition says it's the first RL of this type scaled to a multi trillion parameter base.
Can I call SWE-2 through an API?
Not as of the launch. It's available in Devin Desktop and the Devin CLI, and Cognition says it's rolling out on Devin Web and Fusion. There's no standalone API or per token price, and nothing to download. If that changes we'll add it to this page.
Is SWE-2 really 64% cheaper than Claude Fable 5.1?
That's Cognition's figure, measured on its own FrontierCode 1.1 Main benchmark at each model's best effort level, using list prices for the frontier models and an internal cost for SWE-2 that it doesn't publish. You can't check it from outside. The comparison you can check is against SWE-1.7, where SWE-2 medium costs 81% less on average and takes 58% fewer turns.
Why does SWE-2 score so low on Terminal-Bench 4?
Cognition doesn't say. Its table shows 27.3% for SWE-2 against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra, while on the older Terminal-Bench 2.1 SWE-2 leads at 92.8%. Our guess is that Terminal-Bench 4's longer, messier shell tasks fall outside what the RL environments cover, but that's a guess, and the vendor is the only party that can confirm it.
Does SWE-2 inherit Chinese censorship from Kimi K3?
Cognition tested for it. On 145 politically sensitive questions about China, asked in three languages and judged by GPT-5.6 Luna, SWE-2 passed 98.0% overall, with the lowest rate in Simplified Chinese at 95.2%. That's one judge on a small set, so treat it as a signal rather than a verdict.






















