• Latest
  • Trending
  • All
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
Answer card stating that the public download pages for the VMware Virtual Disk Development Kit on developer.broadcom.com began returning 404 errors on 25 August 2026 with no announcement or deprecation notice, that Broadcom support tells customers the kit is no longer available for use or download, and that release lines 7.0.3.1, 8.x and 9.x are all affected.

Broadcom pulled VDDK 8.0 and 9.0, and the 404 is the only notice

8 September 2026
Answer card stating that OpenAI published its research acceleration measurements on 6 September 2026, that as of mid August 2026 its research organisation logged 3.1 agent workdays of coding agent runtime for every workday of human labour normalised to a standard eight hour day, and that OpenAI states this should not be read as a 3.1 times productivity gain because it measures runtime rather than delivered output.

OpenAI’s 3.1 agent-workdays per human day is not a 3.1x gain

7 September 2026
Answer card stating that Mullvad announced on 3 September 2026 that it is shutting down its public encrypted domain name system servers on 2 November 2026 and sponsoring the Quad9 Foundation instead, with 194.242.2.2 and its five sibling addresses all going away, and virtual private network customers unaffected.

Mullvad’s DNS servers go dark on 2 November, and Quad9 blocks no ads

5 September 2026
OpenAI announcement image for GPT-6 Astra, a spiral galaxy of white, blue and amber points of light curling around a bright core on a near black star field.

GPT-6 Astra lists at $10 and $50, 2.5x what GPT-5.6 Sol costs

6 September 2026
Google's official announcement image for the release, reading Introducing Gemini 3.8 Flash and 3.8 Flash Cyber in black type over a pale blue background with a blurred white chevron and the four colour Gemini spark below.

Gemini 3.8 Flash keeps the price and the 1 January cliff

3 September 2026
Answer card stating that Anthropic announced Enterprise Frontier Safeguards on 1 September 2026, that activity data used for misuse monitoring moves into cloud storage the customer controls under the customer own encryption keys, that Anthropic charges nothing for the feature while the cloud provider bills storage and egress, and that the phased rollout starts later in autumn 2026 with interim zero data retention on Fable 5 and Fable 5.1 for eligible customers.

Anthropic moves retention into your own cloud, for 30 days

3 September 2026
Official Google diagram of a client connection in three numbered steps: a DNS lookup with a query and an address, a TLS ClientHello and ServerHello, then a content exchange with a website. A callout on the DNS step reads 25% of global web traffic is now protected by encrypted DNS, and a callout beside an Android phone on the ClientHello step reads Android 17 supports ECH GREASE by default.

Android 17 hides the SNI, not your DNS or destination

3 September 2026
Still frame from the Claude Fable 5.1 launch video showing model-designed protein binders in orange docked against twelve grey target proteins, rendered as ESMFold2 structure predictions.

Claude Fable 5.1 breaks forced tool use, cuts cache 75%

1 September 2026
Answer card stating that on 31 August 2026 the European Commission designated ChatGPT a Very Large Online Search Engine under the Digital Services Act, the first conversational AI service classified that way, because it answers user prompts and queries including by searching the web, with OpenAI having declared roughly 159.1 million average monthly users in the European Union for ChatGPT search.

The EU now calls ChatGPT a very large search engine

3 September 2026
Answer card stating that on 31 August 2026 the Department of War added OpenAI ChatGPT Mil and Starshield AI Grok for Government to the GenAI.mil portal alongside Google Gemini, all three accredited at Impact Level 5 for Controlled Unclassified Information, with 1.7 million unique users onboarded out of roughly 3 million eligible personnel, and ChatGPT Mil currently serving GPT-5.4 Terra with GPT-5.6 Terra said to be rolling out.

ChatGPT Mil and Grok reached IL5 on GenAI.mil

3 September 2026
Answer card stating that Anthropic opened a research preview of the Model Hardware Standard on 27 August 2026, standardising the driver layer between an operating system and a laboratory instrument with read and write primitives plus discovery and safety limits, reachable through MCP as well as a command line and code files, with no public specification published.

Anthropic’s Model Hardware Standard is gated, and sits under MCP

3 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Friday, September 11, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

by stephane
11 September 2026
in Dev
0
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.
491
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Somebody has already pasted the line into your team chat: within one point of Fable 5.1, 64% cheaper. Cognition shipped SWE-2 on 10 September 2026, a coding model post-trained from the 2.8 trillion parameter Kimi K3, and on FrontierCode 1.1 Main it scores 50.0% against 50.9% for Claude Fable 5.1. That's the good row. The other row is Terminal-Bench 4, where SWE-2 lands at 27.3% and Fable 5.1 at 55.8%. Same table, same launch post. Only one of them made the headline.

The short answer

SWE-2 is live in Devin Desktop and the Devin CLI as of 10 September 2026, with Devin Web and Fusion following. There's no public API and no per token price, and the weights stay closed, so the "64% cheaper than Fable 5.1" claim is Cognition's own cost accounting on its own benchmark, and you can't reproduce it. It tops the table on Terminal-Bench 2.1 at 92.8%, sits a point under Fable 5.1 on FrontierCode 1.1 Main, and collapses to 27.3% on Terminal-Bench 4, where Fable 5.1 posts 55.8% and GPT-6 Astra 57.9%. The post doesn't explain that gap.

50.0%FrontierCode 1.1 Main, Fable 5.1 at 50.9%
27.3%Terminal-Bench 4, Fable 5.1 at 55.8%
2.8Tparameters in the Kimi K3 base
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from the 2.8 trillion parameter Kimi K3 base, that it scores 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1, that it scores 27.3 percent on Terminal-Bench 4 against 55.8 percent for Fable 5.1, that Cognition says it is 64 percent cheaper than Fable 5.1, and that it is available only inside Devin with no public API.
Two rows from one table. The launch post quotes the first.

What the table says, both halves of it

Start with the number Cognition wants you to read. FrontierCode 1.1 Main is Cognition's own benchmark, published in July 2026, and on it SWE-2 posts 50.0% at its best effort level. Fable 5.1 posts 50.9%, GPT-6 Astra 53.3%, GPT-5.6 Sol 47.5%. The Kimi K3 base sits at 44.2%, so the reinforcement learning bought 5.8 points, and the previous SWE-1.7 is down at 42.0%. DeepSWE 1.1 reads the same way: 73.0% for SWE-2, 74.1% for Astra. On Terminal-Bench 2.1 SWE-2 actually tops the table at 92.8%, ahead of Fable 5.1 at 91.4%. Three rows, one of them a win.

Horizontal bar chart of FrontierCode 1.1 Main scores as published by Cognition on 10 September 2026, showing GPT-6 Astra at 53.3 percent, Claude Fable 5.1 at 50.9 percent, SWE-2 at 50.0 percent, the Kimi K3 base model at 44.2 percent and the previous SWE-1.7 at 42.0 percent.
Cognition's benchmark, Cognition's harness. Its own model in third.

Then there's the fourth row. Terminal-Bench 4 is the newer, harder terminal benchmark, and it's the one OpenAI leaned on the week before to sell GPT-6 Astra at $10 and $50. On it SWE-2 scores 27.3%. Astra scores 57.9%, Fable 5.1 scores 55.8%, and even GPT-5.6 Sol, which Cognition beats on the headline benchmark, scores 37.3%. The Kimi K3 base is at 21.5%, so the RL lifted it by about six points here too, and SWE-2 still lands under half of the two frontier models. The post prints the row and says nothing about it. Not a footnote, not a caveat.

Honestly, a gap that size isn't noise, and I'd want it explained before I believed the rest. The charitable reading is that Terminal-Bench 4 rewards long, messy shell sessions that Cognition's RL environments don't cover yet. The less charitable one is that a model tuned hard against one lab's benchmark family looks exactly like this: great where it trained, ordinary where it didn't. I might be wrong about which applies. Cognition could settle it in a paragraph, and chose not to.

Horizontal bar chart of Terminal-Bench 4 scores as published by Cognition on 10 September 2026, showing GPT-6 Astra at 57.9 percent, Claude Fable 5.1 at 55.8 percent, GPT-5.6 Sol at 37.3 percent, SWE-2 at 27.3 percent and the Kimi K3 base model at 21.5 percent.
Same launch table, no commentary attached.

64% cheaper than what, and paid how

Now the money. The exact sentence is that SWE-2 reaches 50.0% on FrontierCode 1.1 Main "within one point of Fable 5.1 while being 64% cheaper", and that it "comes within a few points of GPT-6 Astra at a quarter of the cost". The footnote says costs assume list pricing for all models. Fine for Fable 5.1 and Astra, which have rate cards. SWE-2 doesn't. You can't buy it by the token. It lives inside Devin, which bills in its own compute units, and Cognition hasn't published a per token rate for it, any more than it did for SWE-1.7.

So the 64% is two public rate cards against one internal number nobody outside Cognition can see. That's not a lie, and I don't think it is one. It's just unfalsifiable from where you're sitting, which for a purchasing decision amounts to the same thing. The comparison you can check is internal: on the same benchmark, SWE-2 at medium effort scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less on average, and it makes its first real edit after a median of 18 steps against 48. Those we'd trust, because it's the same vendor measuring itself twice with the same ruler.

Fewer steps matters more than the percentage if you run agents at any volume. An agent that spends 48 turns reading before it touches a file is burning budget on exploration, and cutting that to 18 is the sort of change you notice in a day. Whether it stays at 18 on your repo is a different question. Cognition's chart breaks the steps down by phase, and exploration is the bar that shrank. That fits a model that has learned when to stop looking. It also fits one that has learned what FrontierCode tasks look like.

A Chinese open base, a closed product on top

SWE-2 is the second SWE model Cognition has built on a Moonshot base. SWE-1.7, in July, came from Kimi K2.7. This one comes from Kimi K3, the 2.8 trillion parameter model whose weights landed at 1.4 terabytes in late July. Cognition says this is the first time RL of this kind has been scaled to the multi trillion parameter regime, and the technical half of the post is genuinely about that: cost penalties tuned to the Pareto slope at each effort level, plus a draft model for speculative decoding that had to be retrained online because its acceptance rate kept sliding during RL. The three effort levels, medium, high and max, come out of one training run.

None of that reaches you as weights. Kimi K3 is open, SWE-2 isn't, and there's no licence discussion in the post at all. You get a closed product on an open base, which is a perfectly normal business model and worth naming plainly, because "built on an open model" tends to get shortened to "open" by the time it reaches a slide.

Cognition does something with the Chinese origin that we haven't seen other vendors do, and we'll give it credit for that. It reruns its propaganda and censorship test: 145 questions on politically sensitive topics in China, asked in English, Simplified Chinese and Traditional Chinese, each answer graded pass or fail by a single judge, GPT-5.6 Luna. SWE-2 passes 98.0% overall: 99.8% in English, 95.2% in Simplified Chinese, 99.1% in Traditional Chinese. It also reruns a test where identical coding tasks arrive with Western, Pakistani, Chinese, Tibetan or Falun Gong affiliated customer framings, and reports no statistically significant change in how vulnerable the code came out, for any of the six models tested. Good tests. One judge, from a rival lab, on 145 questions, is a thin basis for the word "trustworthy", though, and we'd rerun it before quoting it.

Our take, for what it's worth: if you already pay for Devin, SWE-2 is the model you'll get and it looks like a real step over SWE-1.7 on every row. If you don't, nothing here gives you a price to compare or a terminal benchmark you'd want to lead with. Wait for a number you can put in a spreadsheet.

Sources

Cognition, Introducing SWE-2: Pushing the Pareto Frontier, 10 September 2026 (the benchmark table, the 64% and quarter of the cost claims, the 18 versus 48 steps, the RL method, the propaganda and framing evaluations, the availability wording). Cognition, SWE-1.7: Frontier Intelligence at a Fraction of the Cost, July 2026 (the Kimi K2.7 base of the previous model). OfficeChai, Cognition Releases SWE-2, 10 September 2026 (independent confirmation of the release, the base model and the rollout surfaces). BenchLM, SWE-2 model page (confirms no published pricing and a 1M context listing, with partial benchmark coverage).

Frequently asked questions

What is SWE-2?

Cognition's coding model for Devin, released on 10 September 2026. It's post-trained with reinforcement learning from Kimi K3, Moonshot's 2.8 trillion parameter open weight model, and ships three effort levels, medium, high and max, trained in a single run. Cognition says it's the first RL of this type scaled to a multi trillion parameter base.

Can I call SWE-2 through an API?

Not as of the launch. It's available in Devin Desktop and the Devin CLI, and Cognition says it's rolling out on Devin Web and Fusion. There's no standalone API or per token price, and nothing to download. If that changes we'll add it to this page.

Is SWE-2 really 64% cheaper than Claude Fable 5.1?

That's Cognition's figure, measured on its own FrontierCode 1.1 Main benchmark at each model's best effort level, using list prices for the frontier models and an internal cost for SWE-2 that it doesn't publish. You can't check it from outside. The comparison you can check is against SWE-1.7, where SWE-2 medium costs 81% less on average and takes 58% fewer turns.

Why does SWE-2 score so low on Terminal-Bench 4?

Cognition doesn't say. Its table shows 27.3% for SWE-2 against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra, while on the older Terminal-Bench 2.1 SWE-2 leads at 92.8%. Our guess is that Terminal-Bench 4's longer, messier shell tasks fall outside what the RL environments cover, but that's a guess, and the vendor is the only party that can confirm it.

Does SWE-2 inherit Chinese censorship from Kimi K3?

Cognition tested for it. On 145 politically sensitive questions about China, asked in three languages and judged by GPT-5.6 Luna, SWE-2 passed 98.0% overall, with the lowest rate in Simplified Chinese at 95.2%. That's one judge on a small set, so treat it as a signal rather than a verdict.

Tags: benchmarksClaude Fable 5.1coding agentsCognitionDevinGPT-6 AstraKimi K3newsSWE-2
Share196Tweet123
stephane

stephane

  • Trending
  • Comments
  • Latest
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
Answer card: Apple released iOS 26.6 and iPadOS 26.6 on 27 July 2026 with a release note covering bug fixes, security updates and an optimized Spotlight index to prepare for iOS 27, the index the rebuilt Siri reads for personal context.

iOS 26.6 is out: the Spotlight index it quietly builds

27 July 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
Answer card stating that the public download pages for the VMware Virtual Disk Development Kit on developer.broadcom.com began returning 404 errors on 25 August 2026 with no announcement or deprecation notice, that Broadcom support tells customers the kit is no longer available for use or download, and that release lines 7.0.3.1, 8.x and 9.x are all affected.

Broadcom pulled VDDK 8.0 and 9.0, and the 404 is the only notice

8 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.