• Latest
  • Trending
  • All
Answer card: on 5 August 2026 Anthropic confirmed it is building an in-house custom silicon team to co-design chips and Claude models, while keeping its multi-chip approach across AWS, Google, Nvidia and AMD, with no tape-out date, no foundry partner and no product announced.

Anthropic’s custom silicon team: no chip, no timeline

7 August 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Monday, September 21, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Anthropic’s custom silicon team: no chip, no timeline

by stephane
7 August 2026
in Dev
0
Answer card: on 5 August 2026 Anthropic confirmed it is building an in-house custom silicon team to co-design chips and Claude models, while keeping its multi-chip approach across AWS, Google, Nvidia and AMD, with no tape-out date, no foundry partner and no product announced.
494
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Six job listings. That's most of the announcement, and honestly it's more informative than the headlines stacked on top of it. On 5 August Anthropic confirmed it's building a custom silicon team to co-design chips and Claude models, then kept every existing supplier in place: AWS, Google, Nvidia and AMD all stay. No tape-out date. No foundry named. What we can actually verify is the careers page, which lists a Silicon Engineer, a Hardware Systems Architect, a Technical Program Manager for Silicon and a few more, at $320,000 to $485,000. One listing asks the hire to support first silicon bring-up and debug when it arrives. When it arrives. That's a team being assembled, not a chip being shipped, and the difference matters if you're sizing Claude spend for next year.

The short answer

On 5 August Anthropic confirmed it’s building a custom silicon team to co-design chips with Claude models. Everything else stays: AWS, Google, Nvidia and AMD remain in the plan, and no tape-out date, foundry or product was named. The strongest evidence is the careers page, where the listings ask for engineers who have shipped silicon and mention first silicon bring-up as a future event. The 50 percent inference saving doing the rounds is outlet framing, not an Anthropic number.

0Anthropic-designed chips serving Claude
$485ktop of the published silicon band
4chip suppliers explicitly kept
Answer card: on 5 August 2026 Anthropic confirmed it is building an in-house custom silicon team to co-design chips and Claude models, while keeping its multi-chip approach across AWS, Google, Nvidia and AMD, with no tape-out date, no foundry partner and no product announced.
The one-card version. A team, a salary band, and a chip that doesn't exist yet.

What was confirmed, and what was inferred

The confirmation came through Business Insider and was quickly matched by everyone else. The quote is short: Anthropic works from the chip level up with its silicon partners, and is now deepening that investment by building a custom silicon team. The stated goal is co-design, meaning hardware and Claude models shaped around each other so the models run faster and more efficiently at the scale customers need.

Then the important half. Anthropic said it keeps a multi-chip approach across AWS, Google, Nvidia and AMD. Read that carefully and it’s a hedge with four names in it. Nothing is being replaced.

What nobody got was a date. No tape-out, no sampling window, no node, no foundry. Reuters reported back in April that Anthropic was exploring chip design, and The Information reported talks with Samsung in July, but talks aren’t a manufacturing deal and Anthropic hasn’t confirmed one. So when a headline says Anthropic is building chips, the accurate version is that Anthropic is building the team that would build chips.

Checklist comparing what Anthropic confirmed on 5 August 2026 about its custom silicon team, including the co-design goal, the continued multi-chip approach and the published salary band, against what it left undisclosed including the tape-out date, the foundry partner, the inference or training target and any effect on Claude pricing.
Five things Anthropic said. Four things it didn't.

The job board is the most honest document

This is the part I’d point at if I only had thirty seconds. Anthropic’s careers page currently carries a Silicon Engineer role, a Hardware Systems Architect, a Technical Program Manager for Silicon, a Hardware Lab Manager, a TPU Kernel Engineer and a Research Engineer working on chip design with reinforcement learning. The published band runs from $320,000 to $485,000.

The wording gives away the stage better than any press coverage. One listing asks for someone who has shipped silicon and is comfortable making consequential calls without a large organization behind them. Another says the hire will support first silicon bring-up and debug when it arrives. Bring-up is what happens weeks after a wafer comes back from a fab. Writing it as a future conditional means there’s no wafer, and there’s no fab commitment either.

The breadth of disciplines is the other tell. Front-end design, pre-silicon verification, physical design, design-for-test, analog and mixed-signal, packaging with signal and power integrity. That’s not one team. That’s the org chart of a chip company, being recruited from scratch, by a company whose entire staff would fit inside one floor of a traditional silicon vendor. I might be wrong about the pace, but hiring across that many disciplines simultaneously usually means year one, not year three.

Does it change your Claude bill

No. Not this year, and I’d be surprised at next year too.

Custom accelerators are a fixed-cost bet before they’re a variable-cost win. You pay for design tools, mask sets, verification and bring-up long before a single token gets cheaper, and the payoff only lands if the volume is enormous and the model architecture stays stable enough for the silicon to still fit it two years later. That second condition is the interesting one for a lab shipping new frontier models every few months.

The 50 percent per-token saving that showed up in a few write-ups is worth flagging. We couldn’t find it in anything Anthropic said. It reads like an analyst estimate that got repeated until it acquired quotation marks, which is exactly the kind of number that ends up in someone’s budget spreadsheet as a fact. If your capacity planning already runs close to the edge, the pattern to watch is the one we saw when a $1.8 million Claude bill came in 860 percent over forecast: the cost lever that actually works today is caching and routing, not future hardware.

Bar chart comparing Anthropic's contracted compute: roughly 3.5 gigawatts of next generation Google TPU capacity from 2027, well over one gigawatt of Google TPU capacity online during 2026, and zero gigawatts served by silicon Anthropic designed itself.
Scale of what's already signed, against the chip that hasn't been designed.

Put it next to the contracted capacity and the proportions get clear. Anthropic’s October 2025 agreement with Google Cloud gives it access to up to a million TPUs and was expected to bring well over a gigawatt online during 2026, for tens of billions of dollars. The expanded deal reported in April 2026, with Broadcom involved, adds roughly 3.5 gigawatts of next-generation TPU capacity from 2027. Against that, an in-house part is a rounding error for years.

Everyone is doing this now, which is the actual story

OpenAI put its name on the Broadcom-built Jalapeno inference chip in June. Meta has been pushing its own accelerator into production. Google has shipped TPUs for a decade. Amazon has Trainium. Anthropic joining is less a surprise than a confirmation that no frontier lab thinks it can buy its way out of compute constraints on the merchant market alone.

What separates them is where they are on the curve. OpenAI has a named partner and a named part. Anthropic has a hiring page. Both facts are worth knowing, and conflating them is how a reasonable strategic move turns into a headline about Nvidia losing another customer. Nvidia is still in Anthropic’s own list of four.

For anyone building on the API, the practical takeaway is short: change nothing. Watch for a foundry announcement or a stated tape-out. Those are the events that would move a date onto a calendar. A job listing isn’t one.

Sources

  • Anthropic careers page, for the live silicon and hardware roles, checked 7 August 2026.
  • TechCrunch, “Anthropic is hiring an AI chip design team”, 5 August 2026, for the confirmation and the list of existing chip suppliers.
  • Unite.AI, for the spokesperson wording, the quoted job listing phrases, the salary band and the disciplines being recruited.
  • Techzine, for the Reuters and The Information reporting timeline including the Samsung talks.
  • Anthropic, “Expanding our use of Google Cloud TPUs and services”, 23 October 2025, for the one million TPU ceiling and the gigawatt of capacity expected in 2026.
  • Forbes, 6 August 2026, for the 3.5 gigawatts of next-generation TPU capacity from 2027.

Frequently asked questions

What did Anthropic actually announce on 5 August 2026?

That it is building an in-house custom silicon team. A spokesperson said the company works from the chip level up with its silicon partners and is now deepening that investment by building a custom silicon team, with the aim of co-designing hardware and models so Claude runs faster and more efficiently at the scale customers need. Anthropic also said it keeps a multi-chip approach across AWS, Google, Nvidia and AMD. It announced no chip, no tape-out date and no manufacturing partner.

Will this make Claude cheaper?

Not on any timeline Anthropic has given. Several outlets framed the move as targeting roughly half the per-token inference cost, but that figure is coverage framing rather than a number Anthropic published, so treat it as unconfirmed. Custom accelerators are also a fixed-cost play: you pay for design, masks and bring-up years before the unit economics improve. Nothing about your current bill changes because a job listing went up.

Is Anthropic dropping Nvidia?

No. The confirmation explicitly keeps the multi-chip approach across AWS, Google, Nvidia and AMD, and the in-house effort sits alongside those suppliers rather than replacing them. Anthropic already has enormous contracted capacity on Google TPUs, so its own silicon would join a mix rather than take it over.

How does this compare with OpenAI and Meta?

It is a later start on the same path. OpenAI unveiled its Broadcom-built Jalapeno inference chip in June 2026, Meta has been producing its own accelerator silicon, and Google has shipped TPUs for years. Anthropic is at the stage of hiring people who have shipped silicon before, which is roughly where those programs were several product generations ago.

What would count as real progress here?

A named foundry and process node, a stated tape-out or sampling window, and a clear answer on whether the part targets inference or training. Until at least one of those is public, the honest read is that Anthropic has a team, a budget line and a hiring page.

Tags: aianthropicchipshardwarellmnews
Share198Tweet124
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.