• Latest
  • Trending
  • All
Answer card: Fireworks closed a 1.505 billion dollar Series D on July 16 2026 at a 17.5 billion dollar valuation, led by Atreides Management, Index Ventures and TCV with Nvidia among the backers, as an AI inference platform for serving and fine-tuning open-weight models.

Fireworks raised $1.5B, and a raise ships no code

3 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
Answer card stating that the public download pages for the VMware Virtual Disk Development Kit on developer.broadcom.com began returning 404 errors on 25 August 2026 with no announcement or deprecation notice, that Broadcom support tells customers the kit is no longer available for use or download, and that release lines 7.0.3.1, 8.x and 9.x are all affected.

Broadcom pulled VDDK 8.0 and 9.0, and the 404 is the only notice

8 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Monday, September 21, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Fireworks raised $1.5B, and a raise ships no code

by stephane
3 September 2026
in Dev
0
Answer card: Fireworks closed a 1.505 billion dollar Series D on July 16 2026 at a 17.5 billion dollar valuation, led by Atreides Management, Index Ventures and TCV with Nvidia among the backers, as an AI inference platform for serving and fine-tuning open-weight models.
495
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

You read our piece on Qwen 3.8, nodded at the open weights, then hit the same wall we always do. Great, now where do you actually run a near trillion-parameter model without buying a rack of GPUs? On July 16 one of the answers got a lot bigger. Fireworks, the platform plenty of teams lean on to serve and fine-tune open models, closed a $1.505 billion Series D at a $17.5 billion valuation, with Nvidia among the backers. Here's the honest read: a funding round changes nothing about your code today. What it signals is where the money thinks inference is heading, and it isn't toward one giant closed API. It's toward open models you bend to your own data. We've been telling that story model by model. This is the infrastructure bet sitting underneath it.

The short answer

Fireworks, an inference platform for open-weight and fine-tuned models, closed a $1.505 billion Series D on July 16 at a $17.5 billion valuation, with Nvidia among the backers. The raise doesn’t touch your code. What it marks is a bet: companies serving open models tuned to their own data, not one closed API for everyone. That’s the trend we’ve been covering release by release.

$1.505BSeries D raised
$17.5Bpost-money valuation
40T+tokens served per day (company)
Answer card: Fireworks closed a 1.505 billion dollar Series D on July 16 2026 at a 17.5 billion dollar valuation, led by Atreides Management, Index Ventures and TCV with Nvidia among the backers, as a platform for serving and fine-tuning open-weight models.
The raise in one card. Big number, familiar plumbing underneath.

What actually happened

On July 16 Fireworks said it raised a $1.505 billion Series D at a $17.5 billion valuation. The round was jointly led by Atreides Management, Index Ventures and TCV. Nvidia is in there too, along with Lightspeed, Bessemer and Menlo Ventures. The stated plan is dull in the good way: more compute, more engineers, more platform.

The company put some scale numbers next to it. Over $1 billion in annualized revenue. More than 40 trillion tokens served every day. And the line that matters most for where this is going: over 95% of those tokens come from models specialized on a customer’s own data, not off-the-shelf general models. That last figure is Fireworks own, so hold it as a claim, not an audit. It’s still the whole thesis in one stat.

What Fireworks even is

If you’ve only ever hit a closed API, here’s the short version. Fireworks runs open-weight models for you. You send prompts to a serverless endpoint, or you pin a model to dedicated GPUs when traffic justifies it. You can also fine-tune an open model on your own data through managed clusters, with the usual machinery: quantization to shrink the memory bill, autoscaling, distributed training. Named users, per the company and reporting, include Cursor for coding, Harvey in legal, plus Samsung Electronics and GitLab.

It was started in 2022 by former Meta engineers, with Lin Qiao as CEO. So the pedigree is people who ran model infrastructure at scale before, which is the boring detail that actually predicts whether an inference platform stays up.

Bar comparison of Fireworks daily token mix as reported by the company: over 95 percent from models specialized on customer data versus under 5 percent from general-purpose off-the-shelf models.
The bet in one bar. Almost all the traffic is specialized models, per Fireworks own figures.

Why we’re writing about a funding round

Normally we skip raises. A pile of money isn’t a product, and it’s not a benchmark.

This one earns a paragraph because of what it’s a proxy for. We spend a lot of these posts on open weights: Qwen 3.8, GLM 5.2, DeepSeek V4, Inkling. Every single time, the comments boil down to one question. Where do I run this thing? Inkling alone wants 2 TB of VRAM in BF16. You’re not doing that on a laptop.

Serving platforms are one answer, and $1.5 billion of investor money landing on the “open and specialized” answer, rather than the “just use one frontier API” answer, is a real signal about which way the wind is blowing. Lin Qiao’s framing is that every company holds knowledge no one else has, its data and its definition of quality, and should own the model built on it. You don’t have to buy the pitch to notice the capital agrees with it.

Platform or your own metal

So do you reach for something like this, or run the weights yourself? Depends on two things, scale and control.

Checklist separating what the Fireworks raise signals from what it does not change: it confirms investor conviction in open specialized models and funds more capacity, but it ships no feature today, cuts no price you can see, and the 95 percent and 40 trillion figures are company-reported and not audited.
Signal against the headline number. What moves, and what's just a big figure.

If your workload is small, or the data can’t leave your walls, run it local. That’s the whole point of our Qwen 3.7 offline walkthrough: weights on hardware you own, nothing leaving the building. No per-token meter.

If you’ve got bursty production traffic on an open model and you don’t want to own or babysit GPUs, a platform usually wins on total effort, even with the per-token or per-GPU-hour cost and the vendor lock-in that comes with it. The honest caveat: once your prompts and fine-tunes live on someone’s platform, moving off it is real work. Price that in before you commit anything load-bearing.

The honest read

I’m a little wary of cheering a funding round, and I’ll say why. Nothing shipped. No new model, no price cut, no benchmark you can rerun tonight. The scale numbers are the company’s own, and $17.5 billion is a valuation, not revenue. Big raises in a hot market sometimes age badly.

What’s real is the direction. The money is voting for open, customized models served at scale, and that lines up with everything we keep seeing when a new open release drops and the first question is where to run it. If that’s you, this is worth a bookmark, not a migration. Try a serving platform against your own prompts, put the token bill next to what your closed API costs, and let that number decide. It’s the only comparison that shows up on the invoice.

Sources: Fireworks’ Series D announcement, with independent reporting via SiliconANGLE and CNBC, July 2026. Revenue, token-volume and specialization figures are as stated by the company and are not independently audited.

Frequently asked questions

What is Fireworks AI?

Fireworks is a cloud platform for running open-weight and fine-tuned AI models in production. You can serve models through a serverless API or on dedicated GPU deployments, and fine-tune open models on your own data with managed clusters. Customers named by the company and reporting include Cursor, Harvey, Samsung Electronics and GitLab. It was founded in 2022 by former Meta engineers, with Lin Qiao as CEO.

How big was the Fireworks Series D?

Fireworks announced a $1.505 billion Series D on July 16, 2026 at a $17.5 billion valuation. It was jointly led by Atreides Management, Index Ventures and TCV, with Nvidia, Lightspeed, Bessemer, Menlo Ventures and others also in the round. The company says the money goes to more compute and more engineers.

Does this raise change anything for developers today?

No. A funding round does not ship a feature or cut a price on its own. What it tells you is direction: heavy investor money betting that companies will run open, specialized models rather than route everything through one closed API. If you already deploy open weights, treat this as confirmation of the trend, not a reason to switch platforms.

Should I use a platform like Fireworks or run open models myself?

It depends on scale and control. A serving platform saves you from owning GPUs and handles autoscaling and quantization, at a per-token or per-GPU-hour cost and some vendor lock-in. Running it yourself, as we cover in our Qwen 3.7 offline guide, keeps everything local and private but means you buy and babysit the hardware. Small and private, run local. Bursty production traffic on open weights, a platform usually wins.

Is the 95% specialized-tokens figure independently verified?

No. The claim that more than 95% of its 40 trillion daily tokens come from models specialized on customer data is Fireworks own number, from its announcement. It fits the story the round is selling, so we report it as a company figure, not an audited one.

Tags: aiinferencellmnewsopen-sourceself-hosting
Share198Tweet124
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.