• Latest
  • Trending
  • All
Answer card: Ox Alpha appeared on OpenRouter on 20 August 2026 as stealth/ox-alpha with a 1,048,576 token context, 131,072 tokens of output, text image and video input, zero pricing during the preview, and terms stating that prompts and completions are retained by an anonymous provider and not used for training.

A free 1M context coding model appeared with no lab behind it

3 September 2026
The official xAI announcement card for Grok 4.7, white type on a dark grey and navy gradient.

Grok 4.7 keeps $2 and $6, and its gains over 4.6 are xhigh versus high

22 September 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Tuesday, September 22, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

A free 1M context coding model appeared with no lab behind it

by stephane
3 September 2026
in Dev
0
Answer card: Ox Alpha appeared on OpenRouter on 20 August 2026 as stealth/ox-alpha with a 1,048,576 token context, 131,072 tokens of output, text image and video input, zero pricing during the preview, and terms stating that prompts and completions are retained by an anonymous provider and not used for training.
498
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

You open the OpenRouter model list, sort by price, and there it is at the top: ox-alpha, a million tokens of context, zero dollars in and zero dollars out. Free coding model, listed 20 August 2026, and no lab has put its name on it. That is the whole story and also the catch. OpenRouter is explicit that it only routes the requests, that the thing is run by a third party who chose to stay anonymous for the preview, and that prompts and completions are retained by that provider. Not used for training, the terms say. Retained, though. So the question is not whether Ox Alpha is good, because the early spot checks say it is decent. The question is what you are willing to send to an operator you cannot name.

The short answer

Ox Alpha went up on OpenRouter on 20 August as stealth/ox-alpha, a coding and agent model from an operator nobody has identified. It is genuinely capable on the early spot checks. It also retains everything you send it, under terms written by a company with no name on them. Use it on code you would happily publish.

1,048,576token context window
$0in, out and cache reads, for now
27 Augreported end of the free window
Answer card: Ox Alpha appeared on OpenRouter on 20 August 2026 as stealth/ox-alpha with a 1,048,576 token context window, 131,072 tokens of maximum output, text image and video input, zero pricing during the preview reported to end 27 August, and terms stating that prompts and completions are retained by an anonymous provider and not used for training.
Everything on the listing, and the one line most people scrolled past.

What is actually on the listing

The documented part is short, so here it is in full.

Model ID stealth/ox-alpha. Context window 1,048,576 tokens, with a ceiling of 131,072 tokens on the completion. Input accepts text, images and video. Output is text. Function calling works through the usual tools and tool_choice parameters, and JSON response formatting is supported. The description on the page pitches it at long horizon software engineering and workflows that mix text with visual context.

Pricing during the preview is zero on prompt tokens, completion tokens and cache reads. The reporting around the launch puts the free window at roughly a week from 20 August, so it closes around 27 August, and OpenCode pushed it to its users on the same free basis.

When we looked at the model page, the live stats read 6.01 seconds median latency, 23 tokens per second of throughput, and three day uptime of 99.99 percent. Twenty-three tokens a second is not fast. For a reasoning model doing agent work that is survivable, for a chat UI it would feel sluggish.

Comparison chart of a ten task DeepSWE spot check: Ox Alpha at 80 percent, Claude Fable 5 at 65 percent and GPT-5.6 at 52 percent, with a note that ten tasks is a spot check rather than a full benchmark suite.
One tester, ten tasks. Worth knowing, not worth quoting as a benchmark.

The 80 percent number, and what it is worth

The headline everywhere is that Ox Alpha closed out DeepSWE at 80 percent, beating Claude Fable 5 at 65 and GPT-5.6 at 52. Developer Ben Davis ran that, and to his credit he described his own method plainly.

Ten tasks. Pulled from DeepSWE, not the whole suite. And the models were not all given the same number of attempts. That is a spot check, the kind of thing you do on a Thursday evening to decide whether a new model is worth a second look, and it answered that question well. It does not sit alongside the SWE-bench Verified figures the frontier labs publish, which come off a much larger harness and land far higher for everyone involved.

Honestly, the detail from the same tester that impressed us more was the agent run: 69 tool calls, one error, no retry loop. Long tool chains are where cheap models fall apart, usually by looping on a failed call until the budget dies. Not looping is a real signal.

Who is running it

Nobody has said, and the guessing has been thorough.

Davis puts himself at 99 percent certainty that this is Zhipu’s GLM family, on four pieces of evidence. Video encoder token consumption matches GLM-5V-Turbo, at roughly 147 tokens per second of clip. Tokenizer counts match GLM-5.3 exactly, offset by a fixed 75 token wrapper. Audio input gets rejected the same way GLM-5V rejects it. And the output carries about 1.3 emojis per thousand characters, which lines up with the GLM and Qwen house style.

That is careful work and it is still circumstantial. Nobody at Zhipu has confirmed anything, and a tokenizer can be shared or copied. If it is GLM, the timing fits: GLM-5.3 shipped on 14 August with the weights held back, and a cloaked multimodal sibling running free for a week is exactly how you collect a week of real agent traffic before you name the thing.

There is precedent for the whole ritual. In September 2025, two models called Sonoma Sky Alpha and Sonoma Dusk Alpha sat free and anonymous on OpenRouter with a 2M context, climbed the charts, and were confirmed a few weeks later as xAI’s Grok 4 Fast in reasoning and non reasoning form. Same pattern of a leak-shaped launch, then a real name. We wrote about the gap between a confirmed launch and the leak noise around it when Kimi K3 landed, and the lesson holds here.

Checklist separating what is documented about Ox Alpha, including the listing date, context and output limits, modalities, zero pricing and the retention clause, from what is unknown, including the operator identity, the price after the preview, the retention window and any full benchmark run.
Four things you can plan around, three you cannot.

The line in the terms

Here is the sentence, from the model page: prompts and completions for this model are retained by the provider and are not used for training.

Read both halves. Not used for training is the reassuring half, and it is also the half that is easy to promise. Retained by the provider is the operative half. Retained where, for how long, under which country’s disclosure rules, by a legal entity with what name. None of that is published, and there is no one to ask, because the point of a stealth listing is that there is no one to ask.

For open source work, a side project, a benchmark harness, a throwaway script, that is a non-issue. Free frontier-ish coding with a million tokens of context is a genuinely good deal and you should go use it before 27 August.

For your employer’s private repository, it is not a close call. Don’t.

Our own rule while these previews run: point the agent at a scratch checkout, keep credentials out of the environment it can read, and treat every token you send as public. Which, incidentally, is a reasonable default for any model endpoint whose operator you would not name in a compliance review.

What we would do this week

Try it. Give it a real task with a long tool chain, because that is the part the spot check suggested is strong and the part that would show up as weak fastest.

Then write down what you actually get, because the price is going to change. Every cloaked preview ends the same way: a name, a rate card, and a number that is not zero. If Ox Alpha is worth paying for, you want that judgement made on your own workload rather than on ten tasks somebody ran on a Thursday.

I might be wrong about the GLM attribution, incidentally. The tokenizer evidence is strong but it is the kind of strong that has been wrong before.

Sources

Model listing, pricing, context limits and data policy: Ox Alpha on OpenRouter. The ten task DeepSWE spot check and the tokenizer fingerprinting are reported in this analysis of the stealth model and in AI/TLDR’s release note. The Sonoma precedent comes from OpenRouter’s own reveal post.

Frequently asked questions

What is Ox Alpha?

A reasoning model aimed at coding and long running agent work, listed on OpenRouter on 20 August 2026 under the model ID stealth/ox-alpha. It carries a 1,048,576 token context window, returns up to 131,072 tokens, and accepts text, images and video as input. OpenRouter states that it routes requests to the model and is not its developer, owner or provider.

Who made Ox Alpha?

Nobody has said. The provider is listed as Stealth and no lab has claimed it. Community fingerprinting points at Zhipu's GLM family, based on tokenizer counts that match GLM-5.3 with a fixed offset and on video encoder behaviour that matches GLM-5V-Turbo. That is analysis by an outside tester, not a confirmation from anyone who would know.

How long is Ox Alpha free?

The listing shows zero for input, output and cache reads during the preview, and the coverage around the launch reports a window of roughly one week from 20 August, so through 27 August 2026. No price for what comes after has been published, and no rate card, quota or retention window either.

Is it safe to use Ox Alpha on work code?

We would not. The data policy on the listing says prompts and completions are retained by the provider and are not used for training. Retention by an operator whose legal identity, jurisdiction and retention period are all unpublished is a poor fit for proprietary source, customer data or anything under an NDA. Open source work and throwaway experiments are a different matter.

Did the 80 percent DeepSWE score come from a full benchmark run?

No. It came from a ten task spot check by developer Ben Davis, who reported 80 percent for Ox Alpha against 65 percent for Claude Fable 5 and 52 percent for GPT-5.6 on the same ten tasks, with the models not all given the same number of attempts. It is a useful signal and it is not a benchmark result.

Tags: aicodingllmnewsopenrouterprivacy
Share199Tweet125
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
The official xAI announcement card for Grok 4.7, white type on a dark grey and navy gradient.

Grok 4.7 keeps $2 and $6, and its gains over 4.6 are xhigh versus high

22 September 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.