You wired up grok-voice-latest because xAI’s own quickstart does. On 5 August that string began answering from a newer model at $0.08 a minute instead of $0.05, and not one line of your code changed. That’s the gap this page is for. It keeps the current API price of 37 language models, plus the voice and document models that bill by the minute or the page, next to the details a headline number leaves out: the cache-read rate, the context step, the discount with a date on it, the alias that moves. It’s for whoever picks the model string and then has to explain the invoice.
Last verified: 7 October 2026. Every price was read on the vendor’s own page that day, apart from the two rows marked Reported.
How to read this page
Each row comes from the vendor’s pricing or model page, opened on 7 October. Our older articles supplied the questions. They didn’t supply the answers: when one disagrees with the vendor, the vendor wins, and the log says so. Only the two Xiaomi rows lean on press and router listings, because the platform page wouldn’t load for us.
Prices are US dollars per million tokens at the standard tier: short context, global endpoint, no batch discount, no fast-mode premium. Batch halves the bill at Anthropic and OpenAI, and Google matches it; xAI’s Grok 4.7 page says batch isn’t supported. Fast modes and regional processing add to it. None of that is in the numbers, and the FAQ covers it.
Status is one of three things. Confirmed means the vendor’s page says it today. Announced means there’s a price but no model you can call yet. Reported means we couldn’t check the source ourselves.
We left out open-weight models that only carry router prices (Laguna, Hy4, Nemotron and Inkling), because a router’s figure changes with whichever provider answers. Read the Price caveat column before the price columns. That’s where the expiry dates and the context cliffs live.
The price table
37 language models come first, then the voice and document models that don’t bill by the token, then the dates that will move your bill. Rows are grouped by vendor, newest model first. Sort by the caveat, not by the price.
Language models, USD per million tokens
| Vendor | Model | Input $/M tokens | Output $/M tokens | Cache read $/M | Context | Released | Price caveat | Status | Source | Our write-up |
|---|---|---|---|---|---|---|---|---|---|---|
| Anthropic | Claude Fable 5.1claude-fable-5-1 | $10 | $50 | $0.25 | 1M (128K out) | Cache reads at 0.025x base, a quarter of Fable 5’s rate. Thinking always on, so the effort setting sets the real bill. Batch $5/$25. | Confirmed | Anthropic pricing Anthropic models overview | Fable 5.1 cache and tool use | |
| Anthropic | Claude Opus 5.5claude-opus-5-5 | $4 | $20 | $0.20 | 1M (128K out) | Thinking can’t be disabled. Default effort is medium, and Anthropic says it thinks more per turn than Opus 5 at the same effort. Fast mode (preview) $8/$40. | Confirmed | Anthropic pricing Anthropic release notes | Opus 5.5 breaking changes | |
| Anthropic | Claude Sonnet 5.5claude-sonnet-5-5 | $2 | $10 | $0.10 | 1M (128K out) | Cache read cut from $0.20 on 7 Oct. Anthropic’s model page and price table still print $0.20; the release note and the caching section say $0.10. | Confirmed | Anthropic release notes Anthropic pricing | None yet | |
| Anthropic | Claude Haiku 5.5claude-haiku-5-5 | $0.10 up to 100K prompt, $0.50 above | $0.50 up to 100K, $2.50 above | $0.01 up to 100K, $0.05 above | 1M (128K out) | The only current Claude model with a prompt-length step: everything reprices past 100K tokens. Launched the day we checked. | Confirmed | Haiku 5.5 model page Anthropic release notes | The 100K cliff | |
| Anthropic | Claude Opus 5claude-opus-5 | $5 | $25 | $0.50 | 1M (128K out) | Legacy, retirement not before 24 Jul 2027. Thinking on by default. Fast mode $10/$50. | Confirmed, legacy | Anthropic pricing Anthropic model deprecations | Opus 5 launch | |
| Anthropic | Claude Sonnet 5claude-sonnet-5 | $2 | $10 | $0.20 | 1M (128K out) | Launch price became the standard price on 10 Aug. The planned $3/$15 from 1 Sep never applied. | Confirmed, legacy | Anthropic release notes Anthropic pricing | None yet | |
| Anthropic | Claude Fable 5claude-fable-5 | $10 | $50 | $1.00 | 1M (128K out) | Same base price as 5.1, four times the cache-read price. Retirement not before 9 Jun 2027. | Confirmed, legacy | Anthropic pricing Anthropic model deprecations | Fable 5 effort levels | |
| OpenAI | GPT-6 Astragpt-6-astra | $10 | $50 | $1.00 | 1.05M (128K out) | Past 272K input tokens the whole request bills $20/$75. Fast mode 2x; Ultrafast tier $60/$300. First access went to vetted enterprises. | Confirmed | OpenAI pricing | GPT-6 Astra and the 272K cliff | |
| OpenAI | GPT-6.1 Solgpt-6.1-sol | $2 | $10 | $0.10 | 1.05M (128K out) | Past 272K input: $4/$15 on the whole request. Rejects reasoning effort none and minimal. Fast mode $4/$20. Ultrafast tier (from 8 Oct) $12/$60, 6x standard; long-context billing not documented. | Confirmed | OpenAI pricing | GPT-6.1 Sol cache cut | |
| OpenAI | GPT-6 Solgpt-6-sol | $2 | $10 | $0.20 | 1.05M (128K out) | Same rates as 6.1 Sol except cached input costs double. Past 272K: $4/$15. There is no GPT-6 Terra. | Confirmed | OpenAI pricing | GPT-6 Astra and the 272K cliff | |
| OpenAI | GPT-6 Lunagpt-6-luna | $0.10 | $0.50 | $0.01 | 1.05M (128K out) | Past 272K input: $0.20/$0.75 on the whole request. Decisions API (beta): $0.10 in, no output charge, no caching. | Confirmed | OpenAI pricing | GPT-6 Astra and the 272K cliff | |
| OpenAI | GPT-5.6 Solgpt-5.6-sol | $4 | $20 | $0.40 | 1.05M (128K out) | Promo price: was $5/$30 until 21 Aug, promised “at least” through 21 Nov 2026. Past 272K: $8/$30. The ChatGPT version of Sol is tuned differently from the API one. | Confirmed, promo | OpenAI pricing GPT-5.6 Sol model page | Sol in ChatGPT vs the API | |
| OpenAI | GPT-5.6 Terragpt-5.6-terra | $2 | $12 | $0.20 | 1.05M (128K out) | Launched at $2.50/$15, cut 20% on 30 Jul. Past 272K: $4/$18. | Confirmed | OpenAI pricing | None yet | |
| OpenAI | GPT-5.6 Lunagpt-5.6-luna | $0.20 | $1.20 | $0.02 | 1.05M (128K out) | Launched at $1/$6, cut 80% on 30 Jul. Past 272K: $0.40/$1.80. | Confirmed | OpenAI pricing | None yet | |
| OpenAI | o3o3-2025-04-16 | $2 | $8 | $0.50 | 200K (100K out) | Shuts down 11 Dec 2026, along with o3-pro. OpenAI names gpt-5.6-sol as the replacement, at double the input and 2.5x the output price. | Confirmed, retiring | OpenAI pricing OpenAI deprecations | o3 shutdown dates | |
Gemini 3.8 Flashgemini-3.8-flash | $0.75 | $3.75 | $0.075 | 1M (65K out) | Intro rate through 31 Dec 2026, then $1.50/$7.50 (cache $0.15) from 1 Jan 2027. Spends more tokens on purpose at high effort. Thinking level minimal removed. | Confirmed, intro price | Gemini API pricing Gemini API changelog | Gemini 3.8 Flash and 1 January | ||
Gemini 3.7 Flashgemini-3.7-flash | $0.75 | $3.75 | $0.075 | 1M (65K out) | Same intro rate and the same 31 Dec end date as 3.8. Google says it stays supported for efficiency-first work. | Confirmed, intro price | Gemini API pricing | None yet | ||
Gemini 3.6 Flashgemini-3.6-flash | $0.75 | $3.75 | $0.075 | 1M (65K out) | Launched at $1.50/$7.50. The pricing page now shows the same intro rate as 3.7 and 3.8; we couldn’t date the change. | Confirmed, intro price | Gemini API pricing | None yet | ||
Gemini 3.5 Flash-Litegemini-3.5-flash-lite | $0.30 | $2.50 | $0.03 | 1M (65K out) | No intro discount listed. Cache storage adds $1.00 per million tokens per hour. | Confirmed | Gemini API pricing | None yet | ||
Gemini 4 Argonnot listed yet | $2 | $10 | 95% off input (about $0.10) | 1M output limit | Intro price, then $4/$20 with no end date given. Not on the Gemini API pricing page or changelog on 7 Oct. Trusted cyber defenders only for now. | Announced | Google, Gemini 4 Argon Gemini API pricing | Gemini 4 Argon pricing | ||
| xAI | Grok 4.7grok-4.7 | $2 | $6 | $0.50 | 500K | Past 200K the whole request bills $4/$12 (cache $1.00). Reasoning can’t be turned off. The launch table ran 4.7 at xhigh against 4.6 at high. | Confirmed | xAI, Grok 4.7 page xAI models | Grok 4.7 xhigh vs high | |
| xAI | Grok 4.6grok-4.6 | $2 | $6 | $0.50 | 500K | Cache read was $0.30 on 4.5. Same 200K step as 4.7. | Confirmed | xAI models | Grok 4.6 and the 200K cliff | |
| xAI | Grok 4.5grok-4.5 | $2 | $6 | $0.30 | 500K | Cheapest cache read of the three. Past 200K: $4/$12 (cache $0.60). | Confirmed | xAI models | None yet | |
| DeepSeek | DeepSeek-V4.1-Flashdeepseek-flash | $0.15 off-peak, $0.30 peak | $0.60 off-peak, $1.20 peak | $0.003 off-peak, $0.006 peak | 1M (384K out) | Peak is 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, at double. The retired name deepseek-v4-flash is routed here “temporarily”, at these prices. | Confirmed | DeepSeek models and pricing DeepSeek change log | DeepSeek V4.1-Flash | |
| DeepSeek | DeepSeek-V4-Pro-0813deepseek-v4-pro | $0.66 off-peak, $1.32 peak | $1.98 off-peak, $3.96 peak | $0.022 off-peak, $0.044 peak | 1M (384K out) | Same peak rule. A retirement notice for 14 Sep was reversed on 11 Sep, with “further notice” promised. Text only; 500 concurrent requests against 2,500 on Flash. | Confirmed | DeepSeek models and pricing DeepSeek change log | DeepSeek V4.1-Flash | |
| Alibaba | Qwen3.8-Maxqwen3.8-max | $2 | $6 | $0.25 | 1M | Replaced Qwen3.7-Max at $2.50/$7.50. International list price; the $0.25 cache rate comes from our launch coverage, the page only says “context caching discount”. | Confirmed | Alibaba Model Studio pricing | None yet | |
| Alibaba | Qwen3.8-Flashqwen3.8-flash | $0.15 | $0.47 | not listed | 1M | Qwen Cloud showed $0.16 when we looked on 30 Aug; Model Studio’s international price is $0.15. | Confirmed | Alibaba Model Studio pricing | Qwen3.8-Flash-Next | |
| Alibaba | Qwen3.8-Omni-Flashqwen3.8-omni-flash | $0.15 for any input type | $0.47 | $0.016 | 1M | Text output only. The speech variant, qwen3.8-omni-flash-realtime, now has its own row: $0.23 in, $0.93 for audio in, $0.70 text out, $1.87 audio out. | Confirmed | Alibaba Model Studio pricing | None yet | |
| Zhipu (Z.ai) | GLM-5.3glm-5.3 | $1.40 | $4.40 | $0.26 | 1M (128K out) | Thinking can’t be switched off; the lowest setting is low. Same price as GLM-5.2. Z.ai’s docs changelog dates the API listing 18 Aug. | Confirmed | Z.ai pricing Z.ai release notes | None yet | |
| Zhipu (Z.ai) | GLM-5.3-Flashglm-5.3-flash | $0.15 | $0.50 | $0.03 | 1M | Launch price was half ($0.075/$0.25) until 9 Sep 16:00 UTC. FlashX tier $0.37/$1.25. Weights are MIT. | Confirmed | Z.ai pricing | GLM-5.3-Flash launch price | |
| Moonshot | Kimi K3kimi-k3 | $3 | $15 | $0.30 | 1,048,576 | Cache writes cost $3 for five minutes and $6 for an hour, so the one-hour tier doubles the input rate. | Confirmed | Kimi API pricing | None yet | |
| Meta | Muse Spark 1.3muse-spark-1.3 | $1.25 | $4.25 | $0.15 | 1,048,576 | Same price as 1.1 and 1.2. The contributor tier is $0.10/$0.20 but lets Meta train on your prompts. No long-context premium, per Meta. | Confirmed | Meta Model API pricing Meta, Muse Spark 1.3 | None yet | |
| Xiaomi | MiMo-V2.6-Promimo-v2.6-pro | $0.435 | $0.87 | $0.0036 (router) | 1M | Xiaomi swapped retrained weights into the API on 25 Sep under the same name. No older snapshot to pin. | Reported | VentureBeat on MiMo-V2.6 Xiaomi MiMo blog | MiMo-V2.6 retrain | |
| Xiaomi | MiMo-V2.6-Flashmimo-v2.6-flash | $0.14 | $0.28 | $0.0028 (router) | 1M | Same silent swap on 25 Sep. Rates match DeepSeek’s pre-August Flash price to the cent. | Reported | VentureBeat on MiMo-V2.6 Xiaomi MiMo blog | MiMo-V2.6 retrain | |
| Sakana AI | Fugu Maxfugu-max | $2 | $6 | $0.25 | not stated | Flat at any length. Web search or fetch adds $0.007 a call. You pay for every sub-agent token. Not available in the EU or EEA. | Confirmed | Sakana Fugu pricing | Fugu Max and Ultra v2 | |
| Sakana AI | Fugu Ultra v2fugu-ultra-v2.0 | $5 | $30 | $0.50 | not stated | Past 272K context: $10/$45 (cache $1.00). The page doesn’t say if that covers the whole request. Not available in the EU or EEA. | Confirmed | Sakana Fugu pricing | Fugu Max and Ultra v2 | |
| Writer | Palmyra X6palmyra-x6 | $2 | $8 | none listed | 1M | Post-trained on GLM-5.2. Palmyra X4 and X5 are deprecated on 14 Dec 2026. | Confirmed | Writer pricing | Palmyra X6 |
The output price isn’t where the surprises are. Start with cache reads. They run from $0.003 on DeepSeek’s off-peak Flash to $1.00 on GPT-6 Astra and the original Fable 5, and they move inside a single vendor: Fable 5.1 reads cache at a quarter of Fable 5’s rate, while Grok 4.6 raised its read price by 67 percent and left the headline alone. If your agent re-reads a 150K prefix forty times a session, that column is your bill. Then the cliffs. Every GPT-5.6 and GPT-6 row reprices the whole request once the prompt passes 272K tokens, and every Grok row does it past 200K. Haiku 5.5 steps up at 100K. A 900K prompt on GPT-6 Astra costs $18 on input alone.
Then the clocks. I’d size the Flash models at $1.50 and $7.50 and GPT-5.6 Sol at $5 and $30, and treat today’s lower rates as a rebate. I might be wrong about Sol. OpenAI only promises “at least” 21 November, and it may well extend the offer. Our Fable 5.1 piece shows the same trap from the other side: the discount was real, and it sat entirely in one row.
Voice, video and document models
| Vendor | Model | Price | Price caveat | Released | Status | Source | Our write-up |
|---|---|---|---|---|---|---|---|
| OpenAI | gpt-realtime-2.1 and mini | Full: audio $32 in, $64 out; text $4 in, $24 out. Mini: audio $10 in, $20 out; text $0.60 in, $2.40 out. Per million tokens. | 128K context. The older gpt-realtime and gpt-4o-realtime shut down on 20 Jan 2027. | Confirmed | OpenAI pricing OpenAI deprecations | gpt-realtime-2.1 | |
| xAI | grok-voice-think-fast-2.0 (alias grok-voice-latest) | $0.08 per minute of audio. | The alias moved from 1.0 at $0.05 on 5 Aug. xAI’s model list now shows only 2.0. Speech to text: $0.10 an hour batch, $0.20 streaming. | Confirmed | xAI voice guide xAI, Voice Think Fast 2.0 | The grok-voice-latest alias | |
| Gemini 3.8 Live and Live Extended Thinking | Audio $3 in, $12 out per million tokens ($0.005 and $0.018 a minute). Text $0.75 in, $4.50 out. | One price for both. Extended Thinking only accepts non-blocking tools. No caching or structured output. | Confirmed | Gemini API pricing Gemini API changelog | None yet | ||
| Gemini 3.5 Transcribe and Transcribe Live | Recorded: $2.00 audio in, $12.00 text out per million tokens, about $0.005 a minute. Live: $3.50 and $21.00, about $0.009 a minute. | Generally available since 26 Aug per Google’s changelog. OpenAI’s gpt-4o-transcribe, the usual comparison, shuts down 26 Feb 2027. | Confirmed | Gemini API pricing Gemini API changelog | None yet | ||
| Gemini Omni 1.1 Flash (video) | About $0.10 per second at 720p ($17.50 per million video tokens). Draft 360p costs a third. | Ten seconds per generation, so 40 seconds is four billed generations. 1080p and 4K are upscales of a 720p render and cost more per second. | Confirmed | Gemini API pricing Google, Omni 1.1 Flash | Omni 1.1 Flash | ||
| Alibaba | Wan3.0 video (wan3.0-video) | List $0.05, $0.10 and $0.20 per second at 480p, 720p and 1080p. | Alibaba’s page labels these list prices with a “limited-time 30% off” and shows no end date or discounted figure. The faster prime tier is $0.068 to $0.28. | Confirmed | Alibaba Model Studio pricing | Wan 3.0 | |
| Cohere | Parse 5 (parse-v5.0) | $1.50 per 1,000 pages. | One page per request, 8,192-token window. The 79.2 ParseBench headline averages three of five dimensions. | Confirmed | Cohere, Parse | Cohere Parse 5 | |
| TypeSafe AI | Jev 1.13 | $42 per billion input tokens ($0.042 per million). Output is free. | Early access. Text in, probabilities out, never a sentence. 64K-token budget per request. | Confirmed | TypeSafe docs | Jev 1.13 |
Per-minute pricing hides length. xAI’s voice rate rose 60 percent when the alias moved, and calls would have to get 37.5 percent shorter to break even. Gemini’s Omni video bills four ten-second generations for a 40-second clip. Alibaba’s $0.15 audio rate belongs to the text-output model; the speech variant charges $0.93 for audio in. The unit on the sticker is rarely the unit you consume.
Dates that will change your bill
| Date | What happens | What it does to the bill | Status | Source |
|---|---|---|---|---|
| GPT-5.6 Sol’s $4/$20 promotional price is guaranteed “at least” through this date. | OpenAI hasn’t said what follows. The rate before 21 Aug was $5/$30, up 25% on input and 50% on output. | Floor, not an end date | OpenAI pricing OpenAI developer forum, 21 Aug | |
| Claude Sonnet 4.5 retires on the Claude API (deprecated 30 Sep). Anthropic recommends Sonnet 5.5. | Sonnet 4.5 costs $3/$15; Sonnet 5.5 is $2/$10 on a different tokenizer that counts about 30% more tokens for the same text. | Confirmed | Anthropic model deprecations Anthropic release notes | |
| o3, o3-pro and the gpt-5-2025-08-07 snapshot shut down. Replacement: gpt-5.6-sol. | o3 at $2/$8 becomes Sol at $4/$20 while the promo lasts, $5/$30 if it doesn’t. | Confirmed | OpenAI deprecations | |
| Writer deprecates Palmyra X4 and X5. Migration path: Palmyra X6. | X5 at $0.60/$6 moves to X6 at $2/$8, more than three times the input price. | Confirmed | Writer pricing | |
| Introductory rates end on Gemini 3.6, 3.7 and 3.8 Flash. New rate applies from 1 Jan 2027. | Input $0.75 to $1.50, output $3.75 to $7.50, cache read $0.075 to $0.15, batch doubles too. | Confirmed | Gemini API pricing | |
| gpt-realtime, gpt-4o-realtime, gpt-audio and the first realtime-mini shut down. | Replacements are gpt-realtime-2.1, gpt-audio-1.5 and gpt-realtime-2.1-mini, each with its own rate card. | Confirmed | OpenAI deprecations | |
| whisper-1 and the gpt-4o-transcribe family shut down. Successors: gpt-transcribe and gpt-live-transcribe. | Successors list at $0.0045 and $0.017 a minute, against $0.006 for gpt-4o-transcribe. | Confirmed | OpenAI deprecations OpenAI pricing | |
| No date given | Gemini 4 Argon’s introductory $2/$10 gives way to $4/$20. | Doubles both legs. Google gives no end date and no API date either. | No date yet | Google, Gemini 4 Argon |
| No date given | Wan3.0’s limited-time 30% discount ends. | The page prints list prices and the discount label, not the discounted figure or the end of it. | No date yet | Alibaba Model Studio pricing |
Seven of the 9 carry a day, and all 7 land before the end of February. Put 21 November in a calendar first: it isn’t a deadline so much as the first day OpenAI is free to change what Sol costs. The wider list of retirements and breaking changes lives in our retirements calendar. If a cliff pushes you toward running weights yourself, our open-weight models tracker lists what you’d host.
What changed since
Newest first. Each line is one change to a price, a version or a name, with the page that shows it. Where the vendor’s page has overtaken one of our older articles, the line says so.
| Date | What changed | Status | Source |
|---|---|---|---|
OpenAI opens the Decisions API in public beta (6 Oct), gpt-6-luna only: $0.10 per million input tokens, no output, cache-read or cache-write charge, and no caching. Luna’s own row re-read on the OpenAI pricing page and unchanged ($0.10/$0.50, $0.01 cached read, $0.125 cache write); other rows not re-checked this day. | Confirmed | OpenAI Decisions guide and our write-up | |
| OpenAI adds Ultrafast to GPT-6.1 Sol (8 Oct): $12 in, $60 out, $0.60 cached read, $15 cache write, 6x standard. Row re-read on the OpenAI pricing page; other rows not re-checked this day. | Confirmed | OpenAI pricing Ultrafast guide | |
| Anthropic launches Claude Haiku 5.5: $0.10 in and $0.50 out up to 100K prompt tokens, $0.50 and $2.50 above. 1M context. | Confirmed | Anthropic release notes Haiku 5.5 model page | |
| Anthropic cuts Sonnet 5.5 cache reads from $0.20 to $0.10 (0.05x base). Our 29 Sep write-up and Anthropic’s own model page still show $0.20. | Confirmed | Anthropic release notes | |
| OpenAI sets shutdown dates: gpt-5.3-codex, gpt-5.1 and gpt-5.4-nano on 1 Apr 2027, the older speech models on 6 Jan 2027. | Confirmed | OpenAI deprecations | |
| Google announces Gemini 4 Argon at an introductory $2/$10, later $4/$20, for trusted testers only. | Announced | Google, Gemini 4 Argon | |
| Anthropic deprecates Claude Sonnet 4.5; retirement on the Claude API is 30 Nov 2026. | Confirmed | Anthropic release notes Anthropic model deprecations | |
| OpenAI ships GPT-6.1 Sol at GPT-6 Sol’s $2/$10. Cached input falls from $0.20 to $0.10. | Confirmed | OpenAI pricing | |
| Claude Sonnet 5.5 arrives at Sonnet 5’s $2/$10. Thinking disabled now returns a 400; between_tools replaces it. | Confirmed | Anthropic release notes | |
| Xiaomi swaps retrained weights into the MiMo-V2.6 API under the same model names, with no older snapshot to pin. | Reported | Xiaomi MiMo blog | |
| OpenAI adds GPT-6 Sol ($2/$10) and GPT-6 Luna ($0.10/$0.50) beside Astra. | Confirmed | OpenAI pricing | |
| Anthropic ships Claude Opus 5.5 at $4/$20, down from Opus 5’s $5/$25, with cache reads at 0.05x base. | Confirmed | Anthropic release notes Anthropic pricing | |
| xAI ships Grok 4.7 at Grok 4.6’s $2/$6, same cache rate and same 200K step. | Confirmed | xAI, Grok 4.7 page | |
| Xiaomi launches MiMo-V2.6-Pro at $0.435/$0.87 and Flash at $0.14/$0.28. | Reported | VentureBeat on MiMo-V2.6 | |
| Alibaba lists Qwen3.8-Omni-Flash at $0.15 for text, image, video and audio input, against $11 for audio on Qwen3.5-Omni-Plus. | Confirmed | Alibaba Model Studio pricing | |
| Google makes Gemini 3.8 Live and Live Extended Thinking generally available at the old Native Audio rate, $3/$12 per million audio tokens. | Confirmed | Gemini API changelog Gemini API pricing | |
| DeepSeek reverses its V4 Pro retirement a day after announcing it: Pro stays after 14 Sep with billing unchanged. | Confirmed | DeepSeek change log | |
| Sakana adds Fugu Max at $2/$6 and replaces Fugu Ultra with Ultra v2 at the same $5/$30. | Confirmed | Sakana Fugu pricing | |
| DeepSeek releases V4.1-Flash as deepseek-flash at $0.15/$0.60 off-peak. V4 Flash is retired; its old names are routed to the new model at the new prices. | Confirmed | DeepSeek change log DeepSeek models and pricing | |
| OpenAI starts rolling out GPT-6 Astra at $10/$50, with $20/$75 past 272K input tokens. | Confirmed | OpenAI pricing | |
| Google releases Gemini 3.8 Flash at 3.7’s introductory $0.75/$3.75, with the same 31 Dec end date. | Confirmed | Gemini API changelog Gemini API pricing | |
| Meta releases Muse Spark 1.3 at the 1.1 and 1.2 price, $1.25/$4.25. | Confirmed | Meta, Muse Spark 1.3 Meta Model API pricing | |
| Claude Fable 5.1 keeps Fable 5’s $10/$50 but cuts cache reads from $1.00 to $0.25. Forced tool_choice now returns a 400. | Confirmed | Anthropic release notes | |
| Z.ai ships GLM-5.3-Flash at $0.075/$0.25, half off until 9 Sep 16:00 UTC. The list price is $0.15/$0.50. | Confirmed | Z.ai pricing Z.ai release notes | |
| OpenAI cuts GPT-5.6 Sol from $5/$30 to $4/$20 “for the next 3 months”. None of our articles covered the cut. | Confirmed | OpenAI developer forum, 21 Aug OpenAI pricing | |
| DeepSeek starts peak and off-peak billing at 16:00 UTC. Off-peak is half the peak rate. | Confirmed | DeepSeek change log | |
| Google releases Gemini 3.7 Flash at $0.75/$3.75 through 31 Dec, then $1.50/$7.50. Writer ships Palmyra X6 at $2/$8. | Confirmed | Gemini API changelog Writer pricing | |
| xAI ships Grok 4.6 at $2/$6, but the cache read rises from $0.30 to $0.50. | Confirmed | xAI models | |
| Anthropic makes Sonnet 5’s $2/$10 permanent; the $3/$15 planned for 1 Sep won’t apply. Our 1 and 2 July Sonnet 5 comparisons still quote the old plan. | Confirmed | Anthropic release notes | |
| grok-voice-latest moves to Grok Voice Think Fast 2.0: $0.05 to $0.08 a minute. | Confirmed | xAI, Voice Think Fast 2.0 xAI voice guide | |
| Alibaba lists Qwen3.8-Max at $2/$6, cheaper than Qwen3.7-Max at $2.50/$7.50. | Confirmed | Alibaba Model Studio pricing | |
| OpenAI cuts GPT-5.6 Luna by 80% ($1/$6 to $0.20/$1.20) and Terra by 20% ($2.50/$15 to $2/$12). Sol holds at $5/$30 for three more weeks. Priority processing is renamed Fast mode. | Confirmed | OpenAI, GPT-5.6 price-performance post OpenAI pricing | |
| Claude Opus 5 launches at Opus 4.8’s $5/$25 with thinking on by default. DeepSeek’s deepseek-chat and deepseek-reasoner names stop answering. | Confirmed | Anthropic pricing DeepSeek change log | |
| Google releases Gemini 3.6 Flash at $1.50/$7.50 (3.5 Flash was $1.50/$9) and 3.5 Flash-Lite at $0.30/$2.50. | Confirmed | Gemini API changelog Gemini API pricing | |
| Moonshot puts Kimi K3 live at $3/$15 with $0.30 cache reads. | Confirmed | Kimi API pricing | |
| GPT-5.6 reaches general availability as three tiers: Sol $5/$30, Terra $2.50/$15, Luna $1/$6. | Confirmed | OpenAI pricing GPT-5.6 Sol model page | |
| xAI releases Grok 4.5 at $2/$6. | Confirmed | xAI models The Decoder on Grok 4.5 | |
| Claude Fable 5 returns at $10/$50 after the US export-control suspension of 12 to 30 June. | Confirmed | Anthropic pricing | |
| Claude Sonnet 5 launches at an introductory $2/$10. | Confirmed | Anthropic release notes |
FAQ
Which price should I budget on when a model has an introductory rate?
The one that applies after the date. Every Gemini Flash model still on the intro rate goes from $0.75 and $3.75 to $1.50 and $7.50 on 1 January 2027, and cache reads and batch double with it. GPT-5.6 Sol is murkier: $4 and $20 holds “at least” through 21 November, and it cost $5 and $30 before 21 August. I’d budget $5 and $30. Gemini 4 Argon’s later $4 and $20 has no date at all, so size for it from day one.
Is Gemini 3.5 Pro out?
No. It isn’t on Google’s Gemini API pricing page or in the API changelog as of 7 October, and the last official word we found, in July, was that Google was still testing it with partners. The Pro-tier model Google does list is Gemini 3.1 Pro Preview, at $2 and $12 for prompts up to 200K tokens and $4 and $18 above that. Argon has a price but only trusted testers can call it. Any 3.5 Pro price you’ve seen is a leak.
Why does the cache-read column matter more than the output price?
On an agent that re-reads a long prefix every turn, most of the input is cached, so the cache rate sets the bill. It runs from $0.003 to $1.00 per million in this table, and it moves inside a vendor: Fable 5.1 reads at a quarter of Fable 5’s rate, Opus 5.5 at 5 percent of input, Grok 4.6 at 67 percent more than 4.5. Compare cache rates before you compare input prices.
Do these prices include batch, fast mode, data residency or peak hours?
No. They’re standard list prices at short context on the global endpoint. Fast modes multiply the bill: OpenAI’s Fast tier (renamed from Priority on 30 July) is 2x, Ultrafast on GPT-6 Astra and GPT-6.1 Sol is 6x, and Anthropic’s fast mode for Opus 5.5 is 2x at $8 and $40. Regional processing adds 10 percent on OpenAI’s newer models, and US-only inference adds 10 percent on Claude 4.6 and later. DeepSeek doubles its rates at peak hours. Meta is the odd one out: its page says there’s no long-context premium.
What happens to an alias like grok-voice-latest or deepseek-v4-flash?
It moves, and the invoice moves with it. grok-voice-latest has pointed at Grok Voice Think Fast 2.0 since 5 August: $0.08 a minute against $0.05 for 1.0. DeepSeek retired V4 Flash on 10 September but the name still answers, routing you to V4.1-Flash at V4.1 prices, and DeepSeek calls that “temporary”. Xiaomi swapped MiMo-V2.6’s weights under the same API name on 25 September. Pin explicit ids where the vendor lets you, and log the id that answered.
Sources
Every vendor page below was opened on 7 October 2026.
- Anthropic: Anthropic pricing, Anthropic models overview, Anthropic release notes, Anthropic model deprecations, Haiku 5.5 model page.
- OpenAI: OpenAI pricing, OpenAI deprecations, GPT-5.6 Sol model page, OpenAI developer forum, 21 Aug.
- Google: Gemini API pricing, Gemini API changelog, Google, Gemini 4 Argon, Google, Omni 1.1 Flash.
- xAI: xAI models, xAI, Grok 4.7 page, xAI voice guide, xAI, Voice Think Fast 2.0.
- DeepSeek: DeepSeek models and pricing, DeepSeek change log.
- Alibaba: Alibaba Model Studio pricing.
- Z.ai: Z.ai pricing, Z.ai release notes.
- Moonshot: Kimi API pricing.
- Meta: Meta Model API pricing, Meta, Muse Spark 1.3.
- Sakana, Writer, Cohere and TypeSafe: Sakana Fugu pricing, Writer pricing, Cohere, Parse, TypeSafe docs.
- Xiaomi (press and vendor blog): VentureBeat on MiMo-V2.6, Xiaomi MiMo blog.