LLM API pricing: cache reads, cliffs and expiry dates, per model

You wired up grok-voice-latest because xAI’s own quickstart does. On 5 August that string began answering from a newer model at $0.08 a minute instead of $0.05, and not one line of your code changed. That’s the gap this page is for. It keeps the current API price of 37 language models, plus the voice and document models that bill by the minute or the page, next to the details a headline number leaves out: the cache-read rate, the context step, the discount with a date on it, the alias that moves. It’s for whoever picks the model string and then has to explain the invoice.

Last verified: 7 October 2026. Every price was read on the vendor’s own page that day, apart from the two rows marked Reported.

How to read this page

Each row comes from the vendor’s pricing or model page, opened on 7 October. Our older articles supplied the questions. They didn’t supply the answers: when one disagrees with the vendor, the vendor wins, and the log says so. Only the two Xiaomi rows lean on press and router listings, because the platform page wouldn’t load for us.

Prices are US dollars per million tokens at the standard tier: short context, global endpoint, no batch discount, no fast-mode premium. Batch halves the bill at Anthropic and OpenAI, and Google matches it; xAI’s Grok 4.7 page says batch isn’t supported. Fast modes and regional processing add to it. None of that is in the numbers, and the FAQ covers it.

Status is one of three things. Confirmed means the vendor’s page says it today. Announced means there’s a price but no model you can call yet. Reported means we couldn’t check the source ourselves.

We left out open-weight models that only carry router prices (Laguna, Hy4, Nemotron and Inkling), because a router’s figure changes with whichever provider answers. Read the Price caveat column before the price columns. That’s where the expiry dates and the context cliffs live.

The price table

37 language models come first, then the voice and document models that don’t bill by the token, then the dates that will move your bill. Rows are grouped by vendor, newest model first. Sort by the caveat, not by the price.

Language models, USD per million tokens

Current API prices for 37 language models, USD per million tokens, read on 7 October 2026
VendorModelInput $/M tokensOutput $/M tokensCache read $/MContextReleasedPrice caveatStatusSourceOur write-up
AnthropicClaude Fable 5.1
claude-fable-5-1
$10$50$0.251M (128K out)Cache reads at 0.025x base, a quarter of Fable 5’s rate. Thinking always on, so the effort setting sets the real bill. Batch $5/$25.ConfirmedAnthropic pricing
Anthropic models overview
Fable 5.1 cache and tool use
AnthropicClaude Opus 5.5
claude-opus-5-5
$4$20$0.201M (128K out)Thinking can’t be disabled. Default effort is medium, and Anthropic says it thinks more per turn than Opus 5 at the same effort. Fast mode (preview) $8/$40.ConfirmedAnthropic pricing
Anthropic release notes
Opus 5.5 breaking changes
AnthropicClaude Sonnet 5.5
claude-sonnet-5-5
$2$10$0.101M (128K out)Cache read cut from $0.20 on 7 Oct. Anthropic’s model page and price table still print $0.20; the release note and the caching section say $0.10.ConfirmedAnthropic release notes
Anthropic pricing
None yet
AnthropicClaude Haiku 5.5
claude-haiku-5-5
$0.10 up to 100K prompt, $0.50 above$0.50 up to 100K, $2.50 above$0.01 up to 100K, $0.05 above1M (128K out)The only current Claude model with a prompt-length step: everything reprices past 100K tokens. Launched the day we checked.ConfirmedHaiku 5.5 model page
Anthropic release notes
The 100K cliff
AnthropicClaude Opus 5
claude-opus-5
$5$25$0.501M (128K out)Legacy, retirement not before 24 Jul 2027. Thinking on by default. Fast mode $10/$50.Confirmed, legacyAnthropic pricing
Anthropic model deprecations
Opus 5 launch
AnthropicClaude Sonnet 5
claude-sonnet-5
$2$10$0.201M (128K out)Launch price became the standard price on 10 Aug. The planned $3/$15 from 1 Sep never applied.Confirmed, legacyAnthropic release notes
Anthropic pricing
None yet
AnthropicClaude Fable 5
claude-fable-5
$10$50$1.001M (128K out)Same base price as 5.1, four times the cache-read price. Retirement not before 9 Jun 2027.Confirmed, legacyAnthropic pricing
Anthropic model deprecations
Fable 5 effort levels
OpenAIGPT-6 Astra
gpt-6-astra
$10$50$1.001.05M (128K out)Past 272K input tokens the whole request bills $20/$75. Fast mode 2x; Ultrafast tier $60/$300. First access went to vetted enterprises.ConfirmedOpenAI pricingGPT-6 Astra and the 272K cliff
OpenAIGPT-6.1 Sol
gpt-6.1-sol
$2$10$0.101.05M (128K out)Past 272K input: $4/$15 on the whole request. Rejects reasoning effort none and minimal. Fast mode $4/$20. Ultrafast tier (from 8 Oct) $12/$60, 6x standard; long-context billing not documented.ConfirmedOpenAI pricingGPT-6.1 Sol cache cut
OpenAIGPT-6 Sol
gpt-6-sol
$2$10$0.201.05M (128K out)Same rates as 6.1 Sol except cached input costs double. Past 272K: $4/$15. There is no GPT-6 Terra.ConfirmedOpenAI pricingGPT-6 Astra and the 272K cliff
OpenAIGPT-6 Luna
gpt-6-luna
$0.10$0.50$0.011.05M (128K out)Past 272K input: $0.20/$0.75 on the whole request. Decisions API (beta): $0.10 in, no output charge, no caching.ConfirmedOpenAI pricingGPT-6 Astra and the 272K cliff
OpenAIGPT-5.6 Sol
gpt-5.6-sol
$4$20$0.401.05M (128K out)Promo price: was $5/$30 until 21 Aug, promised “at least” through 21 Nov 2026. Past 272K: $8/$30. The ChatGPT version of Sol is tuned differently from the API one.Confirmed, promoOpenAI pricing
GPT-5.6 Sol model page
Sol in ChatGPT vs the API
OpenAIGPT-5.6 Terra
gpt-5.6-terra
$2$12$0.201.05M (128K out)Launched at $2.50/$15, cut 20% on 30 Jul. Past 272K: $4/$18.ConfirmedOpenAI pricingNone yet
OpenAIGPT-5.6 Luna
gpt-5.6-luna
$0.20$1.20$0.021.05M (128K out)Launched at $1/$6, cut 80% on 30 Jul. Past 272K: $0.40/$1.80.ConfirmedOpenAI pricingNone yet
OpenAIo3
o3-2025-04-16
$2$8$0.50200K (100K out)Shuts down 11 Dec 2026, along with o3-pro. OpenAI names gpt-5.6-sol as the replacement, at double the input and 2.5x the output price.Confirmed, retiringOpenAI pricing
OpenAI deprecations
o3 shutdown dates
GoogleGemini 3.8 Flash
gemini-3.8-flash
$0.75$3.75$0.0751M (65K out)Intro rate through 31 Dec 2026, then $1.50/$7.50 (cache $0.15) from 1 Jan 2027. Spends more tokens on purpose at high effort. Thinking level minimal removed.Confirmed, intro priceGemini API pricing
Gemini API changelog
Gemini 3.8 Flash and 1 January
GoogleGemini 3.7 Flash
gemini-3.7-flash
$0.75$3.75$0.0751M (65K out)Same intro rate and the same 31 Dec end date as 3.8. Google says it stays supported for efficiency-first work.Confirmed, intro priceGemini API pricingNone yet
GoogleGemini 3.6 Flash
gemini-3.6-flash
$0.75$3.75$0.0751M (65K out)Launched at $1.50/$7.50. The pricing page now shows the same intro rate as 3.7 and 3.8; we couldn’t date the change.Confirmed, intro priceGemini API pricingNone yet
GoogleGemini 3.5 Flash-Lite
gemini-3.5-flash-lite
$0.30$2.50$0.031M (65K out)No intro discount listed. Cache storage adds $1.00 per million tokens per hour.ConfirmedGemini API pricingNone yet
GoogleGemini 4 Argon
not listed yet
$2$1095% off input (about $0.10)1M output limitIntro price, then $4/$20 with no end date given. Not on the Gemini API pricing page or changelog on 7 Oct. Trusted cyber defenders only for now.AnnouncedGoogle, Gemini 4 Argon
Gemini API pricing
Gemini 4 Argon pricing
xAIGrok 4.7
grok-4.7
$2$6$0.50500KPast 200K the whole request bills $4/$12 (cache $1.00). Reasoning can’t be turned off. The launch table ran 4.7 at xhigh against 4.6 at high.ConfirmedxAI, Grok 4.7 page
xAI models
Grok 4.7 xhigh vs high
xAIGrok 4.6
grok-4.6
$2$6$0.50500KCache read was $0.30 on 4.5. Same 200K step as 4.7.ConfirmedxAI modelsGrok 4.6 and the 200K cliff
xAIGrok 4.5
grok-4.5
$2$6$0.30500KCheapest cache read of the three. Past 200K: $4/$12 (cache $0.60).ConfirmedxAI modelsNone yet
DeepSeekDeepSeek-V4.1-Flash
deepseek-flash
$0.15 off-peak, $0.30 peak$0.60 off-peak, $1.20 peak$0.003 off-peak, $0.006 peak1M (384K out)Peak is 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, at double. The retired name deepseek-v4-flash is routed here “temporarily”, at these prices.ConfirmedDeepSeek models and pricing
DeepSeek change log
DeepSeek V4.1-Flash
DeepSeekDeepSeek-V4-Pro-0813
deepseek-v4-pro
$0.66 off-peak, $1.32 peak$1.98 off-peak, $3.96 peak$0.022 off-peak, $0.044 peak1M (384K out)Same peak rule. A retirement notice for 14 Sep was reversed on 11 Sep, with “further notice” promised. Text only; 500 concurrent requests against 2,500 on Flash.ConfirmedDeepSeek models and pricing
DeepSeek change log
DeepSeek V4.1-Flash
AlibabaQwen3.8-Max
qwen3.8-max
$2$6$0.251MReplaced Qwen3.7-Max at $2.50/$7.50. International list price; the $0.25 cache rate comes from our launch coverage, the page only says “context caching discount”.ConfirmedAlibaba Model Studio pricingNone yet
AlibabaQwen3.8-Flash
qwen3.8-flash
$0.15$0.47not listed1MQwen Cloud showed $0.16 when we looked on 30 Aug; Model Studio’s international price is $0.15.ConfirmedAlibaba Model Studio pricingQwen3.8-Flash-Next
AlibabaQwen3.8-Omni-Flash
qwen3.8-omni-flash
$0.15 for any input type$0.47$0.0161MText output only. The speech variant, qwen3.8-omni-flash-realtime, now has its own row: $0.23 in, $0.93 for audio in, $0.70 text out, $1.87 audio out.ConfirmedAlibaba Model Studio pricingNone yet
Zhipu (Z.ai)GLM-5.3
glm-5.3
$1.40$4.40$0.261M (128K out)Thinking can’t be switched off; the lowest setting is low. Same price as GLM-5.2. Z.ai’s docs changelog dates the API listing 18 Aug.ConfirmedZ.ai pricing
Z.ai release notes
None yet
Zhipu (Z.ai)GLM-5.3-Flash
glm-5.3-flash
$0.15$0.50$0.031MLaunch price was half ($0.075/$0.25) until 9 Sep 16:00 UTC. FlashX tier $0.37/$1.25. Weights are MIT.ConfirmedZ.ai pricingGLM-5.3-Flash launch price
MoonshotKimi K3
kimi-k3
$3$15$0.301,048,576Cache writes cost $3 for five minutes and $6 for an hour, so the one-hour tier doubles the input rate.ConfirmedKimi API pricingNone yet
MetaMuse Spark 1.3
muse-spark-1.3
$1.25$4.25$0.151,048,576Same price as 1.1 and 1.2. The contributor tier is $0.10/$0.20 but lets Meta train on your prompts. No long-context premium, per Meta.ConfirmedMeta Model API pricing
Meta, Muse Spark 1.3
None yet
XiaomiMiMo-V2.6-Pro
mimo-v2.6-pro
$0.435$0.87$0.0036 (router)1MXiaomi swapped retrained weights into the API on 25 Sep under the same name. No older snapshot to pin.ReportedVentureBeat on MiMo-V2.6
Xiaomi MiMo blog
MiMo-V2.6 retrain
XiaomiMiMo-V2.6-Flash
mimo-v2.6-flash
$0.14$0.28$0.0028 (router)1MSame silent swap on 25 Sep. Rates match DeepSeek’s pre-August Flash price to the cent.ReportedVentureBeat on MiMo-V2.6
Xiaomi MiMo blog
MiMo-V2.6 retrain
Sakana AIFugu Max
fugu-max
$2$6$0.25not statedFlat at any length. Web search or fetch adds $0.007 a call. You pay for every sub-agent token. Not available in the EU or EEA.ConfirmedSakana Fugu pricingFugu Max and Ultra v2
Sakana AIFugu Ultra v2
fugu-ultra-v2.0
$5$30$0.50not statedPast 272K context: $10/$45 (cache $1.00). The page doesn’t say if that covers the whole request. Not available in the EU or EEA.ConfirmedSakana Fugu pricingFugu Max and Ultra v2
WriterPalmyra X6
palmyra-x6
$2$8none listed1MPost-trained on GLM-5.2. Palmyra X4 and X5 are deprecated on 14 Dec 2026.ConfirmedWriter pricingPalmyra X6

The output price isn’t where the surprises are. Start with cache reads. They run from $0.003 on DeepSeek’s off-peak Flash to $1.00 on GPT-6 Astra and the original Fable 5, and they move inside a single vendor: Fable 5.1 reads cache at a quarter of Fable 5’s rate, while Grok 4.6 raised its read price by 67 percent and left the headline alone. If your agent re-reads a 150K prefix forty times a session, that column is your bill. Then the cliffs. Every GPT-5.6 and GPT-6 row reprices the whole request once the prompt passes 272K tokens, and every Grok row does it past 200K. Haiku 5.5 steps up at 100K. A 900K prompt on GPT-6 Astra costs $18 on input alone.

Then the clocks. I’d size the Flash models at $1.50 and $7.50 and GPT-5.6 Sol at $5 and $30, and treat today’s lower rates as a rebate. I might be wrong about Sol. OpenAI only promises “at least” 21 November, and it may well extend the offer. Our Fable 5.1 piece shows the same trap from the other side: the discount was real, and it sat entirely in one row.

Voice, video and document models

Voice, video and document models billed by the minute, the second or the page, read on 7 October 2026
VendorModelPricePrice caveatReleasedStatusSourceOur write-up
OpenAIgpt-realtime-2.1 and miniFull: audio $32 in, $64 out; text $4 in, $24 out. Mini: audio $10 in, $20 out; text $0.60 in, $2.40 out. Per million tokens.128K context. The older gpt-realtime and gpt-4o-realtime shut down on 20 Jan 2027.ConfirmedOpenAI pricing
OpenAI deprecations
gpt-realtime-2.1
xAIgrok-voice-think-fast-2.0 (alias grok-voice-latest)$0.08 per minute of audio.The alias moved from 1.0 at $0.05 on 5 Aug. xAI’s model list now shows only 2.0. Speech to text: $0.10 an hour batch, $0.20 streaming.ConfirmedxAI voice guide
xAI, Voice Think Fast 2.0
The grok-voice-latest alias
GoogleGemini 3.8 Live and Live Extended ThinkingAudio $3 in, $12 out per million tokens ($0.005 and $0.018 a minute). Text $0.75 in, $4.50 out.One price for both. Extended Thinking only accepts non-blocking tools. No caching or structured output.ConfirmedGemini API pricing
Gemini API changelog
None yet
GoogleGemini 3.5 Transcribe and Transcribe LiveRecorded: $2.00 audio in, $12.00 text out per million tokens, about $0.005 a minute. Live: $3.50 and $21.00, about $0.009 a minute.Generally available since 26 Aug per Google’s changelog. OpenAI’s gpt-4o-transcribe, the usual comparison, shuts down 26 Feb 2027.ConfirmedGemini API pricing
Gemini API changelog
None yet
GoogleGemini Omni 1.1 Flash (video)About $0.10 per second at 720p ($17.50 per million video tokens). Draft 360p costs a third.Ten seconds per generation, so 40 seconds is four billed generations. 1080p and 4K are upscales of a 720p render and cost more per second.ConfirmedGemini API pricing
Google, Omni 1.1 Flash
Omni 1.1 Flash
AlibabaWan3.0 video (wan3.0-video)List $0.05, $0.10 and $0.20 per second at 480p, 720p and 1080p.Alibaba’s page labels these list prices with a “limited-time 30% off” and shows no end date or discounted figure. The faster prime tier is $0.068 to $0.28.ConfirmedAlibaba Model Studio pricingWan 3.0
CohereParse 5 (parse-v5.0)$1.50 per 1,000 pages.One page per request, 8,192-token window. The 79.2 ParseBench headline averages three of five dimensions.ConfirmedCohere, ParseCohere Parse 5
TypeSafe AIJev 1.13$42 per billion input tokens ($0.042 per million). Output is free.Early access. Text in, probabilities out, never a sentence. 64K-token budget per request.ConfirmedTypeSafe docsJev 1.13

Per-minute pricing hides length. xAI’s voice rate rose 60 percent when the alias moved, and calls would have to get 37.5 percent shorter to break even. Gemini’s Omni video bills four ten-second generations for a 40-second clip. Alibaba’s $0.15 audio rate belongs to the text-output model; the speech variant charges $0.93 for audio in. The unit on the sticker is rarely the unit you consume.

Dates that will change your bill

Dated changes that move an API bill, nearest first, undated items last
DateWhat happensWhat it does to the billStatusSource
GPT-5.6 Sol’s $4/$20 promotional price is guaranteed “at least” through this date.OpenAI hasn’t said what follows. The rate before 21 Aug was $5/$30, up 25% on input and 50% on output.Floor, not an end dateOpenAI pricing
OpenAI developer forum, 21 Aug
Claude Sonnet 4.5 retires on the Claude API (deprecated 30 Sep). Anthropic recommends Sonnet 5.5.Sonnet 4.5 costs $3/$15; Sonnet 5.5 is $2/$10 on a different tokenizer that counts about 30% more tokens for the same text.ConfirmedAnthropic model deprecations
Anthropic release notes
o3, o3-pro and the gpt-5-2025-08-07 snapshot shut down. Replacement: gpt-5.6-sol.o3 at $2/$8 becomes Sol at $4/$20 while the promo lasts, $5/$30 if it doesn’t.ConfirmedOpenAI deprecations
Writer deprecates Palmyra X4 and X5. Migration path: Palmyra X6.X5 at $0.60/$6 moves to X6 at $2/$8, more than three times the input price.ConfirmedWriter pricing
Introductory rates end on Gemini 3.6, 3.7 and 3.8 Flash. New rate applies from 1 Jan 2027.Input $0.75 to $1.50, output $3.75 to $7.50, cache read $0.075 to $0.15, batch doubles too.ConfirmedGemini API pricing
gpt-realtime, gpt-4o-realtime, gpt-audio and the first realtime-mini shut down.Replacements are gpt-realtime-2.1, gpt-audio-1.5 and gpt-realtime-2.1-mini, each with its own rate card.ConfirmedOpenAI deprecations
whisper-1 and the gpt-4o-transcribe family shut down. Successors: gpt-transcribe and gpt-live-transcribe.Successors list at $0.0045 and $0.017 a minute, against $0.006 for gpt-4o-transcribe.ConfirmedOpenAI deprecations
OpenAI pricing
No date givenGemini 4 Argon’s introductory $2/$10 gives way to $4/$20.Doubles both legs. Google gives no end date and no API date either.No date yetGoogle, Gemini 4 Argon
No date givenWan3.0’s limited-time 30% discount ends.The page prints list prices and the discount label, not the discounted figure or the end of it.No date yetAlibaba Model Studio pricing

Seven of the 9 carry a day, and all 7 land before the end of February. Put 21 November in a calendar first: it isn’t a deadline so much as the first day OpenAI is free to change what Sol costs. The wider list of retirements and breaking changes lives in our retirements calendar. If a cliff pushes you toward running weights yourself, our open-weight models tracker lists what you’d host.

What changed since

Newest first. Each line is one change to a price, a version or a name, with the page that shows it. Where the vendor’s page has overtaken one of our older articles, the line says so.

Price and version changes since 30 June 2026, newest first
DateWhat changedStatusSource
OpenAI opens the Decisions API in public beta (6 Oct), gpt-6-luna only: $0.10 per million input tokens, no output, cache-read or cache-write charge, and no caching. Luna’s own row re-read on the OpenAI pricing page and unchanged ($0.10/$0.50, $0.01 cached read, $0.125 cache write); other rows not re-checked this day.ConfirmedOpenAI Decisions guide and our write-up
OpenAI adds Ultrafast to GPT-6.1 Sol (8 Oct): $12 in, $60 out, $0.60 cached read, $15 cache write, 6x standard. Row re-read on the OpenAI pricing page; other rows not re-checked this day.ConfirmedOpenAI pricing
Ultrafast guide
Anthropic launches Claude Haiku 5.5: $0.10 in and $0.50 out up to 100K prompt tokens, $0.50 and $2.50 above. 1M context.ConfirmedAnthropic release notes
Haiku 5.5 model page
Anthropic cuts Sonnet 5.5 cache reads from $0.20 to $0.10 (0.05x base). Our 29 Sep write-up and Anthropic’s own model page still show $0.20.ConfirmedAnthropic release notes
OpenAI sets shutdown dates: gpt-5.3-codex, gpt-5.1 and gpt-5.4-nano on 1 Apr 2027, the older speech models on 6 Jan 2027.ConfirmedOpenAI deprecations
Google announces Gemini 4 Argon at an introductory $2/$10, later $4/$20, for trusted testers only.AnnouncedGoogle, Gemini 4 Argon
Anthropic deprecates Claude Sonnet 4.5; retirement on the Claude API is 30 Nov 2026.ConfirmedAnthropic release notes
Anthropic model deprecations
OpenAI ships GPT-6.1 Sol at GPT-6 Sol’s $2/$10. Cached input falls from $0.20 to $0.10.ConfirmedOpenAI pricing
Claude Sonnet 5.5 arrives at Sonnet 5’s $2/$10. Thinking disabled now returns a 400; between_tools replaces it.ConfirmedAnthropic release notes
Xiaomi swaps retrained weights into the MiMo-V2.6 API under the same model names, with no older snapshot to pin.ReportedXiaomi MiMo blog
OpenAI adds GPT-6 Sol ($2/$10) and GPT-6 Luna ($0.10/$0.50) beside Astra.ConfirmedOpenAI pricing
Anthropic ships Claude Opus 5.5 at $4/$20, down from Opus 5’s $5/$25, with cache reads at 0.05x base.ConfirmedAnthropic release notes
Anthropic pricing
xAI ships Grok 4.7 at Grok 4.6’s $2/$6, same cache rate and same 200K step.ConfirmedxAI, Grok 4.7 page
Xiaomi launches MiMo-V2.6-Pro at $0.435/$0.87 and Flash at $0.14/$0.28.ReportedVentureBeat on MiMo-V2.6
Alibaba lists Qwen3.8-Omni-Flash at $0.15 for text, image, video and audio input, against $11 for audio on Qwen3.5-Omni-Plus.ConfirmedAlibaba Model Studio pricing
Google makes Gemini 3.8 Live and Live Extended Thinking generally available at the old Native Audio rate, $3/$12 per million audio tokens.ConfirmedGemini API changelog
Gemini API pricing
DeepSeek reverses its V4 Pro retirement a day after announcing it: Pro stays after 14 Sep with billing unchanged.ConfirmedDeepSeek change log
Sakana adds Fugu Max at $2/$6 and replaces Fugu Ultra with Ultra v2 at the same $5/$30.ConfirmedSakana Fugu pricing
DeepSeek releases V4.1-Flash as deepseek-flash at $0.15/$0.60 off-peak. V4 Flash is retired; its old names are routed to the new model at the new prices.ConfirmedDeepSeek change log
DeepSeek models and pricing
OpenAI starts rolling out GPT-6 Astra at $10/$50, with $20/$75 past 272K input tokens.ConfirmedOpenAI pricing
Google releases Gemini 3.8 Flash at 3.7’s introductory $0.75/$3.75, with the same 31 Dec end date.ConfirmedGemini API changelog
Gemini API pricing
Meta releases Muse Spark 1.3 at the 1.1 and 1.2 price, $1.25/$4.25.ConfirmedMeta, Muse Spark 1.3
Meta Model API pricing
Claude Fable 5.1 keeps Fable 5’s $10/$50 but cuts cache reads from $1.00 to $0.25. Forced tool_choice now returns a 400.ConfirmedAnthropic release notes
Z.ai ships GLM-5.3-Flash at $0.075/$0.25, half off until 9 Sep 16:00 UTC. The list price is $0.15/$0.50.ConfirmedZ.ai pricing
Z.ai release notes
OpenAI cuts GPT-5.6 Sol from $5/$30 to $4/$20 “for the next 3 months”. None of our articles covered the cut.ConfirmedOpenAI developer forum, 21 Aug
OpenAI pricing
DeepSeek starts peak and off-peak billing at 16:00 UTC. Off-peak is half the peak rate.ConfirmedDeepSeek change log
Google releases Gemini 3.7 Flash at $0.75/$3.75 through 31 Dec, then $1.50/$7.50. Writer ships Palmyra X6 at $2/$8.ConfirmedGemini API changelog
Writer pricing
xAI ships Grok 4.6 at $2/$6, but the cache read rises from $0.30 to $0.50.ConfirmedxAI models
Anthropic makes Sonnet 5’s $2/$10 permanent; the $3/$15 planned for 1 Sep won’t apply. Our 1 and 2 July Sonnet 5 comparisons still quote the old plan.ConfirmedAnthropic release notes
grok-voice-latest moves to Grok Voice Think Fast 2.0: $0.05 to $0.08 a minute.ConfirmedxAI, Voice Think Fast 2.0
xAI voice guide
Alibaba lists Qwen3.8-Max at $2/$6, cheaper than Qwen3.7-Max at $2.50/$7.50.ConfirmedAlibaba Model Studio pricing
OpenAI cuts GPT-5.6 Luna by 80% ($1/$6 to $0.20/$1.20) and Terra by 20% ($2.50/$15 to $2/$12). Sol holds at $5/$30 for three more weeks. Priority processing is renamed Fast mode.ConfirmedOpenAI, GPT-5.6 price-performance post
OpenAI pricing
Claude Opus 5 launches at Opus 4.8’s $5/$25 with thinking on by default. DeepSeek’s deepseek-chat and deepseek-reasoner names stop answering.ConfirmedAnthropic pricing
DeepSeek change log
Google releases Gemini 3.6 Flash at $1.50/$7.50 (3.5 Flash was $1.50/$9) and 3.5 Flash-Lite at $0.30/$2.50.ConfirmedGemini API changelog
Gemini API pricing
Moonshot puts Kimi K3 live at $3/$15 with $0.30 cache reads.ConfirmedKimi API pricing
GPT-5.6 reaches general availability as three tiers: Sol $5/$30, Terra $2.50/$15, Luna $1/$6.ConfirmedOpenAI pricing
GPT-5.6 Sol model page
xAI releases Grok 4.5 at $2/$6.ConfirmedxAI models
The Decoder on Grok 4.5
Claude Fable 5 returns at $10/$50 after the US export-control suspension of 12 to 30 June.ConfirmedAnthropic pricing
Claude Sonnet 5 launches at an introductory $2/$10.ConfirmedAnthropic release notes

FAQ

Which price should I budget on when a model has an introductory rate?

The one that applies after the date. Every Gemini Flash model still on the intro rate goes from $0.75 and $3.75 to $1.50 and $7.50 on 1 January 2027, and cache reads and batch double with it. GPT-5.6 Sol is murkier: $4 and $20 holds “at least” through 21 November, and it cost $5 and $30 before 21 August. I’d budget $5 and $30. Gemini 4 Argon’s later $4 and $20 has no date at all, so size for it from day one.

Is Gemini 3.5 Pro out?

No. It isn’t on Google’s Gemini API pricing page or in the API changelog as of 7 October, and the last official word we found, in July, was that Google was still testing it with partners. The Pro-tier model Google does list is Gemini 3.1 Pro Preview, at $2 and $12 for prompts up to 200K tokens and $4 and $18 above that. Argon has a price but only trusted testers can call it. Any 3.5 Pro price you’ve seen is a leak.

Why does the cache-read column matter more than the output price?

On an agent that re-reads a long prefix every turn, most of the input is cached, so the cache rate sets the bill. It runs from $0.003 to $1.00 per million in this table, and it moves inside a vendor: Fable 5.1 reads at a quarter of Fable 5’s rate, Opus 5.5 at 5 percent of input, Grok 4.6 at 67 percent more than 4.5. Compare cache rates before you compare input prices.

Do these prices include batch, fast mode, data residency or peak hours?

No. They’re standard list prices at short context on the global endpoint. Fast modes multiply the bill: OpenAI’s Fast tier (renamed from Priority on 30 July) is 2x, Ultrafast on GPT-6 Astra and GPT-6.1 Sol is 6x, and Anthropic’s fast mode for Opus 5.5 is 2x at $8 and $40. Regional processing adds 10 percent on OpenAI’s newer models, and US-only inference adds 10 percent on Claude 4.6 and later. DeepSeek doubles its rates at peak hours. Meta is the odd one out: its page says there’s no long-context premium.

What happens to an alias like grok-voice-latest or deepseek-v4-flash?

It moves, and the invoice moves with it. grok-voice-latest has pointed at Grok Voice Think Fast 2.0 since 5 August: $0.08 a minute against $0.05 for 1.0. DeepSeek retired V4 Flash on 10 September but the name still answers, routing you to V4.1-Flash at V4.1 prices, and DeepSeek calls that “temporary”. Xiaomi swapped MiMo-V2.6’s weights under the same API name on 25 September. Pin explicit ids where the vendor lets you, and log the id that answered.

Sources

Every vendor page below was opened on 7 October 2026.