An AI API pricing comparison lines up what each provider charges to send tokens to a model and get tokens back, so you can see which model costs what before you build on it.

That comparison changed 4 times in one week. Anthropic cut a price on September 1. Google shipped a model on September 2 that already has a price rise scheduled. Meta shipped another the same day. OpenAI released GPT-6 Astra on September 3. Four moves, 4 days, 3 of the 4 biggest providers.

The headline numbers are easy to find and easy to misread. Both flagships that shipped this week list the exact same price and land very different bills. One of the cheapest models on the market costs the most of all once you read what you pay it with. This is the part a pricing comparison has to get right, because the sticker is not the bill.

We pay attention to this at ChatFuse because we route every prompt across more than 130 models from OpenAI, Anthropic, Google, Meta and others, and the price of each one feeds that decision. When a provider moves a number, it moves which model answers your next prompt. So at ChatFuse we keep the whole table current as ordinary maintenance, and this is what it looked like in September 2026.

AI API pricing comparison The dearest and cheapest output tokens this week were 13x apart.
Flagship output $50 GPT-6 Astra and Claude Fable 5.1, per 1M tokens
Flash output $3.75 Gemini 3.8 Flash, per 1M tokens through December 31
The spread 13x dearest to cheapest output token this week
Compiled by ChatFuse from vendor pricing pages, September 2026.

What does an AI API pricing comparison actually measure?

An AI API pricing comparison measures 4 separate prices per model, not one. You pay for input tokens, which is everything you send, including the prompt and any context. You pay for output tokens, which is everything the model generates. And on most providers you pay a cache write and a cache read for reusing context you have already sent. Output almost always costs several times what input costs, and cache reads are the cheapest of the 4. A comparison that quotes a single number per model is hiding at least 3 others.

If tokens themselves are new to you, we wrote a separate piece on how credits map to tokens. Here we are comparing the raw provider prices, which is the layer underneath.

How do the September 2026 flagship prices compare?

The list output prices split into 2 clear tiers. GPT-6 Astra and Claude Fable 5.1 both sit at $50 per 1M output tokens. Gemini 3.8 Flash and Meta Muse Spark 1.3 sit near $4. That is a spread of about 13x for a token of output, before any of the details that move the real cost.

Output price, per 1M tokens Two flagships at $50, two challengers near $4.
GPT-6 Astra $50
Claude Fable 5.1 $50
Meta Muse Spark 1.3 $4.25
Gemini 3.8 Flash $3.75
Compiled by ChatFuse from OpenAI, Anthropic, Google and Meta pricing, September 2026. Gemini and Muse standard tier.

Input follows the same shape. GPT-6 Astra and Claude Fable 5.1 charge $10 per 1M input tokens. Gemini 3.8 Flash charges $0.75. Meta's standard tier is $1.25. The flagships are built for the hardest work and priced for it. The cheaper models are built for volume.

Why do two models at the same price cost different amounts?

Because the sticker price hides the cache price and the context tiers. GPT-6 Astra and Claude Fable 5.1 both read $10 for input and $50 for output. But Astra charges $1 to read a cached token where Fable 5.1 charges $0.25, so a workload that reuses a lot of context pays Astra 4 times as much for the same cached read. And Astra has a second price tier for long prompts that Fable 5.1 does not.

Same sticker, different bill Two models at $10 and $50. The gap lives in the cache column.
GPT-6 Astra Dearer cache, doubles on long prompts
  • Input $10, output $50 per 1M
  • Cache read $1 per 1M
  • Long prompts: $20 input, $75 output
  • Cache write $12.50 per 1M
Claude Fable 5.1 Cheaper cache, one flat tier
  • Input $10, output $50 per 1M
  • Cache read $0.25 per 1M
  • No long context surcharge
  • Cache write $12.50 per 1M
Compiled by ChatFuse from the OpenAI and Anthropic pricing pages, September 2026.

Anthropic cut that Fable cache read from $1 to $0.25 on September 1, a 75% cut on one of the 4 prices while the 2 headline prices did not move at all. If you were comparing on sticker alone, you would have missed it.

What is the long context surcharge on GPT-6 Astra?

GPT-6 Astra has a 1,050,000 token context window, and past a threshold reported to sit near 272,000 input tokens in a single request, OpenAI bills the whole request at its long context tier. That is $20 for input and $75 for output, against $10 and $50 for short prompts. Cached reads double too, from $1 to $2. So the same model can cost you double on a prompt that is only different in length. The window is a headline feature. The second price tier attached to it is not.

Are AI prices going up or down?

Both, on schedules the providers publish in advance. Gemini 3.8 Flash costs $0.75 for input today, and Google's own pricing page already lists $1.50 from January 1, 2027, a scheduled doubling written down months ahead. In the same week Anthropic cut a Fable cache read by 75%. Prices are not drifting. They are being moved, up and down, on purpose.

One week of price moves Four changes in 4 days, from 3 of the 4 biggest providers.
1
September 1: Anthropic cuts a Fable cache read Claude Fable 5.1 cache reads drop to $0.25 from $1, a 75% cut, while $10 and $50 hold.
2
September 2: Google schedules a doubling Gemini 3.8 Flash ships at $0.75 input, with $1.50 already listed from January 1, 2027.
3
September 2: Meta prices a discount tier on your data Muse Spark 1.3 Contributor tier lands near $0.10 input, in exchange for training on your prompts.
4
September 3: OpenAI ships GPT-6 Astra at $10 and $50 With a second, long context tier that bills input and cache at 2x and output at 1.5x.
Compiled by ChatFuse from vendor pricing pages, September 2026. Meta tier pricing per public trackers.

We compared 2 of these families head to head when Fable 5 and GPT-5.6 launched, and the prices in that post are already out of date, which is the point. A model you priced a quarter ago is not the model you are paying for now, and some of them retire entirely on a date you did not set.

Is the cheapest AI model the cheapest bill?

Not always, and Meta Muse Spark 1.3 is the clearest example. Its standard tier is already cheap at $1.25 for input and $4.25 for output. Its Contributor tier is far cheaper, near $0.10 for input, roughly 21 times less than a flagship. The catch is in the name. The Contributor tier is discounted because Meta trains on the prompts and completions you send it. The cheapest token on the table is the one you pay for with your data.

We do not run your prompts through a tier like that. At ChatFuse your data is not training anyone's next model, and the cost saving comes from routing instead. Most prompts do not need a $50 model. A short reply, a rewrite, a lookup, a classification, all of it runs well on a model priced near $4, and the flagship only sees the prompts that actually need it. That is where a real bill gets smaller, and we wrote up the routing side of it in energy efficient model routing.

How does ChatFuse price across all these models?

ChatFuse puts more than 130 models behind one subscription and routes each prompt to a model that fits the task, so your cost tracks the work rather than the most expensive model on the menu. The ChatFuse Orchestrator reads each prompt and sends it to the model that can handle it for the least money, and when a provider changes a price we have already accounted for it. You are not reconciling 4 pricing pages every quarter. We are. If you want the frame for judging whether any of this is paying off, we set it out in measuring AI ROI.

Frequently asked questions

What is the cheapest AI API in 2026?

Among current models, Google Gemini 3.8 Flash is one of the cheapest at $0.75 for input and $3.75 for output per 1M tokens through December 31, 2026. Meta Muse Spark 1.3 is cheaper still on its Contributor tier at about $0.10 for input, but that discount comes from Meta training on your prompts, so the lowest price carries a data cost the others do not.

How much does GPT-6 Astra cost per token?

GPT-6 Astra costs $10 per 1M input tokens and $50 per 1M output tokens on short prompts, with cached reads at $1. On requests past a threshold reported near 272,000 input tokens it moves to a long context tier of $20 for input and $75 for output. Its context window is 1,050,000 tokens, with up to 128,000 tokens of output.

Is Claude Fable 5.1 more expensive than GPT-6 Astra?

On the sticker they match, both at $10 for input and $50 for output per 1M tokens. Fable 5.1 is cheaper on cached reads, at $0.25 against Astra's $1, and it has no long context surcharge. So for workloads that reuse context or send long prompts, Fable 5.1 usually costs less for the same headline price.

Why did my AI bill go up without changing anything?

Usually one of 2 reasons. Either your prompts crossed a context tier, like GPT-6 Astra's jump past roughly 272,000 tokens, or a scheduled price change took effect, like Gemini 3.8 Flash doubling on January 1, 2027. Providers publish both in advance, so a bill can rise on a date you did not act on.

Does using multiple AI models cost more?

Not with routing. ChatFuse includes more than 130 models in one subscription and sends each prompt to a model that suits it, so cheap tasks run on cheap models and only the hard prompts reach a flagship. Spreading across models usually lowers the bill rather than raising it, and it removes the risk of a single provider's price change. See pricing for what a subscription covers.

Every number in this comparison has a date on it, because every one of them will move. The comparison worth trusting is not a snapshot of who is cheapest today. It is a system that keeps up with the changes for you.

Start free with ChatFuse and let the routing decide which model, at which price, answers each prompt.

Back to Blog

Written by Nico

Share

Comments

Loading comments…

Secure signup continues in a new tab.