AI credits vs tokens is the difference between the unit a platform sells you and the unit a model actually bills in: tokens are fragments of text the model processes, and credits are a platform's own currency that converts into those tokens at a rate it sets.
The confusion is not accidental. Nearly every AI product prices in something other than the underlying unit, and the translation is where the cost of your usage actually gets decided. At ChatFuse we price in credits and this post explains what sits behind them.
- Roughly 3 to 4 characters each
- Counted on input and output
- Priced differently per model
- Output usually costs far more
- A single unit across many models
- Converts into tokens at a set rate
- Lets you compare tasks, not models
- Expiry rules vary, so read them
What exactly is a token?
A token is a chunk of text a model reads or writes, usually 3 to 4 characters, so a token is often part of a word rather than a whole one. A page of prose runs somewhere around 500 to 800 tokens depending on the language and how unusual the vocabulary is.
Both directions count. The model is billed for everything you send it, which includes your prompt plus any conversation history and attached files, and separately for everything it writes back. Output typically costs several times more than input.
Why does a long conversation get expensive?
Because the whole conversation is resent every turn. A model has no memory between messages, so continuing a thread means sending the entire history again as input. Turn 20 in a conversation is not billed like turn 1; it carries all 19 previous exchanges with it.
That is also why attaching a large document to a chat and then asking 10 follow up questions costs far more than most people expect. The document is included in each of those 10 requests.
Two habits follow from this and both are easy. Start a new conversation when the subject changes rather than continuing one thread all day, because a thread carries its whole history forward whether or not any of it is still relevant. And ask the several questions you have about a document together rather than one at a time, since that sends the document once instead of repeatedly.
What actually drives your AI costs?
Four things, and the model choice is only one of them.
The spread between models is the thing worth internalising. Asking a frontier model for the capital of France costs many times what a small model charges for the identical correct answer, and most requests in normal use are closer to that than to hard reasoning.
This is the part that surprises teams reviewing their first month of usage. The expensive line is rarely the hard work. It is the volume of ordinary requests, every one of them billed at the rate of the most capable model available, because nothing was deciding otherwise.
Why do platforms use credits instead of tokens?
Because token pricing is unusable for a person deciding whether to send a message. Every model has a different rate, input and output differ, and the rates change. Nobody can price a question in their head.
Credits normalise that into one number so the platform absorbs the volatility and you get a predictable monthly figure. When a provider changes its rates, or a model is retired and replaced by one priced differently, that lands on the platform rather than on your invoice.
The tradeoff is a layer of abstraction between what you spend and what it costs, which is precisely why the conversion rate and the expiry rules deserve reading before you commit to a plan.
How does routing change what you pay?
It decides which price list applies to each request. If everything goes to the most capable model available, every trivial request is billed at frontier rates, and most requests are trivial.
The ChatFuse Orchestrator reads each prompt and routes it across more than 100 models from OpenAI, Anthropic, Google and Meta, so a quick factual question goes somewhere cheap and a hard reasoning task goes somewhere capable. The same logic drives the energy argument in energy efficient AI model routing, because compute and cost move together.
Do unused credits roll over?
At ChatFuse, subscription credits reset each billing cycle and do not roll over, while add on credit packs you buy separately never expire. The system spends the expiring ones first, so you are not accidentally burning the permanent balance while a monthly allowance goes unused.
That distinction is worth checking with any provider, because the two behave very differently at renewal and the difference is rarely on the pricing page.
The question to ask is what happens to an unused balance on the day a plan renews or is cancelled. Some providers reset it, some carry it, and some do one for the monthly allowance and the other for purchased packs. ChatFuse does the latter, which is the arrangement that penalises a quiet month least.
Frequently asked questions
What is the difference between AI credits and tokens?
Tokens are the unit models bill in, roughly 3 to 4 characters of text, counted separately for what you send and what the model writes. Credits are a platform's own unit that converts into tokens at a rate the platform sets, giving you one predictable number instead of a different rate per model.
How many tokens is a page of text?
Roughly 500 to 800 tokens for a page of ordinary English prose, though it varies with language and vocabulary. Technical text with unusual terms uses more tokens per word than plain writing.
Why does my AI usage cost more in long conversations?
Because the entire conversation history is resent with every message. Models hold no memory between turns, so a long thread means paying to resend everything that came before it each time you reply.
Does using a cheaper model give worse answers?
For simple requests, generally not in any way you would notice. A small model answering a factual question or reformatting text produces the same result as a frontier model at a fraction of the cost. The difference shows up on complex reasoning, which is why routing per task beats picking one model for everything.
How can I reduce AI costs without using it less?
Start fresh conversations rather than continuing long ones, avoid reattaching large files to every follow up, and use a platform that routes by task instead of sending everything to the most expensive model. See the pricing page for what a single subscription covers.
The useful mental model is that you are not buying answers, you are buying processed text in both directions, and the price of that text depends entirely on which model handled it. Once that is clear, most cost surprises stop being surprising, and the ones that remain are usually a long conversation nobody thought to restart.
Start free with ChatFuse and see what 5,000,000 credits actually covers before you decide.
Comments
Loading commentsโฆ