Cost your workload across every live model
Every live model, ranked for your workload
| # | Model | Context | $ / 1M in | $ / 1M out | List / month | Your cost / month | Saved |
|---|---|---|---|---|---|---|---|
| Loading live prices... | |||||||
What this tool does
Almost every LLM cost calculator on the internet multiplies your token count by a list price. That is the number on the pricing page, and it is very close to the number nobody pays. A production workload sits on top of four adjustments, and each one moves the bill in a direction you cannot predict from the sticker alone.
The first is prompt caching. When the same prefix is sent on every call, providers charge a cache read at a small fraction of the input rate, commonly a tenth and on some models a hundredth. The second is cache misses, and this is the trap. A single current timestamp or request ID near the top of the prompt invalidates the whole prefix, so an application that looks heavily cached on paper bills almost entirely at the full input rate. The third is reasoning tokens, which bill at the output rate and are easy to leave out of a spreadsheet. The fourth is the long-context tier, where a provider charges more per token once a single request crosses a threshold, often 200,000 or 272,000 tokens, and applies the higher rate to the entire request rather than only to the part above the line.
This calculator applies all four and then ranks every live model by the result for your specific token mix. Change the input size, the cache hit rate or the call volume and the order changes, because there is no single cheapest model. A model with a mediocre input rate and an excellent output rate wins on generation-heavy work. A model with a tiny input rate and a long context window wins on retrieval work with a large fixed prefix. A free model wins everything, right up until the point where its rate limits, availability or terms stop you shipping it.
How to use it
- Pick a workload preset, or type your own token counts. The presets are honest starting points: a chatbot turn is roughly 2,000 input and 500 output tokens, a retrieval answer with a document chunk is closer to 6,000 in, and a summarisation job over a long document can be 200,000 in on a single call.
- Set the cache hit rate to the share of your input tokens that genuinely repeat. A fixed system prompt on every call in an agent loop is often 80 to 95 percent. If you are not sure, measure it, because this single number has more leverage on the bill than the model choice.
- Turn on the batch discount only if the work can tolerate the provider's asynchronous window. Half the price is worth nothing if your user is waiting on a response.
- Set a monthly budget to see which models clear it. Rows dim when they do not, and the summary names the cheapest one that fits.
- Replace your estimates with real counts from the provider's usage field before you commit. Once you have measured numbers from your own corpus, the arithmetic here is exact.
Where the number comes from
Rates are pulled live from a public model index that aggregates what providers publish, and the exact fetch time is printed above the results. That is the whole point of the page. Read the two rate columns in the table and you can see why a stored table is worthless: the cheapest input rate is roughly a tenth of a dollar per million tokens and the most expensive is a hundred and fifty, a spread of more than fifteen thousand times across the models live today. A hardcoded rate card is not a little out of date, it is wrong, and it is usually wrong in the direction that makes your architecture look more expensive than it is. Nothing here is written to disk on the server, and the arithmetic runs in your browser, so the same inputs always produce the same answer.
Frequently asked questions
How do I calculate LLM API cost?
Multiply input tokens by the model's input rate, multiply output tokens by its output rate, and add the two. Both rates are USD per token, so divide by one million to get the headline figure providers publish. Reasoning tokens are billed as output. That is the floor. On top of it sit prompt caching, which discounts repeat input at a much lower rate, batch mode, which OpenAI, Anthropic and Google each discount by half, and long-context tiers, where providers charge more per token once a single request crosses a threshold such as 200,000 tokens.
Why is my bill higher than the pricing page says?
Three things move the number. Reasoning tokens bill at the output rate and are easy to forget. Cached input that misses the cache, usually caused by a timestamp or a random ID near the top of the prompt, bills at the full input rate. And a request that crosses a provider's long-context threshold, commonly 200,000 or 272,000 input tokens, jumps to a higher tier for the whole request rather than only for the tokens above the line. This calculator applies all three, and shows the gap between the list price and the effective price for your workload.
What is the cheapest LLM API right now?
It depends on your workload rather than on the model alone, which is why a single answer is rarely correct. Several models are effectively free, and among the paid ones the cheapest input rate is usually a small or fast model from a Chinese or open-weights lab, while the cheapest output rate and the best quality per dollar sit with different vendors. Enter your own token counts and monthly volume above and the table ranks every live model for your specific mix.
How much does prompt caching save?
On a workload with heavy prefix reuse, a great deal. Cache reads are commonly priced at a tenth of the input rate, and on some models at a hundredth, so a 90 percent cache hit rate on a repeated system prompt can cut the input side of the bill by around 90 percent. It only applies to a stable prefix. Anything that changes on every call near the top of the prompt, such as a current timestamp or a request ID, invalidates the prefix and the whole prompt bills at the full rate.
Are these prices live or do they go stale?
Live. Every page load pulls the current rate card from a public model index rather than from a table typed into this page, and the exact fetch time is shown above the results. That matters because the spread across the live table runs from about ten cents to a hundred and fifty dollars per million input tokens, a factor of more than fifteen thousand, and a rate card is the kind of thing that gets revised rather than deprecated. Any calculator carrying hardcoded numbers drifts out of date within weeks. Nothing is written to disk on the server.
Why not just ask an AI chatbot what a model costs?
Because language models answer pricing questions from memory, and provider prices are some of the fastest-moving numbers on the internet. A chatbot will state a confident per-million-token figure that was true eight months ago, and it will not mention cache tiers, batch discounts or long-context thresholds. It also cannot do the arithmetic across hundreds of models for your exact token mix. Everything here is computed from the live rate card in your browser, so the same inputs always give the same answer.
Is the token estimate from pasted text exact?
No, and no browser tool can make it exact without shipping each provider's tokenizer. The estimate uses the standard four-characters-per-token approximation, and the same text can tokenize 20 to 30 percent heavier on one model than another. Use it to get in the right range, then replace it with a real count from your provider's usage field before you sign a contract. Once you have measured counts on your own corpus, this tool becomes exact because it stops guessing.
Related tools
AI Prompt Engineer to cut the input side before you cut the price, Online Calculator for the general arithmetic, and Compound Interest Calculator if you are projecting what the saved budget compounds into.
Comments & Ratings