Two ways to pay for a model: token-based pricing and subscription plan
You can't work with a model directly, so you'll either need to use its native interface or a so-called agent that connects to it via API. Less experienced users often mistake this second option for direct interaction, especially when the agent's name just happens to match the model from the same company.
A third party working via API pays with tokens. On average, the prices on publicly available rate cards look attractive: 1 million tokens (~750,000 words, or roughly 3,000 pages) runs anywhere from $1.50 to $25. Users often assume the same pricing applies when working with models' interfaces, even when they're using monthly plans.

Token-based API pricing
Let's start with the simpler case: working through the API. A user picks an agent, pays a fee for its technology, and settles up with the model. They hand over a credit card and watch their balance. That's transparent enough, right? The problems only start when the user orders deep research. A model's search–read–search again loop burns through tokens at 4–15 times the rate of a normal chat conversation. Fair enough: It is Deep Research, after all.
August 2026 Price list
Here is the official rate per million tokens as of August 10, 2026. It's immediately clear that input and output tokens are priced in a completely different league.
| Model | Input / 1M | Output / 1M |
|---|---|---|
| Claude Fable 5 | $10 | $50 |
| Claude Opus 5 | $5 | $25 |
| Claude Sonnet 5 | $2* | $10* |
| GPT-5.5 | $5 | $30 |
| Gemini 3.1 Pro | $2 | $12 |
| Gemini 3.5 Flash | $1.50 | $9 |
* introductory price through August 31, 2026; standard rate from September 1 — $3 / $15
This means that editing your own text — sending the LLM long passages and getting short recommendations back — is significantly cheaper than having the model write from scratch, where you send a brief prompt and receive a large block of text in return. Translation costs, meanwhile, amount to exactly the sum of two line items: input and output tokens are roughly equal.
The peculiarities of translation pricing
With translation, you can immediately forget about the rate card being accurate. All non-English words contain between 2 and 2.5 times more tokens than their English equivalents. (Officially, this is because English is the LLM's native language, the one in which most of its training took place.) That's an average, though. There is no rate card broken down by language.
Opaque pricing for thinking
If the opacity around translations and deep research feels insufficient for even the most honest and transparent form of API pricing, here's some great news: reasoning is just as opaque. When the model "thinks," it generates a chain of reasoning that is usually invisible and costs the output rate. In Google's pricing table, the relevant column is explicitly labeled "Output price (including thinking tokens)," while OpenAI states that reasoning tokens "are billed as output tokens." To cover their tracks completely, each of the Big Three has introduced right-click controls for adjusting the "strength" or "depth" of reasoning. Deeper thinking costs more, that much is clear, but how much more? As you might expect, it all depends on the depth. No rate card will save you here, not that one exists.
Even that wasn't enough for Anthropic. The company added a fourth dimension of opacity and recruited the tokenizer to do it.

The best way to raise the price without touching the price tag
A tokenizer is the algorithm that chops your text into tokens before the model ever reads it. The rate isn't constant: if 1,000 words produced 1300 tokens yesterday, today that same text might yield 1800. For instance, Anthropic swapped out the tokenizer in Opus 4.7 and admitted that the same input can now produce up to 1.35x more tokens than before. Each one still costs between $5 and $25, but you need more, which raises your total bill.
Anthropic assures users that, on average, consumption didn't increase on their own code benchmarks, and that the model became more efficient in other respects. Finer-grained tokens may be more precise and require fewer computations, but you have no way to verify that on your own text — and you'll notice the extra third of tokens in your bill immediately.
The risks of cheaper models
In any case, the per-token price alone is a poor guide. Lingjiao Chen's paper "The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More" shows that cheaper reasoning models ended up costing more than expected in 21.8% of cases and the gap could be as high as 28x. Gemini 3 Flash, for example, is 78% cheaper than GPT-5.2, yet its actual cost across real-world tasks comes out 22% higher.
The authors attribute the spread to wildly uneven token consumption during reasoning: for the same prompt, one model can consume 900% more reasoning tokens than another. ChatGPT also has a well-known tendency to ramble, and if a model consistently produces 25% more words in the form of padding, restated points, and exhausting extended metaphors, it can comfortably cut its headline price by 10%, come out looking like the winner in every comparison chart, and still make more money.
"Listed API pricing is an unreliable proxy for actual cost," Chen et al. conclude.

The hard way to find out the truth
You can track your own spending under the most transparent billing system available — API plus a rate card — but it's a bit tedious and requires some care. Start by noting your meter reading, then run a set of standard tasks. Note the new reading. Run the same tasks again. Note the reading again. After the third or fourth round, if your measurements are reasonably consistent, you can relax.
You're probably better off not trying to compare two or three other models at the same time. That would mean scoring response quality in a separate column, and fatigue might cloud your judgment at that point.
This method, tiresome as it is, gives you at least some hope of getting a meaningful answer from API-based token billing. With a flat subscription, things are considerably worse.

Subscriptions: familiar meters with a lot of fine print
A subscription is the standard way to pay for a model when you're working with it through its own native agent. There's no unlimited plan like mobile carriers offer and there won't be: the cost of building AI models and running servers is significantly higher than that of operating mobile carriers or conventional SaaS.
How do you cap a particularly heavy user? You set limits. What do you measure those limits in? Over what time window? Do you warn users about them? Do you let them know when the limits actually kick in? Well, none of that really matters. What matters is that every model goes its own way — that its metrics are impossible to compare with a competitor's. Naturally, you make the whole scheme as complicated as possible.
When you have to change these limits for the worse, do so as quietly as you can, ideally by having some engineer post on X. The backlash will come regardless, but it'll be contained within a niche community that lives for exactly this kind of drama. Since the media and social-media landscape in 2026 is thoroughly fragmented, the noise from even a very loud scandal fades fast, especially when arguing back requires spreadsheets in hand and calculations that the average user frankly can't be bothered to work through.
Some users have no idea limits exist at all. "I think my model gets tired by evening. Especially after long sessions." It doesn't get tired. It gets switched off, and a cheaper, dumber model gets quietly slipped in its place. Sometimes a fleeting message announces this. Sometimes not.
ChatGPT swaps in a weaker model
ChatGPT's Plus plan theoretically lets you send up to 160 messages in 3 hours for $20 a month. But that's only for some ideal use case with no context or reasoning. Sure, you could spend 3 hours discussing the meaning of life, the complexities of human relationships, or your dreams. Actually working on this plan for 3 hours is nearly impossible. Even an hour is a stretch, depending on what you're doing.
When a user hits the cap, the advanced ChatGPT quietly falls back to a weaker model. "After reaching this limit, chats will switch to the mini version of the model until the limit resets" — that's what OpenAI's help documentation has to say about it. You keep typing; a cheaper model is already answering. Quality has dropped with no noticeable warning. You think you're still talking to the flagship, which is just a little slow today, but you're actually interacting with the mini.
Gemini burns through limits on complex requests
Gemini does warn users, but it doesn't splash a giant banner across the screen screaming "You're approaching your limit — think twice!" Instead, like ChatGPT, it quietly downgrades to a weaker model and displays a small line of text: "Now using a lighter model to keep you going."
Don't expect a transparent spending report from this AI, either. In May 2026, Google announced several new pricing tiers, and the new logic rolled out across the entire Gemini lineup. The principle was simple: basic requests barely dent your quota, while complex tasks can burn through it fast. Limits reset every five hours, but everything ultimately ran into a weekly ceiling. As a result, some users barely noticed any change, while others quickly hit restrictions in their normal workflows. What was especially frustrating was that two outwardly similar prompts could "cost" wildly different amounts with no explanation. In response to user outcry, Google adjusted the new limits — but did not go back to the old system.
How Claude plays the user
Claude's reputation in this space looks somewhat better. It does show usage, for instance — not in tokens, not in requests, but as a percentage. You can navigate to Account⇒Settings⇒Usage and check a progress bar showing how much of your five-hour window and weekly allowance remain.
But Anthropic, which works hard to appear more principled than the rest, has the same problems. In March 2026, it quietly reworked the mechanics of Claude's five-hour limits: during weekday peak hours — 5 to 11 a.m. — the same session began burning through quota faster than before, even though the weekly limits officially hadn't changed. There was no press release, just an engineer's post on X: "You'll move through your 5-hour session limits faster than before." Users reacted badly, but, as with Google, Anthropic's leadership didn't abandon the idea, just tweaked it. It's a consistent pattern: companies prefer to quietly cut limits first, then explain their actions after the fact as policy optimization and an effort to balance demand against available compute.
Lawsuits
Not all users are willing to accept this quietly. Some are taking it to court. Karl Khan, a Washington state resident, has sued Anthropic for misleading consumers. He argues that the Claude Max 5x plan at $100 per month and the Claude Max 20x plan at $200 per month were advertised as offering five and twenty times more access than Claude Pro, respectively, but the actual limits were far lower and nearly impossible to predict. By his account, a single five-hour session burned through 15% of his weekly allowance. He asked the court to certify the case as a class action so that other Claude Max subscribers could join.
For now, this is only an allegation, not a verdict, but the very fact that subscription limits are ending up in litigation shows that users are losing patience with opacity.

Which is better: a subscription or paying as you go?
There's no clear-cut answer here. It depends on how you work and how much. If you keep hitting the ceiling of your subscription plan — whether it costs $20 or $200 — you're constantly at risk of a silent model swap at the worst possible moment. If you never actually use your full allowance and hover somewhere in the middle, you're paying double the going rate per token and effectively subsidizing either the shareholders or the power users. Which is worse is for each person to decide.
Questions.
How many tokens are in 1,000 words?
What's cheaper — a subscription or the API?
Why does my API bill fluctuate even when my usage seems similar?
Which model is cheapest for drafts?
Sources
References cited in this piece. Last verified on the published or revision date.
- 01
- 02
- 03
- 04
- 05
- 06
- 07
- 08
- 09
- 10
- 11