How AI is priced
Free tiers, monthly subscriptions, business plans and pay-per-token APIs explained, with what a document actually costs and when paying is worth it.
Five ways to pay
AI pricing looks complicated, but it comes down to a handful of models. Many companies use several at once.
| Model | How you pay | Who it suits | Watch out for |
|---|---|---|---|
| Free plan | Nothing; sign up and go | Everyday questions, occasional writing | Low limits; newer models and advanced features are often missing; some free plans carry ads |
| Personal subscription | A fixed monthly fee, usually cheaper if paid yearly | People who use AI daily and keep hitting the free limits | “Unlimited” usually comes with conditions; every plan has limits |
| Team / enterprise | Per person (per seat) per month, sometimes plus usage | Companies and teams | Admin controls, and work content isn’t used for training by default (see Privacy and safety) |
| Pay-as-you-go API | Prepay or add a card; charged per token | Developers building software | No monthly fee, but costs grow with use; set a spending limit |
| Usage credits | Buy a pool of credits that features draw down | People who occasionally exceed their plan; image, video and agent tools | How fast credits go depends on the task, which is hard to predict |
Tokens: the unit AI is priced in
A token is a small chunk of text that a model reads or writes. A short common English word is usually one token; longer words get split into pieces, and punctuation and spaces count too. APIs bill per token, and subscription limits are based on tokens behind the scenes.
How much text is a token? The official rules of thumb are approximate and mostly about English:
| Source | What it says |
|---|---|
| OpenAI developer docs | In English, 1 token is about 4 characters or 0.75 words |
| Anthropic (Claude) pricing docs | In English, 1 token is about 4 characters or 0.75 words; the exact count varies by language and content |
| Google Gemini API docs | 1 token is about 4 characters; 100 tokens are about 60–80 English words |
| DeepSeek API docs | 1 English character is about 0.3 tokens; 1 Chinese character is about 0.6 tokens; ratios differ between models |
Be careful with other languages. Of these four, only DeepSeek gives a ratio for Chinese, and it says the figure applies to its own models. OpenAI, Anthropic and Google publish English rules of thumb only. The same Chinese paragraph can come out as quite different token counts in different models.
Models from the same company differ too. Anthropic’s pricing docs say Claude 4.7 and later models use a new tokenizer that produces about 30% more tokens for the same text. For an exact number, use the provider’s token-counting tool or read the usage figures returned with each request.
There’s also a short definition in the glossary.
Input, output and cached tokens
An API bill usually has three lines, with very different prices:
- Input tokens: everything you send the model, including your question, any documents, earlier messages in the conversation and system instructions. In a long chat, the earlier messages are sent again each turn, so every turn costs more than the last.
- Output tokens: what the model writes. These cost much more; on Claude’s price list, output is five times the input price. For reasoning models that “think” before answering, the thinking counts as output too: Google’s price list labels its output price “including thinking tokens”.
- Cached tokens: if you send the same large block of text repeatedly, such as the same long document, the provider can keep it ready and charge very little to read it again. On Claude, writing to the cache costs 1.25 times the input price (for a five-minute cache), and each later read costs a small fraction of it.
- Claude Sonnet 5.5 input (per million tokens)
- $2
- Claude Sonnet 5.5 output (per million tokens)
- $10
- Claude Sonnet 5.5 cache read (per million tokens)
- $0.10
- Batch processing (non-urgent jobs)
- −50%
Source: Claude API pricing docs, checked 8 October 2026. Prices change; check the official page.
A worked example: summarising a 20-page document via the API
Here is the arithmetic using Claude’s official price list. Say you have a 20-page English report of about 500 words a page, and you ask for a summary of roughly one page.
- 1Estimate the words20 pages × 500 words = about 10,000 words.
- 2Convert to tokensAt 1 token ≈ 0.75 words: 10,000 ÷ 0.75 ≈ 13,300 tokens. The newer tokenizer adds about 30%: ≈ 17,300. Add your instructions and round to 18,000 input tokens.
- 3Estimate the outputA one-page summary: call it 1,000 output tokens.
- 4Multiply by the priceCost = tokens × price ÷ 1,000,000.
With Claude Sonnet 5.5 ($2 input, $10 output per million tokens):
- Input: 18,000 × $2 ÷ 1,000,000 = $0.036
- Output: 1,000 × $10 ÷ 1,000,000 = $0.010
- Total: about $0.046, under five cents.
The same job on other models from the same price list:
Standard prices per million input / output tokens: Haiku 5.5 $0.10 / $0.50 (prompts up to 100,000 tokens), Opus 5.5 $4 / $20. Source: Claude API pricing docs, checked 8 October 2026.
What if you ask ten follow-up questions? Each question sends the whole document again as input. Still on Sonnet 5.5, with answers of about 500 tokens each:
| Without caching | With caching (questions within five minutes of each other) | |
|---|---|---|
| The document | 10 × 18,000 × $2 ÷ 1M = $0.36 | First write: 18,000 × $2.50 ÷ 1M = $0.045; nine reads: 162,000 × $0.10 ÷ 1M ≈ $0.016 |
| The answers | 5,000 × $10 ÷ 1M = $0.05 | $0.05 |
| Total | about $0.41 | about $0.11 |
Two lessons. A single document costs very little through the API; what gets expensive is large volumes of repeated calls. And in long conversations, most of the cost is often the context sent again and again, not the new line you typed. These are estimates; the real token count is whatever the API reports.
Personal plans from the three biggest assistants
This table lists only prices we confirmed on official pages, as of October 2026, in US dollars before tax. Prices can differ by country and when you subscribe through an app store, so check the official page.
| Tier | ChatGPT | Claude | Google AI (Gemini) |
|---|---|---|---|
| Free | Yes; OpenAI says text chats are unlimited (subject to abuse guardrails), with limits on uploads, images, deep research and more | Yes | Yes |
| Entry | Go (price on the official page; may include ads) | — | Google AI Plus: $4.99/month, 2× the usage of Free |
| Standard | Plus: $20/month | Pro: $20/month ($17/month billed yearly), at least 5× the usage of Free | Google AI Pro: $19.99/month, 4× the usage of Free |
| Heavy use | Pro: $100, $200 or $500/month | Max: from $100/month, 5× or 20× the usage of Pro | Ultra: $99.99/month (5× Pro) or $199.99/month (20× Pro) |
| Teams | Business: $20 per user/month billed yearly, $25 monthly; 2+ users | Team standard seat: $20 per seat/month billed yearly, $25 monthly | Gemini Enterprise Business edition: from $21 per seat/month |
Official pages: ChatGPT (the monthly prices come from the Codex pricing page in OpenAI’s developer docs), Claude and Google AI. For more on each product, see our ChatGPT, Claude and Gemini pages.
Chinese assistants: the basics are mostly free
The assistants most used in China generally don’t charge for basic chat. Paid features tend to be agents, long tasks and generation:
- DeepSeek: free on the web and in the app. Its API is billed per token, with off-peak hours at half the peak price.
- Doubao: TechNode reported in May 2026 that ByteDance had begun testing three paid tiers aimed at compute-heavy uses such as slides, data analysis and video. The company said the free version would remain available for everyday use and that the paid plans were still being tested.
- Kimi: its help centre lists four memberships at ¥49, ¥99, ¥199 and ¥699 a month, cheaper if paid yearly. All member features draw on one pool of credits based on actual token use; credits refresh monthly and don’t roll over, and you can switch on pay-as-you-go top-ups when they run out.
For Qwen and others, check the official site.
Cheaper per token isn’t always cheaper per task
When people compare API prices they look at the cost per million tokens. What matters is what it costs to get the job done. At least three things get in the way:
- Models use different numbers of tokens for the same job. Anthropic says Opus 5.5 has a lower API price than Opus 5 (input down from $5 to $4) and also uses fewer tokens to finish the same work; together, typical costs fall by about 40% (see our Opus 5.5 story). Sonnet 5.5 has exactly the same price as Sonnet 5, yet Anthropic says most tasks cost up to 30% less, again because it uses fewer tokens (see Sonnet 5.5).
- Tokenizers differ. As noted above, Claude’s newer tokenizer counts the same text as about 30% more tokens. Two models with the same price per token can produce different bills.
- Mistakes mean doing it again. A cheap model that needs three tries, plus your time checking and fixing, may cost more than an expensive one that gets it right first time. In one internal test rewriting a piece of software, Anthropic reported that Opus 5.5 finished the job for 51% less than the pricier Fable 5.1. It’s the company’s own test, but the lesson holds: compare total cost, not unit price.
A practical approach: take a task you really do repeatedly, run it a few times on two or three models, and note how often each succeeds, how long it takes and what it costs including retries.
When is paying worth it?
Start free. Consider upgrading when:
- You keep hitting limits, waiting for resets several times a day while real work piles up.
- You need a stronger model or a specific feature, such as very long files, deep research, letting AI operate your computer, or coding (Claude Code is included in every paid Claude plan; paid ChatGPT plans include more Codex usage than Free).
- You use it for work. With company material, the data terms of a team or enterprise plan usually matter more than the extra allowance.
Habits that save money:
- Subscribe to one at a time. The standard tiers cost about the same; paying for three is usually waste. Use one for a month or two before adding another.
- Pay monthly to try, yearly once you’re sure. On Claude’s official page, Pro costs about 15% less paid yearly.
- Heavy users: do the maths. If you only use AI intensively a few days a month, credits or pay-as-you-go may beat jumping to a $100+ tier.
- Developers: set spending limits, and use caching and batching. Claude, OpenAI and Google all offer roughly half price for non-urgent batch jobs.
- Remember to cancel. Subscriptions renew automatically, and if you subscribed through Apple’s or Google’s app store, you cancel there.
Sources
- Claude: plans and pricing (individual, team, enterprise, API)
- Claude API docs: pricing (caching, batch, token rule of thumb)
- ChatGPT: plan comparison (features and limits)
- OpenAI developer docs: Codex pricing (lists ChatGPT Plus, Pro and Business prices)
- OpenAI developer docs: API pricing
- OpenAI developer docs: key concepts (token rule of thumb)
- Google: Google AI plans (US prices)
- Google Gemini API docs: understand and count tokens
- Google Gemini API docs: pricing
- DeepSeek API docs: token and token usage
- Kimi Help Center: membership pricing
- TechNode: ByteDance tests paid subscriptions for Doubao (6 May 2026)
- Wikipedia: DeepSeek (chatbot)