How we calculate AI API costs
Every number on our AI pricing pages comes from the providers’ published prices and a small set of stated assumptions. This page explains them, so you can judge how far to trust a figure and adjust it for your own use.
Where the prices come from
API prices come directly from each provider’s official pricing page, using the standard (global, pay-as-you-go) tier in US dollars. We record the regular input, cached input, output, and batch prices for each current model, plus any long-context price that applies above a prompt-length threshold. Subscription prices for the consumer chat apps come from each provider’s plans page; where a provider’s page couldn’t be read directly, we cross-checked recent published pricing.
- Anthropic: platform.claude.com/docs/en/about-claude/pricing
- OpenAI: developers.openai.com/api/docs/pricing
- Google: ai.google.dev/gemini-api/docs/pricing
- DeepSeek: api-docs.deepseek.com/quick_start/pricing
- xAI: docs.x.ai/docs/models
All prices were last checked on October 10, 2026. Providers change prices and release models often, so we recheck at least monthly. If you spot a price that’s changed, the provider’s own page is always the final word.
Which models we include
For each provider we include its current recommended models across three tiers: flagship (most capable), balanced (the everyday model most applications should use), and fast (cheapest, for simple high-volume work). Providers also sell older models, specialized models for images, audio, and video, and premium variants; we leave those out so the comparison stays focused on the text models most businesses actually choose between. DeepSeek and xAI don’t sell a model that fits the balanced tier, so they appear only in the flagship and fast tiers.
How we estimate tokens
A token is a chunk of text. For English, the widely used rule of thumb is that one token is about four characters, or three-quarters of a word. We use roughly 1.33 tokens per word to size each job. It’s an estimate: each provider splits text differently, so the same job can use somewhat more or fewer tokens on different models, and Anthropic notes that its newer Claude models use around 30% more tokens for the same text than its older ones.
For that reason, treat the calculator as a way to compare options and get the order of magnitude right, not as an invoice. Once you’ve picked a model, send a sample of real requests and read the token counts from your provider’s usage dashboard.
The everyday jobs we price
Rather than ask you to think in tokens, the calculator and pricing pages use a set of common business jobs. Each has an input size (instructions plus whatever the model reads) and an output size (what it writes), and a typical monthly volume for a small business:
| Job | Input tokens | Output tokens | Default volume |
|---|---|---|---|
| Draft an emailInstructions plus a short thread in; a 300-word reply out. | 800 | 400 | 1,000 emails |
| Answer a support ticketTicket, customer history, and help-center excerpts in; a reply out. | 2,500 | 350 | 2,000 tickets |
| Summarize a 10-page documentAbout 5,000 words in; a one-page summary out. | 7,000 | 700 | 300 documents |
| Write a 1,500-word articleA brief and outline in; a full draft out. | 1,200 | 2,000 | 100 articles |
| Write a product descriptionProduct facts in; 150–200 words out. | 400 | 280 | 5,000 descriptions |
| Chatbot conversation (10 turns)Each turn resends the conversation so far. | 12,000 | 2,000 | 3,000 conversations |
| AI coding or research agent taskMany steps reading files or web pages; long context. | 150,000 | 10,000 | 200 tasks |
A few notes on the sizes. Input includes the instructions you send with every request, not just the material being worked on; for a support bot that’s often several hundred tokens of guidance plus excerpts from your help center. The chatbot conversation is input-heavy because each turn resends the whole conversation so far. The agent task represents a long, multi-step job that reads many files or pages; real agent tasks vary enormously, from a few thousand to millions of tokens.
How the cost is calculated
For each model, the monthly cost is:
Three adjustments apply when relevant:
- Long-context pricing. If a single request’s input is over a model’s threshold (for example 272,000 tokens on OpenAI’s GPT-6 models, 200,000 on Gemini 3.1 Pro and Grok, or 100,000 on Claude Haiku 5.5), the higher long-context prices apply to the whole request, as the providers specify.
- Batch. If you choose “can wait,” we apply each provider’s published batch prices, typically half the standard rate. Providers without a batch tier keep their standard prices.
- DeepSeek off-peak. DeepSeek’s prices are halved outside its peak hours. We use peak prices by default; the off-peak option halves DeepSeek’s cost.
What the numbers leave out
- Tool and feature fees: web search, code execution, file storage, and cache storage fees are charged separately by some providers.
- Cache write costs: some providers charge extra to write to the cache; we only model the cheaper cached reads.
- Reasoning overhead: models that think before answering bill that thinking as output, which can make output much longer than the visible answer.
- Retries and failures: real applications resend some requests.
- Regional and premium tiers: data-residency endpoints, priority tiers, and fast modes cost more.
- Images, audio, and video: these are priced differently and aren’t included.
- Promotions and free tiers: we show promotional prices where they’re the current list price and note when they end. Free-tier allowances aren’t subtracted.
For a production budget, add a margin of 20–50% to account for these, or measure a week of real traffic and scale from that.
Who maintains this
These pages are written and maintained by Clarence T. Archibald IV for the Free Claude Cowork Course, which teaches business owners and teams to use AI in their daily work. We’re independent and aren’t paid by any of the providers listed. Questions or corrections are welcome through our contact page.
Frequently asked questions
How many tokens is a word?
In English, one token is roughly three-quarters of a word, so 1,000 tokens is about 750 words and 1,000 words is about 1,330 tokens. The exact number depends on the provider and the text: code, numbers, and non-English languages usually use more tokens per word.
Why do providers count tokens differently?
Each provider uses its own tokenizer, the method that splits text into tokens. The same paragraph can be a different number of tokens on different models. Anthropic says its newer Claude models produce about 30% more tokens for the same text than its older ones, for example. That’s why per-token prices are a good guide but not an exact comparison.
How often are the prices updated?
We check every price against the providers’ official pricing pages at least monthly, and whenever a major model launches. The date of the latest check is shown on every page; the current data was checked on October 10, 2026.
Do the prices include tax?
No. All prices are each provider’s published list prices in US dollars, before tax, on their standard global tier. Enterprise agreements, committed-use discounts, and cloud marketplace prices can differ.