LLM API Pricing Comparison
Every major model, priced side by side. 12 models from OpenAI, Anthropic and Google — input, cached input and output rates per 1M tokens, plus context windows. Prices verified against official provider pages as of 2026-08-25; click any column header to sort.
Cheapest input
GPT-5 nano
$0.050 /1M in
Cheapest output
GPT-5 nano
$0.40 /1M out
Largest context
Gemini 3.5 Flash
1000K tokens
All 12 models, sorted by input price
| Model | Context | Input /1M | Cached /1M | Output /1M | Verified |
|---|---|---|---|---|---|
| GPT-5 nano OpenAI | 400K | $0.050 | $0.005 | $0.40 | 2026-08-25 |
| GPT-4o mini OpenAI | 128K | $0.15 | $0.037 | $0.60 | 2026-08-25 |
| GPT-5.4 nano OpenAI | 400K | $0.20 | $0.020 | $1.25 | 2026-08-25 |
| GPT-5 mini OpenAI | 400K | $0.25 | $0.025 | $2.00 | 2026-08-25 |
| Gemini 3.1 Flash-Lite Google | 1000K | $0.25 | — | $1.50 | unverified |
| GPT-5.4 mini OpenAI | 400K | $0.75 | $0.075 | $4.50 | 2026-08-25 |
| Claude Haiku 4.5 Anthropic | 200K | $1.00 | $0.10 | $5.00 | 2026-08-25 |
| GPT-5 OpenAI | 400K | $1.25 | $0.13 | $10.00 | 2026-08-25 |
| Gemini 3.5 Flash Google | 1000K | $1.50 | — | $9.00 | 2026-08-25 |
| GPT-4o OpenAI | 128K | $2.50 | $0.63 | $10.00 | 2026-08-25 |
| Claude Sonnet 5 Anthropic | 200K | $3.00 | $0.30 | $15.00 | 2026-08-25 |
| Claude Opus 5 Anthropic | 200K | $5.00 | $0.50 | $25.00 | 2026-08-25 |
Prices are per 1 million tokens in USD, from official provider pricing pages. A green dot means the row was checked against the provider's own page on the date shown; an amber dot means the figure comes from a secondary source and awaits verification. Llama and other open-weight models are served at provider-dependent rates and are excluded until we can verify them.
How we keep this table honest
Most pricing trackers quietly rot: providers change rates, launch tiers, retire models — and the table keeps showing last year's numbers. We do three things differently. Every row carries its own verification date, checked against the provider's official pricing page. Rows we haven't verified are labeled as such instead of pretending confidence. And the whole table gets re-checked on the first of every month — including after major model launches, when prices tend to move.
Want to sanity-check your own numbers before committing to a provider? Count your real token usage with the AI Token Counter — it runs the exact OpenAI tokenizer in your browser, so the counts match your bill.
Frequently asked questions
What is the cheapest LLM API right now?
As of August 2026, the cheapest verified production-grade API is GPT-5 nano at $0.050 per 1M input tokens and $0.40 per 1M output tokens. If your workload is output-heavy, GPT-5 nano is the cheapest per output token at $0.40 per 1M. Prices are verified against official provider pricing pages.
How much does GPT-5 cost per million tokens?
GPT-5 costs $1.25 per 1M input tokens and $10.00 per 1M output tokens, with cached input at $0.125. The smaller tiers are far cheaper: GPT-5 mini runs $0.25 in / $2.00 out, and GPT-5 nano runs $0.05 in / $0.40 out — 25× cheaper than the flagship on input.
How much does Claude cost per million tokens?
Anthropic’s current lineup: Claude Opus 5 at $5.00 in / $25.00 out, Claude Sonnet 5 at $3.00 in / $15.00 out, and Claude Haiku 4.5 at $1.00 in / $5.00 out. All Claude models include a 200K context window, and cached input costs 10% of the base input rate.
What does cached input pricing mean?
Prompt caching lets you reuse the input you’ve already sent — a long system prompt, a document, or conversation history — at a steep discount, typically 10% of the normal input price. If your requests repeat most of their context (agents, RAG, multi-turn chat), caching routinely cuts input costs by 60–90%. The cached column in the table shows each model’s discounted rate.
Which LLM has the largest context window?
Gemini 3.5 Flash leads with a 1000K-token context window. Google’s Gemini family ships 1M tokens across the lineup, while OpenAI’s GPT-5 family offers 400K and Anthropic’s Claude models offer 200K. More context costs more per request in practice, because you pay for every input token you send.
How should I pick a model based on price?
Match the model tier to the task difficulty, not the other way around. Route high-volume, simple tasks (classification, extraction, routing) to nano-tier models; save flagship models for hard reasoning. Then check your input/output ratio: agent workloads that generate long outputs should compare output prices first — that’s usually 80% of the bill. Our API Cost Calculator (shipping next) computes this for your exact volumes.
Next in the toolkit
The API Cost Calculator — turn these per-token rates into a real monthly bill for your request volumes — plus Context Window Comparison and the Token ↔ Words Converter. See all tools