Skip to content
LLM Toolkit

LLM API Pricing Comparison

Every major model, priced side by side. 12 models from OpenAI, Anthropic and Google — input, cached input and output rates per 1M tokens, plus context windows. Prices verified against official provider pages as of 2026-08-25; click any column header to sort.

Cheapest input

GPT-5 nano

$0.050 /1M in

Cheapest output

GPT-5 nano

$0.40 /1M out

Largest context

Gemini 3.5 Flash

1000K tokens

All 12 models, sorted by input price

Model Context Input /1M Cached /1M Output /1M Verified
GPT-5 nano OpenAI 400K $0.050 $0.005 $0.40 2026-08-25
GPT-4o mini OpenAI 128K $0.15 $0.037 $0.60 2026-08-25
GPT-5.4 nano OpenAI 400K $0.20 $0.020 $1.25 2026-08-25
GPT-5 mini OpenAI 400K $0.25 $0.025 $2.00 2026-08-25
Gemini 3.1 Flash-Lite Google 1000K $0.25 $1.50 unverified
GPT-5.4 mini OpenAI 400K $0.75 $0.075 $4.50 2026-08-25
Claude Haiku 4.5 Anthropic 200K $1.00 $0.10 $5.00 2026-08-25
GPT-5 OpenAI 400K $1.25 $0.13 $10.00 2026-08-25
Gemini 3.5 Flash Google 1000K $1.50 $9.00 2026-08-25
GPT-4o OpenAI 128K $2.50 $0.63 $10.00 2026-08-25
Claude Sonnet 5 Anthropic 200K $3.00 $0.30 $15.00 2026-08-25
Claude Opus 5 Anthropic 200K $5.00 $0.50 $25.00 2026-08-25

Prices are per 1 million tokens in USD, from official provider pricing pages. A green dot means the row was checked against the provider's own page on the date shown; an amber dot means the figure comes from a secondary source and awaits verification. Llama and other open-weight models are served at provider-dependent rates and are excluded until we can verify them.

How we keep this table honest

Most pricing trackers quietly rot: providers change rates, launch tiers, retire models — and the table keeps showing last year's numbers. We do three things differently. Every row carries its own verification date, checked against the provider's official pricing page. Rows we haven't verified are labeled as such instead of pretending confidence. And the whole table gets re-checked on the first of every month — including after major model launches, when prices tend to move.

Want to sanity-check your own numbers before committing to a provider? Count your real token usage with the AI Token Counter — it runs the exact OpenAI tokenizer in your browser, so the counts match your bill.

Frequently asked questions

What is the cheapest LLM API right now?

As of August 2026, the cheapest verified production-grade API is GPT-5 nano at $0.050 per 1M input tokens and $0.40 per 1M output tokens. If your workload is output-heavy, GPT-5 nano is the cheapest per output token at $0.40 per 1M. Prices are verified against official provider pricing pages.

How much does GPT-5 cost per million tokens?

GPT-5 costs $1.25 per 1M input tokens and $10.00 per 1M output tokens, with cached input at $0.125. The smaller tiers are far cheaper: GPT-5 mini runs $0.25 in / $2.00 out, and GPT-5 nano runs $0.05 in / $0.40 out — 25× cheaper than the flagship on input.

How much does Claude cost per million tokens?

Anthropic’s current lineup: Claude Opus 5 at $5.00 in / $25.00 out, Claude Sonnet 5 at $3.00 in / $15.00 out, and Claude Haiku 4.5 at $1.00 in / $5.00 out. All Claude models include a 200K context window, and cached input costs 10% of the base input rate.

What does cached input pricing mean?

Prompt caching lets you reuse the input you’ve already sent — a long system prompt, a document, or conversation history — at a steep discount, typically 10% of the normal input price. If your requests repeat most of their context (agents, RAG, multi-turn chat), caching routinely cuts input costs by 60–90%. The cached column in the table shows each model’s discounted rate.

Which LLM has the largest context window?

Gemini 3.5 Flash leads with a 1000K-token context window. Google’s Gemini family ships 1M tokens across the lineup, while OpenAI’s GPT-5 family offers 400K and Anthropic’s Claude models offer 200K. More context costs more per request in practice, because you pay for every input token you send.

How should I pick a model based on price?

Match the model tier to the task difficulty, not the other way around. Route high-volume, simple tasks (classification, extraction, routing) to nano-tier models; save flagship models for hard reasoning. Then check your input/output ratio: agent workloads that generate long outputs should compare output prices first — that’s usually 80% of the bill. Our API Cost Calculator (shipping next) computes this for your exact volumes.

Next in the toolkit

The API Cost Calculator — turn these per-token rates into a real monthly bill for your request volumes — plus Context Window Comparison and the Token ↔ Words Converter. See all tools