LLM Context Window Comparison
Every major model’s context limit, visualized — with the word counts and, crucially, what it costs to actually fill each window. 12 models, verified August 2026.
Largest window
Gemini 3.5 Flash
1M tokens
That’s about
750K words
≈1,500 pages
Cost to fill it once
Gemini 3.5 Flash input
$1.50 per request
Context windows, side by side
Windows, words, and the cost of filling them
| Model | Context | ≈ Words | ≈ Pages | Cost to fill once |
|---|---|---|---|---|
| Gemini 3.5 Flash Google | 1M | 750K | 1,500 | $1.50 |
| Gemini 3.1 Flash-Lite Google | 1M | 750K | 1,500 | $0.250 |
| GPT-5 OpenAI | 400K | 300K | 600 | $0.500 |
| GPT-5 mini OpenAI | 400K | 300K | 600 | $0.100 |
| GPT-5 nano OpenAI | 400K | 300K | 600 | $0.020 |
| GPT-5.4 mini OpenAI | 400K | 300K | 600 | $0.300 |
| GPT-5.4 nano OpenAI | 400K | 300K | 600 | $0.080 |
| Claude Opus 5 Anthropic | 200K | 150K | 300 | $1.00 |
| Claude Sonnet 5 Anthropic | 200K | 150K | 300 | $0.600 |
| Claude Haiku 4.5 Anthropic | 200K | 150K | 300 | $0.200 |
| GPT-4o OpenAI | 128K | 96K | 192 | $0.320 |
| GPT-4o mini OpenAI | 128K | 96K | 192 | $0.019 |
Word estimates assume English prose at 0.75 words per token (1.33 tokens per word) and 500 words per page. "Cost to fill once" is list-price input cost for a single request that uses the entire window — with caching, repeated long-context requests drop to roughly 10% of that.
Frequently asked questions
What is an LLM context window?
The context window is the maximum number of tokens a model can process in a single request — system prompt, conversation history, retrieved documents and the model’s own output all count against it. Exceed it and the API either truncates the oldest content or returns an error. Bigger windows let you stuff entire codebases, long documents, or multi-hour agent histories into one request.
Which model has the largest context window?
Gemini 3.5 Flash currently leads with 1M tokens — roughly 750K words or about 1,500 pages of text. Google’s Gemini lineup ships 1M across the family, OpenAI’s GPT-5 family offers 400K, and Anthropic’s Claude models offer 200K.
How many words fit in 128K tokens?
About 96,000 words — or roughly 190 pages. English text averages 1.3 tokens per word (0.75 words per token), so a 128K-token window holds ≈ 96K words. Code is denser: expect 30–50% fewer "useful" tokens after comments and whitespace. Our GPT-5 entry holds 300K words at 400K tokens.
Does a bigger context window cost more?
The window itself is free — you pay per token actually sent, not for the capacity. But in practice big contexts are expensive: sending ${fmtTokens(largest.contextWindow)} input tokens to fill Gemini’s window costs ${fmtCost((largest.contextWindow / 1_000_000) * (largest.inputPer1M ?? 0))} at list price, every single request. Prompt caching (10% of normal input cost on Claude and GPT-5) is how production systems survive long contexts.
What happens if my prompt exceeds the context window?
The API rejects the request with a context-length error — it does not silently truncate in most cases. You either trim history, summarize older turns, switch to a bigger-window model, or move to a RAG setup that retrieves only relevant chunks instead of stuffing everything.
Is a bigger context window always better?
Not necessarily. Effective context use often degrades deep into very long inputs — models can miss details buried in the middle of huge prompts ("lost in the middle" behavior). For retrieval-heavy workloads, a well-chunked RAG pipeline over a 200K model frequently beats dumping 1M tokens of raw text. Pick the window that fits your real document sizes, not the biggest number on the box.
Continue with the toolkit
Check how many tokens your document actually uses with the AI Token Counter, compare per-token rates on the Pricing Comparison, or turn volumes into a monthly bill with the API Cost Calculator.