Skip to content
LLM Toolkit

LLM Context Window Comparison

Every major model’s context limit, visualized — with the word counts and, crucially, what it costs to actually fill each window. 12 models, verified August 2026.

Largest window

Gemini 3.5 Flash

1M tokens

That’s about

750K words

≈1,500 pages

Cost to fill it once

Gemini 3.5 Flash input

$1.50 per request

Context windows, side by side

Gemini 3.5 Flash (Google) 1M · ≈750K words
1M
Gemini 3.1 Flash-Lite (Google) 1M · ≈750K words
1M
GPT-5 (OpenAI) 400K · ≈300K words
400K
GPT-5 mini (OpenAI) 400K · ≈300K words
400K
GPT-5 nano (OpenAI) 400K · ≈300K words
400K
GPT-5.4 mini (OpenAI) 400K · ≈300K words
400K
GPT-5.4 nano (OpenAI) 400K · ≈300K words
400K
Claude Opus 5 (Anthropic) 200K · ≈150K words
200K
Claude Sonnet 5 (Anthropic) 200K · ≈150K words
200K
Claude Haiku 4.5 (Anthropic) 200K · ≈150K words
200K
GPT-4o (OpenAI) 128K · ≈96K words
128K
GPT-4o mini (OpenAI) 128K · ≈96K words
128K

Windows, words, and the cost of filling them

Model Context ≈ Words ≈ Pages Cost to fill once
Gemini 3.5 Flash Google 1M 750K 1,500 $1.50
Gemini 3.1 Flash-Lite Google 1M 750K 1,500 $0.250
GPT-5 OpenAI 400K 300K 600 $0.500
GPT-5 mini OpenAI 400K 300K 600 $0.100
GPT-5 nano OpenAI 400K 300K 600 $0.020
GPT-5.4 mini OpenAI 400K 300K 600 $0.300
GPT-5.4 nano OpenAI 400K 300K 600 $0.080
Claude Opus 5 Anthropic 200K 150K 300 $1.00
Claude Sonnet 5 Anthropic 200K 150K 300 $0.600
Claude Haiku 4.5 Anthropic 200K 150K 300 $0.200
GPT-4o OpenAI 128K 96K 192 $0.320
GPT-4o mini OpenAI 128K 96K 192 $0.019

Word estimates assume English prose at 0.75 words per token (1.33 tokens per word) and 500 words per page. "Cost to fill once" is list-price input cost for a single request that uses the entire window — with caching, repeated long-context requests drop to roughly 10% of that.

Frequently asked questions

What is an LLM context window?

The context window is the maximum number of tokens a model can process in a single request — system prompt, conversation history, retrieved documents and the model’s own output all count against it. Exceed it and the API either truncates the oldest content or returns an error. Bigger windows let you stuff entire codebases, long documents, or multi-hour agent histories into one request.

Which model has the largest context window?

Gemini 3.5 Flash currently leads with 1M tokens — roughly 750K words or about 1,500 pages of text. Google’s Gemini lineup ships 1M across the family, OpenAI’s GPT-5 family offers 400K, and Anthropic’s Claude models offer 200K.

How many words fit in 128K tokens?

About 96,000 words — or roughly 190 pages. English text averages 1.3 tokens per word (0.75 words per token), so a 128K-token window holds ≈ 96K words. Code is denser: expect 30–50% fewer "useful" tokens after comments and whitespace. Our GPT-5 entry holds 300K words at 400K tokens.

Does a bigger context window cost more?

The window itself is free — you pay per token actually sent, not for the capacity. But in practice big contexts are expensive: sending ${fmtTokens(largest.contextWindow)} input tokens to fill Gemini’s window costs ${fmtCost((largest.contextWindow / 1_000_000) * (largest.inputPer1M ?? 0))} at list price, every single request. Prompt caching (10% of normal input cost on Claude and GPT-5) is how production systems survive long contexts.

What happens if my prompt exceeds the context window?

The API rejects the request with a context-length error — it does not silently truncate in most cases. You either trim history, summarize older turns, switch to a bigger-window model, or move to a RAG setup that retrieves only relevant chunks instead of stuffing everything.

Is a bigger context window always better?

Not necessarily. Effective context use often degrades deep into very long inputs — models can miss details buried in the middle of huge prompts ("lost in the middle" behavior). For retrieval-heavy workloads, a well-chunked RAG pipeline over a 200K model frequently beats dumping 1M tokens of raw text. Pick the window that fits your real document sizes, not the biggest number on the box.

Continue with the toolkit

Check how many tokens your document actually uses with the AI Token Counter, compare per-token rates on the Pricing Comparison, or turn volumes into a monthly bill with the API Cost Calculator.