AI & MLDeep
Intermediate

Context Windows & Long Context

11 min read

Learn
Deep Reading
Estimated 11 mins
Prereq
Intermediate
Basic ML concepts helpful
Interactive
Live Simulator
Contains interactive viz

What a context window is

The context window is the maximum amount of text an LLM can consider at once — the conversation history, any documents you've provided, and its own generated response all share this single budget. Once you exceed it, the model simply cannot see anything beyond the limit; earlier content has to be truncated, summarized, or dropped entirely.

Context windows are measured in tokens, not words or characters. A token is a chunk of text — often a word, part of a word, or punctuation — produced by the model's tokenizer. In English, a rough rule of thumb is that 1,000 tokens correspond to roughly 750 words, though this varies with vocabulary, formatting, and language (code and non-English text often tokenize less efficiently, using more tokens per unit of meaning).

python

Why context is expensive: quadratic attention

The transformer's self-attention mechanism, which lets every token attend to every other token, is what gives LLMs their power — but it's also what makes long context computationally expensive. Computing attention requires comparing every token against every other token, so the number of comparisons grows with the *square* of sequence length: double the context length, and the compute (and memory) required for attention roughly quadruples.

This quadratic scaling is why context windows didn't simply start at a million tokens — early transformer-based models were limited to a few thousand tokens partly because naive attention at longer lengths was computationally prohibitive. Getting to today's long-context models required real architectural and systems engineering work, not just "turning a dial."

Note

Because attention cost grows quadratically with sequence length, a request with 100,000 tokens of context doesn't just cost proportionally more than one with 10,000 tokens — it costs disproportionately more, and takes disproportionately longer to process for the attention computation specifically (though real-world systems use various optimizations that soften this in practice). This is a direct reason long-context API calls are typically priced higher per token than short ones.

How million-token context windows became possible

Several complementary techniques let modern models reach context windows of hundreds of thousands to over a million tokens, despite the quadratic cost of naive attention. Efficient attention implementations like FlashAttention restructure the computation to be far more memory-efficient on GPU hardware, without changing the underlying math — dramatically reducing the practical cost even though the theoretical scaling is unchanged.

Sparse and local attention patterns have some layers or heads attend only to a subset of tokens (nearby tokens, or a fixed set of "global" tokens) rather than the full sequence, trading a small amount of modeling power for large efficiency gains. Better positional encoding schemes, such as rotary position embeddings (RoPE) and extensions to them, allow models to generalize to sequence lengths beyond what they were originally trained on. Architectural changes like grouped-query attention reduce the memory footprint of the key/value cache that has to be stored for every token during generation, which becomes a major bottleneck at long context lengths regardless of the attention computation itself.

Together, these let providers offer context windows that would have been computationally infeasible with the original 2017 transformer architecture running naively at scale.

Context rot: why more context isn't always better

A million-token context window doesn't mean a model uses all million tokens equally well. In practice, models exhibit context rot — a measurable decline in how reliably they retrieve and reason over information as the amount of context grows, even well within the stated limit. Information placed in the middle of a very long context is often used less reliably than information near the beginning or end, a pattern sometimes called "lost in the middle."

This matters practically: dumping an entire codebase or document library into context because the window technically allows it often performs worse than curating the specific, relevant context a task actually needs. Long context is a capability that expands what's *possible*, not a substitute for good context engineering — retrieval, summarization, and relevance filtering remain valuable even when the raw token budget would allow everything to be included.

Note

A large context window lets a model see more at once, but it's not persistent memory — once a conversation ends or the window is exceeded, that information is gone unless explicitly saved somewhere else (a database, a file, a RAG index). Confusing "the model can hold a lot of text right now" with "the model remembers this permanently" is a common source of surprise in production systems.

Strategies for working within context limits

StrategyHow it worksBest for
Full long contextPut all relevant material directly in the promptSmall-to-moderate document sets, one-off deep analysis
RAG (retrieval)Retrieve only the most relevant chunks at query timeLarge knowledge bases, production Q&A systems
Summarization / compressionCondense older conversation turns into summariesLong-running chat sessions, agent memory
Chunking + iterationProcess a large document in sequential smaller passesVery large single documents exceeding even long context

See attention at work

This attention visualization shows how each token computes a weighted relationship to every other token — the exact mechanism whose cost grows quadratically with sequence length, and which context rot and long-context optimizations both try to manage more gracefully.

🧠 Attention Visualizer

Interactive
The
cat
sat
on
the
mat
The
0.35
0.14
0.14
0.13
0.11
0.12
cat
0.15
0.34
0.14
0.13
0.13
0.11
sat
0.12
0.15
0.34
0.13
0.13
0.13
on
0.11
0.13
0.15
0.35
0.14
0.13
the
0.13
0.11
0.13
0.15
0.35
0.14
mat
0.12
0.13
0.11
0.13
0.16
0.35
Hover over cells to view exact attention weights.
Low
High
Each row shows how much that token attends to every other token. Rows sum to 1.

What's next

Context windows determine how much text a model can reason over at once — but for many tasks, the more effective strategy isn't maximizing context, it's retrieving the right context. The RAG lesson covers how retrieval systems select exactly the relevant material a query needs, rather than relying on window size alone.

I build these systems professionally.

Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.

Need custom AI or MarTech setup? Let's build together.