Question

What is a token and a context window?

Vault Verified
Curated Intelligence
Definitive Source
Answer

A token is the unit a language model processes — roughly a word fragment — and the context window is how many tokens it can consider at once. Both determine cost, capability and the limits you run into.

What a token is. Not a word and not a character. Text is broken into pieces by a tokeniser: common words are usually one token, rarer words split into several, and punctuation and spaces are tokens too. A frequently cited approximation for English is around three-quarters of a word per token, or four characters.

Why tokenisation is not neutral:

Languages differ substantially. Tokenisers trained predominantly on English split other languages into more tokens for the same meaning — so the same text costs more and consumes more context in some languages than others, which is a genuine inequity.

Code, numbers and unusual strings tokenise inefficiently.

It explains specific failures. Character-level tasks — counting letters, reversing words, handling spelling — are hard because the model does not see characters. This is why questions about letter counts produce surprising errors from systems that handle complex reasoning well.

The context window. The maximum tokens the model can attend to in one request, including both input and output. Everything the model can use must fit: system instructions, conversation history, supplied documents, and the response being generated.

What happens at the limit. Older material is truncated or summarised. This is why a long conversation appears to forget earlier detail — it is no longer present, rather than being forgotten.

Why bigger windows are not simply better:

Cost and latency rise with length.

Attention degrades across long contexts. Research has repeatedly found models perform worse on information in the middle of a long context than at the beginning or end — sometimes called "lost in the middle" — so filling a large window is not equivalent to the model using it all well.

More context can dilute the relevant material.

The practical implication: retrieving the relevant portion and supplying that is usually better than supplying everything, even where everything fits.

Related Questions