Tech●●●●●Difficulty 3 of 5

How does an AI agent remember, and look things up before it answers?

A model only 'sees' what fits in its window. Everything else has to be summarized, retrieved or handed back to it.

▶ Start the story

A model remembers only what is inside its context window, and for anything beyond that it needs a way to look things up. The context window is the maximum number of tokens the model can use at one time when generating an output. In practical terms it is the material the model can see while answering; anything outside the window is not directly available unless it is summarized, retrieved or provided again.

The windows have grown. The small GPT-2 model had a context window of only about 1,000 tokens, while by the mid-2020s long-context systems reported windows of hundreds of thousands to millions of tokens. But bigger is not simply better. A paper called Lost in the Middle found that performance is often highest when the relevant information is at the beginning or end of the input, and degrades significantly when it sits in the middle, even for long-context models. Longer windows also bring slower responses and higher costs.

~1,000 tokens

The small GPT-2 model's context window, against windows of up to millions of tokens in the mid-2020s

So agents also look things up. Retrieval-augmented generation, or RAG, lets a model retrieve and use new information from external data sources: it first refers to a specified set of documents, then answers. Typically the documents are turned into embeddings, lists of numbers, and stored in a vector database; given a question, a retriever selects the most relevant ones to add to the prompt.

RAG is no cure-all. It lets users check the cited sources, but it does not prevent hallucinations.

Quiz me

0/3

  1. 1.What happens to information that falls outside a model's context window?
  2. 2.What did the Lost in the Middle paper find?
  3. 3.What is the basic idea of retrieval-augmented generation?

Recap

The window is the desk; retrieval fetches the book; neither guarantees a correct answer.

💡 A trick to remember it · Window is the desk, retrieval is the library run: you can only work with what lands on the desk.

Surprising fact · Models can be worse at using information from the middle of a long input than from its start or end.

Sources (5)

No source, no claim. Every fact in this lesson (16 claims) cites at least one of these.

  1. [1]Context window · Wikipedia
  2. [2]Lost in the Middle: How Language Models Use Long Contexts · arXiv (Liu et al.)
  3. [3]Retrieval-augmented generation · Wikipedia
  4. [4]Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks · arXiv (Lewis et al.)
  5. [5]Large language model · Wikipedia
More lessons in 💻 Tech (3) See all tech lessons →

One more light on your map.

Get one lesson like this every day, about the things you love. Free, in two or five minutes.

Get the share card for this lesson ↗