Tokens, Context Windows and Why the Model Forgot Your Name

· 2 min read · Syed Omar Faruk Towaha
Tokens, Context Windows and Why the Model Forgot Your Name

You spent an hour with a chatbot planning a project. It knew your stack, your deadline, your manager's name and the fact that you hate YAML. Then, around message 140, it suggested a YAML config and asked what language you're using. It's like a friend who listened intently all evening and then called you by your cousin's name.

The explanation is two words: context window.

First, tokens

Models don't read words; they read tokens, which are chunks of text. In English, a token is roughly three-quarters of a word: "unbelievable" might be split into "un", "believ", "able." Spaces, punctuation and code symbols are tokens too.

Token costs
The same meaning can cost very different numbers of tokens.

This has some practical consequences:

The context window

The context window is the maximum number of tokens a model can look at in one go: system instructions, the whole conversation so far, any documents you pasted, and the answer it's writing. Everything the model "remembers" in a chat is just text inside that window. There's no separate memory unless the product adds one.

When a conversation gets longer than the window, the application has to do something: drop the oldest messages, or summarise them. Either way, details from the start can fall out.

Long chats
Early details get trimmed or compressed so new messages fit.

Even with today's very large windows, there's a softer problem: models can pay less attention to details buried in the middle of a huge context. Having room for 500 pages doesn't mean every page gets equal focus.

How to work with it, not against it

  1. Start fresh for new tasks. A clean chat with a good summary beats a 200-message thread.
  2. Restate key constraints when they matter: "Reminder: Python 3.12, no YAML, deadline Friday."
  3. Ask for a running summary in long sessions, then use it to start a new conversation.
  4. Paste only what's relevant. Ten focused paragraphs beat an entire wiki.
  5. Put instructions at the start and repeat critical ones near the end of very long prompts.

The human parallel

Honestly, the context window is the most human thing about these models. Talk to me for two hours about a complicated project and I, too, will forget what we decided in the first ten minutes. The difference is that I'll pretend I remember. The model, at least, will just ask again.

// related

// prefer the terminal?

Open the terminal blog and type read tokens-context-windows-why-the-model-forgot.