Independent project. Not a U.S. government website.

USASI

Explainer · AI basics

Tokens and context windows

Why do AI tools count tokens, and what happens when a conversation gets too long?

Beginner6 min readReviewed Oct 8, 2026

General information, not legal or professional advice. All explainers

How this page was made
  • Researched and written with AI assistance from primary sources, which are listed at the end with the date they were read.
  • Source-checked: a separate AI fact-check pass compared each sentence with its source and corrected what did not match (fact-check report (external site: github.com)).
  • Automated checks passed: links, structure, and formatting are validated before every publish.
  • Not individually reviewed by a person before publication. What these review levels mean

Key takeaways

  • A token is a chunk of text, often a word or part of one; different models can use different tokenizers, so the same text can count differently.
  • In a chat, the context window must hold the system prompt, the whole conversation, attached files, tool definitions, and the reply being generated.
  • When a conversation outgrows the window, APIs may refuse or truncate, and some tools drop or summarize older turns, so early material may drop out of view.
On this page

A token is a small chunk of text, often a word or a piece of one, and it is the unit a language model reads and writes. AI tools count tokens because that is what the model actually processes, so input limits, reply limits, and usage bills are all measured in tokens. Different models can use different tokenizers, so the same sentence can produce different counts. The context window is the most tokens a model can handle in one request, and in a chat it must hold the instructions, the conversation so far, any attached files, and the reply. When a conversation outgrows it, an API may refuse the request or cut the reply short, and some tools drop or summarize older material, so the model may no longer see what was said early on.

What a token is

OpenAI's tiktoken README (external site: github.com) explains that language models do not see text the way people do; they see a sequence of numbers called tokens. A tokenizer converts text into those numbers and back. OpenAI's token counting cookbook (external site: github.com) shows "tiktoken is great!" split into six tokens by the o200k_base encoding, which it lists for GPT-4o: "t", "ikt", "oken", " is", " great", and "!". The space travels with the word after it. Common pieces get their own tokens; the README says "encoding" will often split into "encod" and "ing".

Google's Gemini token guide (external site: ai.google.dev) says a Gemini token is about four characters, and 100 tokens about 60 to 80 English words. Google's machine learning course (external site: developers.google.com) adds that tokenization is language-specific, so the ratio differs across languages. Images and audio become tokens too: the Gemini guide counts an image 384 pixels or smaller on both sides as 258 tokens, and audio at 32 tokens per second.

Why counts differ between models

Each tokenizer has its own vocabulary, the fixed list of pieces it can produce. The cookbook runs several OpenAI encodings on the same text. "antidisestablishmentarianism" is 5 tokens in the older r50k_base encoding and 6 in cl100k_base and o200k_base, split at different points. "2 + 2 = 4" is 5 tokens in r50k_base and 7 in the other two. A short Japanese phrase falls from 14 tokens in r50k_base to 9 in cl100k_base and 8 in o200k_base. tiktoken's model table (external site: github.com) maps GPT-4 to cl100k_base, GPT-4o to o200k_base, and gpt-oss to o200k_harmony.

Vocabulary sizes differ too. The gpt-oss model card (external site: arxiv.org) gives 201,088 tokens for o200k_harmony, while the configuration file (external site: huggingface.co) for Ai2's Olmo 3 7B sets 100,278. Counts can shift within one provider: Anthropic's token counting page (external site: platform.claude.com) says Claude models from version 4.7 onward use a newer tokenizer that produces roughly 30 percent more tokens for the same text, and advises recounting for the model you plan to use. A count holds only for the tokenizer it was measured with.

Input limits, output limits, and the context window

The context window is the overall limit. OpenAI's conversation state guide (external site: developers.openai.com) defines it as the maximum number of tokens in a single request, counting input, output, and reasoning tokens; Google's guide calls it the combined limit of input and output tokens. A reply also has its own cap. OpenAI's example, gpt-4o-2024-08-06, has a 128k-token window but generates at most 16,384 output tokens. Anthropic's context window page (external site: platform.claude.com) lists a 1M-token window for many of its models, with up to 128k output tokens per request on those, and 200k for the rest.

Open models state limits in the model card or configuration. The Olmo 3 7B card (external site: huggingface.co) lists a context length of 65,536 tokens, and the gpt-oss card says the context length of its dense attention layers was extended to 131,072 tokens. Software running a model locally may use less than the maximum: Ollama's context length page (external site: docs.ollama.com) says its default depends on graphics memory (VRAM), from 4k tokens with less than 24 GiB to 256k with 48 GiB or more, and that a larger context needs more memory.

What fills the window in a chat

Anthropic's page lists what counts: the system prompt (instructions set by the app or developer), every message in the conversation, including tool results, images, and documents, the definitions of any tools the model can use, and the reply being generated, including the model's extended thinking. Each turn's input is the whole history plus the new message. Google's guide likewise counts system instructions as input, and counts tools too.

More is not automatically better. Anthropic says accuracy and recall degrade as the token count grows, which it calls context rot. Google's long context guide (external site: ai.google.dev) advises leaving out tokens the model does not need and, for long inputs, putting your question at the end.

When a conversation gets too long

What happens depends on the tool:

  • The request is refused. Anthropic's API returns a "prompt is too long" error when the input alone exceeds the window.
  • The reply is cut off. OpenAI's guide warns that exceeding the window might result in truncated outputs. On newer Claude models, a reply that reaches the limit stops with a stop reason saying so.
  • Old turns are dropped. Anthropic says chat interfaces such as claude.ai can manage the window on a rolling first-in, first-out basis, so the oldest material leaves first.
  • Old turns are summarized. Anthropic's server-side compaction summarizes earlier parts of a conversation so it can continue past the limit. OpenAI's compaction guide (external site: developers.openai.com) describes a compaction step that runs when the token count crosses a threshold the developer sets, carrying forward key prior state.

Google's long context guide names similar strategies: dropping old messages, summarizing, retrieval, and filtering prompts. Either way, the model cannot consult what was dropped, and a summary keeps only part of the detail.

Billing and counting tokens

Hosted APIs bill by tokens: Google's guide says the cost of a Gemini API call depends in part on input and output tokens, and OpenAI's cookbook says usage is priced by token. Long chats cost more per turn because earlier turns are counted again: OpenAI says all previous input tokens in a chained conversation are billed as input tokens. Anthropic bills thinking tokens as output tokens. USASI does not publish prices; check the provider's pricing page and note the date.

To count before you send:

  • tiktoken counts OpenAI tokens locally for plain text. OpenAI's token counting guide (external site: developers.openai.com) says local tokenizers do not support images or files and that tools add tokens that are hard to count locally. It offers an endpoint that returns the exact input count, including formatting tokens for message roles.
  • Anthropic's token counting API accepts the same input as a message, including system prompt, tools, images, and PDFs, and returns an estimate that may differ slightly from actual use.
  • Gemini's countTokens returns the input count before sending, and each response reports input, output, and other counts in its usage data.
import tiktoken
enc = tiktoken.get_encoding("o200k_base")
print(len(enc.encode("tiktoken is great!")))  # 6

Worked example: fitting a transcript

Ray is a fictional reader invented for this page. Ray wants the decisions from a 20,000-word meeting transcript, using a model run locally with Ollama on a graphics card with 16 GiB of memory.

  1. Estimate the size. At 60 to 80 words per 100 tokens, 20,000 words is roughly 25,000 to 33,000 tokens. The exact figure depends on the tokenizer.
  2. Find the limit that applies. The model card gives the model's maximum, but Ollama's default with less than 24 GiB of VRAM is 4k tokens, far too small for the transcript.
  3. Leave room. The window must also hold the instructions and the reply, so Ray sets a context length well above the estimate, then runs ollama ps to confirm the CONTEXT value and that the model still runs fully on the GPU.
  4. Or split the job. If the larger setting does not fit in memory, Ray asks about one section of the transcript at a time and combines the lists.
  5. Repeat key instructions. In follow-ups, Ray restates the format wanted rather than relying on the first message.

On a hosted tool, the same steps apply, with the provider's counter and documented limits replacing Ollama's settings.

What you can do next

Sources

All read on October 8, 2026.

Check your understanding

Three quick questions, answered from this page. Nothing you choose is saved or sent anywhere.

  1. 1.Why can the same sentence produce a different token count in two different models?
  2. 2.In a chat, what has to fit inside the context window?
  3. 3.If a chat tool drops or summarizes older turns to stay within the context window, what follows, according to the page?

0 of 3 answered.

Keep learning

Part of New to AI.

Support Us

Help keep USASI useful.

Find the catalog useful? Leave an optional tip to support its upkeep. Tips never affect listings, coverage, or openness assessments.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project