Independent project. Not a U.S. government website.

USASI

Explainer · AI basics

How large language models work

What happens between typing a prompt and getting an answer?

Beginner6 min readReviewed Oct 8, 2026

General information, not legal or professional advice. All explainers

How this page was made
  • Researched and written with AI assistance from primary sources, which are listed at the end with the date they were read.
  • Source-checked: a separate AI fact-check pass compared each sentence with its source and corrected what did not match (fact-check report (external site: github.com)).
  • Automated checks passed: links, structure, and formatting are validated before every publish.
  • Not individually reviewed by a person before publication. What these review levels mean

Key takeaways

  • A model is a file of learned numbers, called weights, plus an architecture; training adjusts the weights so the model gets better at predicting the next token.
  • A reply is built one token at a time: the model scores every possible next token, a setting such as temperature shapes the pick, and the loop repeats.
  • Fluent output can still be wrong; Anthropic's guide says its techniques reduce hallucinations but do not eliminate them, so validate critical information.
On this page

A large language model is a file of learned numbers, called weights, plus an architecture: a fixed recipe for the calculations that turn those numbers and your text into a prediction. In training, the model was shown a very large amount of text, and its weights were adjusted again and again so it got better at predicting the next token, a word or piece of a word. When you send a prompt, the model scores every possible next token, one is chosen and added to the text, and the cycle repeats until the reply is finished. Chat assistants are base models given further training to follow instructions. Because the model produces likely text rather than looking facts up, a fluent answer can still be wrong, as developers say in their own documentation.

What a model is made of

Ai2's Olmo 3 7B (external site: huggingface.co) has every file in public view. Its file list (external site: huggingface.co) includes three large .safetensors files holding the weights, the numbers learned in training; an index file (external site: huggingface.co) counts 7,298,011,136 of them. A short config.json (external site: huggingface.co) describes the architecture: how many layers of calculation are stacked (32), how many different tokens the model knows (100,278), and how much text it can handle at once (65,536 tokens). The folder also holds the tokenizer files that turn text into tokens.

The model card describes Olmo 3 as a Transformer-style autoregressive language model. The Transformer was proposed in the 2017 paper Attention Is All You Need (external site: arxiv.org) as an architecture built entirely on a mechanism called attention. Google's Machine Learning Crash Course (external site: developers.google.com) explains self-attention as asking, for each token, how much every other token affects its meaning: in "The animal didn't cross the street because it was too tired," attention links "it" to "animal" more than to "street." Autoregressive means the model produces one piece at a time and feeds each piece back in as input for the next.

What training does

Models read tokens, not letters. Google's introduction to language models (external site: developers.google.com) says a token can be a word, part of a word, or a single character; "unwatched" might become "un", "watch", and "ed". Tokens and context windows covers them in depth.

Training needs a very large amount of text: Olmo 3 7B's first training stage used 5.93 trillion tokens, according to its model card. The basic task is a guessing game: the model predicts the next token, the guess is compared with the token that actually comes next, and the weights are adjusted so the right answer becomes slightly more likely. Hugging Face's guide to causal language modeling (external site: huggingface.co) describes this kind of model as predicting the next token while seeing only the tokens before it, and Google's course explains that the model's errors guide how an algorithm called backpropagation updates the parameter values. Over trillions of tokens, the weights come to capture patterns that, the course says, include a good deal about grammar, words, and idioms. This first phase is called pretraining. Its data stops at a cutoff date; the Olmo 3 card gives December 2024.

From prompt to answer

Using a trained model is called inference. A prompt does not change the weights, as Google's course page on prompting (external site: developers.google.com) notes. A reply is built in a loop:

  1. Tokenize. The prompt is split into tokens, and each token becomes a number.
  2. Score. The model calculates a probability for every token it knows, which the 2017 paper calls next-token probabilities. In Google's example, to fill the blank in "When I hear rain on my roof, I ___ in my kitchen," a model might rate "cook soup" at 9.4% and "nap" at 2.5%.
  3. Choose. Greedy decoding always takes the top choice; Hugging Face's generation guide (external site: huggingface.co) notes that it starts repeating itself on longer outputs. Sampling picks at random, weighted by the probabilities. Google's Gemini prompting guide (external site: ai.google.dev) says temperature controls the degree of randomness: lower values suit less open-ended answers, and higher values can give more diverse or creative results. Top-K and top-P settings narrow the choice to the most likely candidates.
  4. Repeat. The chosen token is added, and the longer text goes back in. Hugging Face's text generation guide (external site: huggingface.co) says this continues until the model produces an end-of-sequence token or reaches a set length.
  5. Detokenize. The tokens are turned back into readable text.

Because of the choosing step, the same prompt can produce different replies; OpenAI's text generation guide (external site: developers.openai.com) calls model output non-deterministic. The Olmo 3 7B Instruct card (external site: huggingface.co) suggests a temperature of 0.6 and a top-p of 0.95, while Google recommends keeping the defaults for Gemini 3 models.

What the model can see at once

The context window is the text a model can take into account while writing a reply. Anthropic's documentation (external site: platform.claude.com) describes it as all the text the model can reference when generating a response, including the response itself, and calls it a working memory, separate from the training data. Every model has a limit in tokens. In a conversation, each new turn includes the history so far, so long chats fill the window. Tokens and context windows explains what happens when it runs out.

Why chat assistants are trained further

A pretrained model, often called a base or foundation model, has learned to continue text. Google's course calls instruction tuning an optional step that can improve a model's ability to follow instructions. Why Language Models Hallucinate (external site: arxiv.org), a 2025 paper by researchers at OpenAI and Georgia Tech, describes current training as two stages, pretraining and post-training, with post-training refining the base model. Olmo shows the stages publicly: Olmo 3 7B Instruct starts from the Olmo 3 7B base model and adds supervised fine-tuning, direct preference optimization, and reinforcement learning, each with its own dataset. Its card shows the chat template, which wraps each message in markers such as <|im_start|>user, so a conversation still reaches the model as one sequence of tokens to continue. Pretraining and post-training covers these stages.

Why fluent answers can still be wrong

Fluency is not evidence of accuracy. The hallucination paper describes hallucinations as plausible but incorrect statements produced instead of an admission of uncertainty. It argues that pretraining leads to some errors even if the training data contained none, including on arbitrary facts with no pattern to learn, such as a birthday that appears only once in the data. It also argues that errors survive post-training partly because most evaluations give points for a correct guess and none for "I don't know," which rewards guessing.

Anthropic's guide to reducing hallucinations (external site: platform.claude.com) suggests explicitly allowing the model to say it does not know, asking it to extract word-for-word quotes from a long document before answering, and having it support each claim with a quote. It says these techniques reduce hallucinations but do not eliminate them, and advises validating critical information. The Olmo 3 cards warn that many statements from Olmo or any LLM are often inaccurate and should be verified.

Worked example: checking an answer

Dana is a fictional reader invented for this page. Dana asks a chat assistant what year the town's public library opened, then asks again in a new chat and gets a different year.

  1. Why two answers? Each reply was sampled token by token, so a different year can come out on a different run. Anthropic's guide notes that inconsistent outputs to the same prompt could indicate hallucinations.
  2. Is either one right? A small library's opening date may appear rarely in training data, the kind of arbitrary fact the paper links to hallucination.
  3. Give the model the source. Dana pastes the library's own history page into the chat, putting it in the context window, and asks the assistant to quote the sentence that gives the opening year, or to say the page does not give one.
  4. Check the quote. Dana finds that sentence on the original page before using the year.
  5. Outcome. The year comes from the library's page; the assistant only helped locate it. The source counts, not how confident the answer sounded.

What you can do next

Sources

All read on October 8, 2026.

Check your understanding

Three quick questions, answered from this page. Nothing you choose is saved or sent anywhere.

  1. 1.What does training change in a large language model?
  2. 2.Why can the same prompt give different replies on different runs?
  3. 3.In the page's worked example, where does the library opening year that Dana finally uses come from?

0 of 3 answered.

Keep learning

Part of New to AI.

Support Us

Help keep USASI useful.

Find the catalog useful? Leave an optional tip to support its upkeep. Tips never affect listings, coverage, or openness assessments.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project