Key takeaways
- RAG searches your documents for passages related to a question and adds them to the prompt, so the model can answer from them without retraining.
- Retrieval can miss the right passage or return conflicting versions, and the model can still misread them, so test retrieval and answers separately.
- An index is another copy of your documents: check a hosted provider's retention terms, or run embeddings and search on your own hardware.
On this page
- Why a model alone does not know your documents
- Embeddings: meaning as numbers
- Chunks and vector search
- Putting passages in the prompt, with citations
- Where it goes wrong, and published fixes
- Privacy when you index private documents
- How to evaluate a RAG setup
- Worked example: a handbook assistant
- What you can do next
- Check your understanding
A language model learns from training data collected up to a cutoff date, so on its own it does not know your private documents or events after that date. Retrieval-augmented generation (RAG) works around this without retraining. Before the model answers, a search step finds passages in your documents that look relevant to the question and adds them to the prompt. The model then answers from those passages and can be asked to cite them. The search usually relies on embeddings, lists of numbers that place texts with similar meanings close together. RAG can still miss the right passage, mix up conflicting ones, or misread them, so it needs testing, and its index is a new copy of your documents that must be protected.
Why a model alone does not know your documents
A model's knowledge stops at a date. Anthropic's models overview (external site: platform.claude.com) lists a "reliable knowledge cutoff" of June 2026 for its current Claude models, and OpenAI's embeddings guide (external site: developers.openai.com) says its text-embedding-3 models "lack knowledge of events that occurred after September 2021." Nor can a model have learned from documents it never saw, such as an internal handbook updated last week.
A 2020 paper by Lewis and colleagues, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (external site: arxiv.org), introduced models it called RAG. For models that store knowledge in their parameters (weights), it says "providing provenance for their decisions and updating their world knowledge remain open research problems." The paper paired a text generator with a searchable index of Wikipedia and showed the index could be "hot-swapped to update the model without requiring any retraining." The term is now used more broadly: Anthropic's Contextual Retrieval (external site: anthropic.com) post describes RAG as a method that "retrieves relevant information from a knowledge base and appends it to the user's prompt."
RAG is not always needed. The same post says a knowledge base under 200,000 tokens, "about 500 pages of material," can simply go into the prompt (see tokens and context windows).
Embeddings: meaning as numbers
OpenAI defines an embedding as "a vector (list) of floating point numbers," where "the distance between two vectors measures their relatedness." Its retrieval guide (external site: developers.openai.com) gives an example: for "When did we go to the moon?", the sentence "The first lunar landing occurred in July of 1969." shares no words with the query yet scores highest on semantic similarity.
Embeddings come from a dedicated embedding model. The EmbeddingGemma 300M model card (external site: ai.google.dev) describes a 300M-parameter model that reads up to 2K tokens and outputs 768 numbers (or 512, 256, or 128), with an "on-device focus" suited to phones, laptops, and desktops. The card gives different prompts for queries and documents, and the Sentence Transformers documentation (external site: sbert.net) recommends separate encode_query and encode_document calls for "asymmetric" search, where a short question must find a longer passage. Both must land in the same vector space, so use one embedding model for both.
Chunks and vector search
Documents are first split into chunks, so search can return a passage instead of a whole file. Anthropic says chunks are "usually no more than a few hundred tokens"; the RAG paper split Wikipedia into "disjoint 100-word chunks," 21 million in all. OpenAI's vector stores chunk files automatically, with settings for chunk size and overlap (text shared by neighboring chunks).
In vector search, each chunk's embedding goes into an index; at question time the question is embedded and the closest chunks come back. FAISS, which the RAG paper used, is a library for "efficient similarity search and clustering of dense vectors," per its README (external site: github.com). Its index types trade search time against search quality and memory. Sentence Transformers explains that approximate search speeds up collections of millions of entries, but "results are not necessarily exact," so some close matches can be missed.
Exact words still matter. Anthropic notes that embedding models "can miss crucial exact matches," such as an error code like "TS-999," which BM25, a ranking method based on matching words, can catch. OpenAI's file search (external site: developers.openai.com) tool uses "semantic and keyword search."
Putting passages in the prompt, with citations
Anthropic's prompting guidance (external site: platform.claude.com) suggests placing long documents above the question, labeling each with its source, and asking the model to "quote relevant parts of the documents first." Anthropic's citations feature (external site: platform.claude.com) returns "valid pointers to the provided documents," and OpenAI's file search marks answers with file citations. A valid pointer shows where text came from, not that the answer reads it correctly, so readers should still open the source.
Where it goes wrong, and published fixes
- The right passage is not retrieved. Anthropic says traditional RAG can "remove context when encoding information": a chunk saying a company's revenue grew may not name the company or quarter. Its Contextual Retrieval technique has a model write a short context, usually 50 to 100 tokens, for each chunk before both embedding and BM25 indexing; the post then adds a reranking model that rescores 150 candidates and keeps 20. Anthropic reports fewer failed retrievals in its tests on code, fiction, and research papers.
- Passages conflict. Old and new versions of a policy can both match. Remove superseded documents, or record dates and filter on them; OpenAI's attribute filtering can restrict a search "to a specific date range."
- The model misreads or is misled. More chunks raise the chance of including the right one, but Anthropic warns "more information can be distracting for models." OWASP's prompt injection entry (external site: genai.owasp.org) says RAG does not fully mitigate injection and describes an attacker editing a document in a RAG repository to mislead answers (see how AI models use tools).
Privacy when you index private documents
If the index is hosted, read the provider's terms for that feature: OpenAI's data controls page (external site: developers.openai.com), for example, says data sent to its vector stores endpoint is not used for training, that its application state is kept "until deleted," and that the endpoint is not eligible for Zero Data Retention. The retrieval guide shows how to give a vector store an expiration policy. A local setup keeps the copy on your hardware: Google says the first version of EmbeddingGemma (external site: ai.google.dev) generates embeddings directly on your hardware and "works without internet connection," and FAISS is a library you run yourself.
Access is a separate question: retrieval returns whatever matches unless the application limits it, so a shared index can surface one team's files in another team's answers. File attributes, such as the "department" tag in OpenAI's retrieval examples, or separate indexes can keep them apart. Terms change; this is general information, not legal advice. See AI and your data.
How to evaluate a RAG setup
Test retrieval and answers separately.
- Build a question set. OpenAI's evaluation guide (external site: developers.openai.com) suggests mixing production data from users, correct answers written by domain experts, and historical logs, and including typical, edge, and adversarial cases. Note which passage answers each question.
- Score retrieval. Recall at k is the share of correct passages found in the top k results; Anthropic measured "1 minus recall@20," the share missed. Sentence Transformers' InformationRetrievalEvaluator (external site: sbert.net) computes recall at k from queries, documents, and known correct matches.
- Score answers. OpenAI's example for questions over documents sets targets for context recall, context precision, and positively rated answers. OWASP's prompt injection entry mentions checking groundedness, meaning whether the answer is supported by the retrieved passages.
- Repeat after every change. OpenAI warns against "vibe-based evals."
Worked example: a handbook assistant
Priya is a fictional reader invented for this page. Priya runs operations at a 40-person nonprofit and wants staff to ask questions about the employee handbook and years of policy memos, some of which replace earlier ones.
- Is RAG needed? The 80-page handbook is well under the roughly 500 pages Anthropic's post says can fit in a prompt, so it can go in directly. The memos add hundreds of pages and conflicting versions, which points toward retrieval.
- Where does the index live? The memos name staff. Priya compares a hosted vector store, checking retention and setting an expiration, with a local embedding model and FAISS index on an office machine.
- How to chunk and search? By section heading, with each memo's title and date attached, and with keyword search alongside embeddings because staff ask about form numbers.
- How to handle conflicts? Superseded memos leave the index.
- Does it work? Priya writes 40 real questions with the correct memo for each, checks recall at 5, reads every answer and its citations, and re-runs the set after each change.
Priya's outcome: handbook in the prompt now, retrieval for memos only if the test questions show the need. Stricter rules on staff data could change that.
What you can do next
- See the catalog records for FAISS and EmbeddingGemma 300M, and browse the data and datasets hub.
- Read how AI models use tools, since document search is often offered as a tool, and tokens and context windows.
- Read hosted or local before choosing where your index lives.
- Look up context window, weights, and self-hosting in the glossary.
Sources
All read on October 8, 2026.
- Lewis et al.: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (external site: arxiv.org) (arXiv 2005.11401; abstract page and v4 PDF (external site: arxiv.org))
- Anthropic: Introducing Contextual Retrieval (external site: anthropic.com) (September 19, 2024), Models overview (external site: platform.claude.com), Prompting best practices: long context prompting (external site: platform.claude.com), Citations (external site: platform.claude.com)
- OpenAI: Vector embeddings (external site: developers.openai.com), Retrieval (external site: developers.openai.com), File search (external site: developers.openai.com), Data controls in the OpenAI platform (external site: developers.openai.com), Evaluation best practices (external site: developers.openai.com)
- Google: EmbeddingGemma model card (external site: ai.google.dev), EmbeddingGemma overview (external site: ai.google.dev)
- Meta: Faiss README (external site: github.com)
- Sentence Transformers: Semantic Search (external site: sbert.net), Evaluation reference (external site: sbert.net)
- OWASP Gen AI Security Project: LLM01:2025 Prompt Injection (external site: genai.owasp.org)
Check your understanding
Three quick questions, answered from this page. Nothing you choose is saved or sent anywhere.
0 of 3 answered.