Retrieval-augmented generation
RAG splits documents into chunks, turns them into embeddings and stores them in a vector index. At question time, the most similar chunks are retrieved and placed in the prompt. Quality depends on chunking, retrieval and whether the model is told to cite and admit gaps.
MAKE IT CONCRETE
A museum app retrieves the three most relevant exhibit notes before answering “Who carved this?”.
Most RAG failures are retrieval failures.
CHECK YOURSELF
