Python course Β· Module 11: RAG and Multi-Agent Systems
RAG - Extended Memory for AI
In this lesson6
Imagine a guide who knows the savanna only from books printed three years ago. You ask about a new trail to the watering hole, and he confidently describes a path that no longer exists. That is how a language model behaves without access to current knowledge. Welcome to the penultimate stage of Python Safari! In this location we will teach the model to check its field notebook before it says anything.
Just as large mammals evolved bigger brains and better memory, language models evolve thanks to RAG (Retrieval-Augmented Generation). It is a technique that gives AI access to external knowledge without retraining the model.
What you will learn
- recognise the LLM limitations that RAG solves,
- describe the three pipeline steps: retrieve, augment, generate,
- build the simplest RAG in plain Python with OpenAI and NumPy,
- split long documents into pieces (chunking),
- measure retrieval quality with precision and recall.
The Problem with Traditional LLMs
Language models have fundamental limitations, much like animals adapted only to a single environment:
Limitations of Base LLMs:
- Knowledge cutoff - knowledge frozen at training time
- Hallucinations - generating false information
- Lack of company context - they don't know your documents
- High fine-tuning costs - adapting the model is expensive
The first limitation matters most: the model knows nothing about what happened after training ended, and it has never seen your private documents. RAG does not change the model itself. It changes what the model gets to read before answering.
What is RAG?
RAG is an architecture that combines information retrieval with text generation. The diagram below follows a question from the user all the way to the final answer:
1βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
2β RAG Pipeline β
3βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
4β β
5β User question β
6β β β
7β βΌ β
8β βββββββββββββββββββ β
9β β Embedding β β Convert to vector β
10β ββββββββββ¬βββββββββ β
11β β β
12β βΌ β
13β βββββββββββββββββββ βββββββββββββββββββ β
14β β Vector Search ββββββΆβ Vector Database β β
15β ββββββββββ¬βββββββββ βββββββββββββββββββ β
16β β β
17β βΌ β
18β βββββββββββββββββββ β
19β β Retrieved Docs β β Top-K similar documents β
20β ββββββββββ¬βββββββββ β
21β β β
22β βΌ β
23β βββββββββββββββββββ β
24β β LLM + Context β β Generate with context β
25β ββββββββββ¬βββββββββ β
26β β β
27β βΌ β
28β Answer β
29βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββRead it from the top. We turn the question into a vector of numbers (an embedding), search a vector database for passages with a similar meaning, take the best few (Top-K) and paste them into the prompt. Only then does the LLM generate the answer. The model does not change, only its "cheat sheet" for this one question does.
Basic RAG Implementation
Let's build the whole pipeline in a few dozen lines. We start with the OpenAI client, NumPy and a tiny knowledge base, a list of sentences that plays the role of our field notebook:
1from openai import OpenAI
2import numpy as np
3
4client = OpenAI()
5
6# Simple knowledge base
7knowledge_base = [
8 "Python Safari is a Python programming course with 12 modules.",
9 "RAG combines document retrieval with text generation by LLMs.",
10 "Embeddings are vector representations of text in semantic space.",
11 "Vector databases store and search embeddings efficiently.",
12 "LlamaIndex is a framework for building RAG applications.",
13 "CrewAI enables the creation of multi-agent systems."
14]OpenAI() reads the API key from the OPENAI_API_KEY environment variable, so we never type it into the code. The knowledge base itself is an ordinary list of strings, nothing magical.
Now two helper functions. The first sends text to the text-embedding-3-small model and returns a vector, the second computes cosine similarity, which tells how much two vectors "point" in the same direction:
1def get_embedding(text: str) -> list[float]:
2 """Generates an embedding for the text."""
3 response = client.embeddings.create(
4 model="text-embedding-3-small",
5 input=text
6 )
7 return response.data[0].embedding
8
9def cosine_similarity(vec1: list, vec2: list) -> float:
10 """Computes cosine similarity of two vectors."""
11 vec1, vec2 = np.array(vec1), np.array(vec2)
12 return np.dot(vec1, vec2) / (np.linalg.norm(vec1) * np.linalg.norm(vec2))Cosine similarity ranges from -1 to 1, and the closer it gets to 1, the closer the meanings of the texts. We will take it apart in the next lesson.
The third step is the search function. For every document we compute its similarity to the question, sort in descending order and return the best top_k:
1def retrieve_relevant_docs(query: str, docs: list, top_k: int = 3) -> list[str]:
2 """Retrieves the most similar documents."""
3 query_embedding = get_embedding(query)
4
5 similarities = []
6 for doc in docs:
7 doc_embedding = get_embedding(doc)
8 similarity = cosine_similarity(query_embedding, doc_embedding)
9 similarities.append((doc, similarity))
10
11 # Sort by similarity
12 similarities.sort(key=lambda x: x[1], reverse=True)
13
14 return [doc for doc, _ in similarities[:top_k]]Notice one weakness: this function computes the embedding of every document on every question. With six sentences it does not hurt, with thousands of documents it would be slow and expensive. That is why in later lessons we compute document embeddings once and keep them in a vector database.
Finally we assemble everything in rag_query, where you can see the three letters of the acronym: Retrieve, Augment, Generate:
1def rag_query(question: str) -> str:
2 """Full RAG pipeline."""
3 # 1. Retrieve - find relevant documents
4 relevant_docs = retrieve_relevant_docs(question, knowledge_base)
5
6 # 2. Augment - build prompt with context
7 context = "\n".join(relevant_docs)
8
9 prompt = f"""Answer the question using ONLY the context below.
10If you cannot answer based on the context, say "I don't know".
11
12Context:
13{context}
14
15Question: {question}
16"""
17
18 # 3. Generate - produce the answer
19 response = client.chat.completions.create(
20 model="gpt-4o-mini",
21 messages=[{"role": "user", "content": prompt}]
22 )
23
24 return response.choices[0].message.content
25
26# Test
27answer = rag_query("What is RAG?")
28print(answer)The key part is the instruction in the prompt: "using ONLY the context below", plus permission to answer "I don't know". That is what limits hallucinations. The gpt-4o-mini model is exactly the same as a moment ago, only the prompt changed.
Chunking - Splitting Documents
Long documents need to be split into smaller pieces (chunks), because an embedding of a whole book blurs the meaning and a prompt has a limited length. The simplest method cuts the text every fixed number of characters, with an overlap:
1from typing import Generator
2
3def simple_chunker(text: str, chunk_size: int = 500, overlap: int = 100) -> Generator[str, None, None]:
4 """Simple function for splitting text into chunks."""
5 start = 0
6 while start < len(text):
7 end = start + chunk_size
8 chunk = text[start:end]
9 yield chunk
10 start = end - overlap # Overlap to preserve contextThe overlap makes sure a sentence cut at a boundary lands in full in at least one chunk. Just keep overlap smaller than chunk_size, otherwise the loop never ends. The last chunk can also be very short, because it is only the tail of the text.
Splitting by sentences preserves meaning better. A regular expression cuts the text after a full stop, exclamation mark or question mark, and then we glue a few sentences together:
1def sentence_chunker(text: str, sentences_per_chunk: int = 5) -> list[str]:
2 """Splits text into chunks by sentences."""
3 import re
4
5 # Simple sentence segmentation
6 sentences = re.split(r'(?<=[.!?])\s+', text)
7
8 chunks = []
9 for i in range(0, len(sentences), sentences_per_chunk):
10 chunk = ' '.join(sentences[i:i + sentences_per_chunk])
11 chunks.append(chunk)
12
13 return chunks
14
15# Usage example
16document = """
17Python is a high-level programming language. It was created by
18Guido van Rossum. It is very popular in Data Science and Machine Learning.
19RAG is a technique combining retrieval with generation. It allows LLMs
20to use external knowledge sources. It is crucial for enterprise AI.
21"""
22
23chunks = sentence_chunker(document, sentences_per_chunk=2)
24for i, chunk in enumerate(chunks):
25 print(f"Chunk {i}: {chunk[:50]}...")Every printed chunk now contains whole sentences, and the document content has not changed, only its packaging has. I recommend starting with sentence splitting, because chunks cut in the middle of a thought produce weaker embeddings.
RAG Quality Metrics
How do you check whether RAG works well? We evaluate two stages separately: did retrieval find the right passages, and does the answer stick to the context. Here is a class skeleton with three typical metrics:
1from dataclasses import dataclass
2
3@dataclass
4class RAGMetrics:
5 """Metrics for evaluating a RAG system."""
6
7 def relevance_score(self, query: str, retrieved_doc: str) -> float:
8 """Is the document relevant to the question?"""
9 # In practice we use an LLM for evaluation
10 pass
11
12 def faithfulness_score(self, answer: str, context: str) -> float:
13 """Is the answer grounded in the context?"""
14 # Checks for hallucinations
15 pass
16
17 def answer_correctness(self, answer: str, ground_truth: str) -> float:
18 """Is the answer correct?"""
19 # Comparison with ground truth
20 passThe methods contain only pass, because in practice relevance, faithfulness and answer correctness are usually judged by a second model acting as a referee. It is a skeleton to fill in, not a finished tool.
Retrieval, on the other hand, can be measured with plain set arithmetic. Precision tells what share of the retrieved documents is relevant, recall tells what share of the relevant documents we found:
1# Simple evaluation example
2def evaluate_retrieval(query: str, retrieved: list[str], relevant: list[str]) -> dict:
3 """Computes precision and recall for retrieval."""
4 retrieved_set = set(retrieved)
5 relevant_set = set(relevant)
6
7 true_positives = len(retrieved_set & relevant_set)
8
9 precision = true_positives / len(retrieved_set) if retrieved_set else 0
10 recall = true_positives / len(relevant_set) if relevant_set else 0
11
12 f1 = 2 * precision * recall / (precision + recall) if (precision + recall) > 0 else 0
13
14 return {
15 "precision": precision,
16 "recall": recall,
17 "f1": f1
18 }Example: if the system returned ["a", "b", "c"] and the relevant ones were ["a", "d"], precision is 1/3, recall is 1/2, and F1, the harmonic mean of both, is 0.4.
RAG is the foundation of modern enterprise AI applications. In the next lesson you will learn about embeddings - the heart of the entire search system - and see why similar meanings end up close to each other.
Remember from this expedition: RAG is a guide who checks the field notebook before every answer.
Spotted a mistake in this lesson?
Check yourself
Answer the questions from this lesson. Pick an answer to see right away whether it is correct.
1. What does the acronym RAG stand for in the context of AI?
2. Which limitation of traditional LLMs does RAG help solve?
Hands-on tasks in the game
- Code editor
Implement a basic RAG pipeline
- Vertical ordering
Arrange the steps in order:
- Click in order
Arrange the RAG pipeline steps:
- Vertical ordering
Arrange the RAG query pipeline steps in the correct order: