Python course Β· Module 11: RAG and Multi-Agent Systems
Embeddings and Vector Search
In this lesson5
A user asks "where do cats rest?", and your notebook contains the sentence "The kitten lies on the carpet". A keyword search finds no shared word here, even though the meaning is almost the same. We need a way to compare meaning, not letters. That way is embeddings.
Embeddings are vector representations of text that let machines "understand" the meaning of words and sentences. Think of them as a footprint: every animal leaves a different track, and a tracker can tell from the shape who walked down the trail. An embedding is such a footprint of a text - a sequence of numbers that describes its semantics.
What are embeddings?
An embedding is a vector of real numbers representing text in a semantic space. Texts with similar meanings get vectors that lie close to each other. The simplest way to generate one is the OpenAI API, with the client.embeddings.create method:
1from openai import OpenAI
2
3client = OpenAI()
4
5def get_embedding(text: str, model: str = "text-embedding-3-small") -> list[float]:
6 """Generates an embedding for the given text."""
7 response = client.embeddings.create(
8 model=model,
9 input=text
10 )
11 return response.data[0].embedding
12
13# Example
14text = "Python is a great programming language"
15embedding = get_embedding(text)
16
17print(f"Vector dimension: {len(embedding)}") # 1536 for text-embedding-3-small
18print(f"First 5 values: {embedding[:5]}")The function returns an ordinary list of floats. The text-embedding-3-small model produces 1536-dimensional vectors by default, and text-embedding-3-large produces 3072. The text itself is not stored anywhere, we only get its numeric footprint. The process always has the same rhythm: prepare the text, send it to the embedding model, receive the vector and store it in a vector database.
Embedding Models
Not every camp has a budget for a paid API. The sentence-transformers library runs open-source models locally, on your own machine. The class below keeps a local model and an OpenAI model in one place:
1from sentence_transformers import SentenceTransformer
2
3# Open-source models
4class EmbeddingModels:
5 """Various embedding models."""
6
7 def __init__(self):
8 # Fast and lightweight
9 self.minilm = SentenceTransformer('all-MiniLM-L6-v2')
10
11 # Larger, better for multilingual
12 self.multilingual = SentenceTransformer('paraphrase-multilingual-mpnet-base-v2')
13
14 def embed_with_openai(self, texts: list[str]) -> list[list[float]]:
15 """OpenAI embeddings - highest quality."""
16 from openai import OpenAI
17 client = OpenAI()
18
19 response = client.embeddings.create(
20 model="text-embedding-3-large", # 3072 dimensions
21 input=texts
22 )
23 return [item.embedding for item in response.data]
24
25 def embed_with_sentence_transformers(self, texts: list[str]) -> list[list[float]]:
26 """Local embeddings - no API costs."""
27 return self.minilm.encode(texts).tolist()all-MiniLM-L6-v2 is fast and light (384 dimensions), paraphrase-multilingual-mpnet-base-v2 is bigger and handles many languages, including Polish. The encode method returns a NumPy array, which is why we call .tolist().
Now let's embed two related sentences about AI and one from a completely different story:
1# Model comparison
2models = EmbeddingModels()
3
4texts = [
5 "Machine learning is a field of AI",
6 "Deep learning uses neural networks",
7 "I like eating pizza"
8]
9
10# Open-source embeddings
11embeddings = models.embed_with_sentence_transformers(texts)
12print(f"Local embeddings: {len(embeddings[0])} dimensions")The output shows 384 dimensions. Notice that the dimension depends on the model, not on the text length: a short sentence and a long paragraph get vectors of the same length.
Similarity Metrics
We have vectors, so we need to compare them. Here are the three most popular measures, each a single line of NumPy:
1import numpy as np
2from typing import Callable
3
4def cosine_similarity(vec1: np.ndarray, vec2: np.ndarray) -> float:
5 """Cosine similarity - most commonly used."""
6 return np.dot(vec1, vec2) / (np.linalg.norm(vec1) * np.linalg.norm(vec2))
7
8def euclidean_distance(vec1: np.ndarray, vec2: np.ndarray) -> float:
9 """Euclidean distance."""
10 return np.linalg.norm(vec1 - vec2)
11
12def dot_product(vec1: np.ndarray, vec2: np.ndarray) -> float:
13 """Dot product."""
14 return np.dot(vec1, vec2)Cosine similarity looks only at the angle between vectors, and it is the metric most often used to compare embeddings. Euclidean distance measures distance, so here a smaller number means more similar. The dot product depends on both the angle and the length of the vectors.
Let's check it on three sentences, two of which are about cats:
1# Demonstration
2def demonstrate_similarity():
3 """Shows how similarity metrics work."""
4 from sentence_transformers import SentenceTransformer
5
6 model = SentenceTransformer('all-MiniLM-L6-v2')
7
8 sentences = [
9 "The cat sits on the mat", # 0
10 "The kitten lies on the carpet", # 1 - similar meaning
11 "Programming in Python", # 2 - different meaning
12 ]
13
14 embeddings = model.encode(sentences)
15
16 print("Cosine similarity:")
17 print(f" Sentence 0 vs 1: {cosine_similarity(embeddings[0], embeddings[1]):.3f}")
18 print(f" Sentence 0 vs 2: {cosine_similarity(embeddings[0], embeddings[2]):.3f}")
19 print(f" Sentence 1 vs 2: {cosine_similarity(embeddings[1], embeddings[2]):.3f}")
20
21demonstrate_similarity()Sentences 0 and 1 should get a clearly higher score than the pairs with the Python sentence, even though "cat" and "kitten" are different words. I am not writing exact numbers, because they depend on the model version.
One more useful option. If you ask the model for normalized vectors (length equal to 1), the dot product becomes exactly the cosine similarity:
1from sentence_transformers import SentenceTransformer
2
3model = SentenceTransformer('all-MiniLM-L6-v2')
4document_text = "Lions hunt most often at dusk"
5
6# Unit-length vector - the dot product directly gives cosine similarity
7embedding = model.encode(document_text, normalize_embeddings=True)
8print(embedding.shape) # (384,)The parameter is called normalize_embeddings, and a single string gives a one-dimensional array. The vector points in the same direction as without normalization, only its length changes.
Semantic Search
Let's put these pieces together into a semantic search engine. First a small container for a result, made with the @dataclass decorator, and a class that indexes the documents once:
1from dataclasses import dataclass
2import numpy as np
3
4@dataclass
5class SearchResult:
6 """Semantic search result."""
7 text: str
8 score: float
9 index: int
10
11class SemanticSearch:
12 """Simple semantic search implementation."""
13
14 def __init__(self, model_name: str = 'all-MiniLM-L6-v2'):
15 from sentence_transformers import SentenceTransformer
16 self.model = SentenceTransformer(model_name)
17 self.documents: list[str] = []
18 self.embeddings: np.ndarray | None = None
19
20 def index_documents(self, documents: list[str]) -> None:
21 """Indexes documents."""
22 self.documents = documents
23 self.embeddings = self.model.encode(documents)
24 print(f"Indexed {len(documents)} documents")index_documents computes the embeddings of all documents with a single encode call and keeps them in a matrix. This fixes the weakness from the previous lesson: we embed documents once, not on every question.
The search method computes the similarity of the question to the whole matrix at once, without a loop:
1 def search(self, query: str, top_k: int = 5) -> list[SearchResult]:
2 """Searches for the most similar documents."""
3 if self.embeddings is None:
4 raise ValueError("No indexed documents!")
5
6 query_embedding = self.model.encode([query])[0]
7
8 # Compute similarities
9 similarities = np.dot(self.embeddings, query_embedding)
10 similarities /= np.linalg.norm(self.embeddings, axis=1)
11 similarities /= np.linalg.norm(query_embedding)
12
13 # Top-K results
14 top_indices = np.argsort(similarities)[::-1][:top_k]
15
16 results = []
17 for idx in top_indices:
18 results.append(SearchResult(
19 text=self.documents[idx],
20 score=float(similarities[idx]),
21 index=int(idx)
22 ))
23
24 return resultsnp.argsort sorts in ascending order, so [::-1] reverses it and [:top_k] takes the best results. The two divisions by norms turn the dot product into cosine similarity.
Time for a test on five documents:
1# Usage example
2search = SemanticSearch()
3
4documents = [
5 "Python is a general-purpose programming language",
6 "Machine Learning uses algorithms to learn from data",
7 "RAG combines retrieval with text generation",
8 "FastAPI is a modern framework for building APIs",
9 "Docker containerizes applications",
10]
11
12search.index_documents(documents)
13
14results = search.search("How to build web applications?")
15for r in results:
16 print(f"Score: {r.score:.3f} | {r.text}")For the question about web applications, the FastAPI document should rank highest, even though it shares no keyword with the question.
Batch Processing
A thousand documents sent one by one means a thousand HTTP requests. The OpenAI API accepts a list of texts in the input parameter, so we send them in batches:
1def batch_embed(texts: list[str], batch_size: int = 100) -> list[list[float]]:
2 """Embeds large amounts of text in batches."""
3 from openai import OpenAI
4 client = OpenAI()
5
6 all_embeddings = []
7
8 for i in range(0, len(texts), batch_size):
9 batch = texts[i:i + batch_size]
10
11 response = client.embeddings.create(
12 model="text-embedding-3-small",
13 input=batch
14 )
15
16 batch_embeddings = [item.embedding for item in response.data]
17 all_embeddings.extend(batch_embeddings)
18
19 print(f"Processed {min(i + batch_size, len(texts))}/{len(texts)}")
20
21 return all_embeddingsThe order of results in response.data matches the order of texts in the batch, which is why we can simply append them with extend.
The second way to save money is a cache. The same text with the same model always produces the same embedding, so we store the result on disk under a hash key:
1# Caching embeddings
2import hashlib
3import json
4from pathlib import Path
5
6class EmbeddingCache:
7 """Cache for embeddings."""
8
9 def __init__(self, cache_dir: str = ".embedding_cache"):
10 self.cache_dir = Path(cache_dir)
11 self.cache_dir.mkdir(exist_ok=True)
12
13 def _get_cache_key(self, text: str, model: str) -> str:
14 """Generates a cache key."""
15 content = f"{model}:{text}"
16 return hashlib.md5(content.encode()).hexdigest()
17
18 def get(self, text: str, model: str) -> list[float] | None:
19 """Retrieves an embedding from the cache."""
20 key = self._get_cache_key(text, model)
21 cache_file = self.cache_dir / f"{key}.json"
22
23 if cache_file.exists():
24 return json.loads(cache_file.read_text())
25 return None
26
27 def set(self, text: str, model: str, embedding: list[float]) -> None:
28 """Saves an embedding to the cache."""
29 key = self._get_cache_key(text, model)
30 cache_file = self.cache_dir / f"{key}.json"
31 cache_file.write_text(json.dumps(embedding))MD5 serves here only as a fast file identifier, not as security, so its cryptographic weaknesses do not matter. The model is part of the key on purpose: after a model change, old vectors do not match new ones. I recommend caching embeddings from day one, because re-indexing then saves both time and money.
Embeddings are the foundation of every RAG system. In the next lesson you will learn about vector databases - specialized databases for storing and searching vectors - which do what our NumPy matrix does, but for millions of documents.
Remember: an embedding is the footprint of a text, and similar meanings leave similar tracks.
Spotted a mistake in this lesson?
Check yourself
Answer the questions from this lesson. Pick an answer to see right away whether it is correct.
1. What is an embedding in the context of NLP?
2. Which metric is most commonly used to compare embeddings?
Hands-on tasks in the game
- Code editor
Write a cosine_similarity function
- Vertical ordering
Arrange the steps in order:
- Horizontal ordering
Arrange the embedding creation method call:
- Click in order
Arrange the embedding call: