Python course Β· Module 11: RAG and Multi-Agent Systems

Vector Databases

9 min read
In this lesson6

In the previous lesson we kept embeddings in a NumPy matrix. With five documents that is enough, but with a million the matrix does not fit in memory, disappears when the program restarts, and scanning it every time takes too long. We need an archive of tracks that survives the night at camp and answers in milliseconds.

Vector databases are specialized databases optimized for storing and searching embeddings. They are like a library with a magical catalog that finds similar books based on their content, not their title.

The diagram groups the popular databases by where they run:

1β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
2β”‚              Vector Databases Landscape                  β”‚
3β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
4β”‚                                                          β”‚
5β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
6β”‚  β”‚   Pinecone   β”‚  β”‚    Qdrant    β”‚  β”‚   Chroma     β”‚  β”‚
7β”‚  β”‚   (Cloud)    β”‚  β”‚  (Self-host) β”‚  β”‚   (Local)    β”‚  β”‚
8β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
9β”‚                                                          β”‚
10β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
11β”‚  β”‚   Weaviate   β”‚  β”‚    Milvus    β”‚  β”‚    FAISS     β”‚  β”‚
12β”‚  β”‚  (Semantic)  β”‚  β”‚  (Enterprise)β”‚  β”‚  (In-memory) β”‚  β”‚
13β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
14β”‚                                                          β”‚
15β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Pinecone runs only in the cloud as a managed service, Qdrant and Milvus can be hosted on your own server, Chroma works great locally for prototypes, and FAISS is a library from Meta that keeps the index in the memory of your process.

Chroma - Local Vector Database

Chroma is a local vector database for RAG applications that you install with a single pip install chromadb. First we create a client, an embedding function and a collection, the equivalent of a table:

1import chromadb
2from chromadb.utils import embedding_functions
3
4# Initialize client
5client = chromadb.Client()  # In-memory
6# or: client = chromadb.PersistentClient(path="./chroma_db")  # Persistent
7
8# Configure embedding function
9openai_ef = embedding_functions.OpenAIEmbeddingFunction(
10    api_key="your-key",
11    model_name="text-embedding-3-small"
12)
13
14# Create collection
15collection = client.create_collection(
16    name="python_safari",
17    embedding_function=openai_ef,
18    metadata={"description": "Python Safari course documentation"}
19)

chromadb.Client() keeps data only in memory, while PersistentClient saves it to a directory on disk. The collection calls openai_ef for every document by itself, so you do not have to compute embeddings manually. Do not type the API key into the code as in the example: in a real project read it from an environment variable.

Now we add documents with metadata and search right away:

1# Add documents
2collection.add(
3    documents=[
4        "Python is a high-level programming language",
5        "RAG combines retrieval with text generation",
6        "Vector databases store embeddings"
7    ],
8    metadatas=[
9        {"topic": "python", "level": "beginner"},
10        {"topic": "ai", "level": "advanced"},
11        {"topic": "databases", "level": "intermediate"}
12    ],
13    ids=["doc1", "doc2", "doc3"]
14)
15
16# Search
17results = collection.query(
18    query_texts=["How to start learning programming?"],
19    n_results=2,
20    where={"level": "beginner"}  # Filter by metadata
21)
22
23print(results)

Every document needs a unique id. The where parameter filters by metadata before vectors are compared, so only the document with level equal to beginner passes here. The result is a dictionary of lists, one per question.

If you already have your own vectors, for example from the get_embedding function in the previous lesson, pass them in the embeddings parameter and Chroma skips the computation:

1# Your own vectors instead of the collection's embedding function
2collection.add(
3    documents=["Elephants walk to the watering hole at dawn"],
4    embeddings=[get_embedding("Elephants walk to the watering hole at dawn")],
5    ids=["doc4"]
6)

The vectors must have the same dimension as the rest of the collection, which is why we use the same text-embedding-3-small model.

Qdrant is a database you run on a server or in a Docker container. Working with it always has four steps: connect the client, create a collection with a vector configuration, add points and search. The first two steps look like this:

1from qdrant_client import QdrantClient
2from qdrant_client.models import Distance, VectorParams, PointStruct
3import numpy as np
4
5# Connect to Qdrant
6client = QdrantClient(host="localhost", port=6333)
7# or: client = QdrantClient(":memory:")  # In-memory
8
9# Create collection
10client.create_collection(
11    collection_name="documents",
12    vectors_config=VectorParams(
13        size=1536,  # Embedding size
14        distance=Distance.COSINE
15    )
16)

VectorParams tells the database how long the vectors will be and which measure to compare them with. size must match the embedding model, here 1536 for text-embedding-3-small.

A point in Qdrant is a vector plus a payload, an arbitrary dictionary of data:

1# Add points
2def add_documents(texts: list[str], embeddings: list[list[float]]):
3    """Adds documents to Qdrant."""
4    points = [
5        PointStruct(
6            id=i,
7            vector=embedding,
8            payload={"text": text}
9        )
10        for i, (text, embedding) in enumerate(zip(texts, embeddings))
11    ]
12
13    client.upsert(
14        collection_name="documents",
15        points=points
16    )

The upsert method inserts new points or overwrites existing ones with the same id, which is why running it again does not create duplicates.

In the current Python client we search with the query_points method. The older search method was marked as deprecated, and new versions of qdrant-client no longer have it:

1# Search
2def search(query_embedding: list[float], limit: int = 5):
3    """Searches for similar documents."""
4    results = client.query_points(
5        collection_name="documents",
6        query=query_embedding,
7        limit=limit
8    ).points
9
10    return [
11        {
12            "text": hit.payload["text"],
13            "score": hit.score
14        }
15        for hit in results
16    ]

query_points returns an object with a points list, and every hit has a score and a payload.

A filter works like where in Chroma, only with a more formal syntax:

1# Filtering
2from qdrant_client.models import Filter, FieldCondition, MatchValue
3
4filtered_results = client.query_points(
5    collection_name="documents",
6    query=query_embedding,
7    query_filter=Filter(
8        must=[
9            FieldCondition(
10                key="category",
11                match=MatchValue(value="python")
12            )
13        ]
14    ),
15    limit=5
16).points

Watch out for a trap: add_documents stores only text in the payload, so a filter on category finds nothing until you add that field to the payload. query_embedding is the question vector you have already computed.

FAISS is not a server but a library. The index lives in the memory of your process. Let's wrap it in a class that also remembers the document texts:

1import faiss
2import numpy as np
3from dataclasses import dataclass
4
5@dataclass
6class FAISSIndex:
7    """Wrapper for FAISS."""
8
9    dimension: int
10    index: faiss.Index = None
11    documents: list[str] = None
12
13    def __post_init__(self):
14        # Different index types
15        # Flat - exact, slower
16        self.index = faiss.IndexFlatL2(self.dimension)
17
18        # IVF - faster, approximate
19        # quantizer = faiss.IndexFlatL2(self.dimension)
20        # self.index = faiss.IndexIVFFlat(quantizer, self.dimension, 100)
21
22        self.documents = []
23
24    def add(self, embeddings: np.ndarray, documents: list[str]):
25        """Adds vectors to the index."""
26        embeddings = np.array(embeddings).astype('float32')
27        self.index.add(embeddings)
28        self.documents.extend(documents)

IndexFlatL2 compares the question with every vector, so the result is exact. The commented-out IndexIVFFlat is faster but approximate, and it needs training with the train method before you add vectors. FAISS accepts only float32, hence the conversion in add.

Searching returns the distances and the numbers of the nearest neighbours:

1    def search(self, query_embedding: np.ndarray, k: int = 5) -> list[tuple[str, float]]:
2        """Searches for k nearest neighbors."""
3        query = np.array([query_embedding]).astype('float32')
4        distances, indices = self.index.search(query, k)
5
6        results = []
7        for idx, dist in zip(indices[0], distances[0]):
8            if 0 <= idx < len(self.documents):
9                results.append((self.documents[idx], float(dist)))
10
11        return results

A smaller L2 distance means greater similarity, the opposite of cosine. When you ask for more results than there are vectors, FAISS returns the index -1, which is why we check 0 <= idx.

A test on random vectors shows the mechanism itself:

1# Usage example
2faiss_index = FAISSIndex(dimension=384)
3
4# Add documents
5embeddings = np.random.rand(100, 384).astype('float32')
6documents = [f"Document {i}" for i in range(100)]
7faiss_index.add(embeddings, documents)
8
9# Search
10query = np.random.rand(384).astype('float32')
11results = faiss_index.search(query, k=5)

In real code, instead of random numbers you plug in embeddings from a 384-dimensional model.

Pinecone - Managed Vector Database

Pinecone is a cloud service: you do not install a server, you create an index through the API. When creating it you specify the dimension, the metric and the region:

1from pinecone import Pinecone, ServerlessSpec
2
3# Initialize
4pc = Pinecone(api_key="your-key")
5
6# Create index
7pc.create_index(
8    name="python-safari",
9    dimension=1536,
10    metric="cosine",
11    spec=ServerlessSpec(
12        cloud="aws",
13        region="us-east-1"
14    )
15)
16
17# Connect to index
18index = pc.Index("python-safari")

ServerlessSpec means a serverless index, billed by usage.

Upsert takes a list of dictionaries with an id, a vector and metadata:

1# Upsert (insert/update)
2index.upsert(
3    vectors=[
4        {
5            "id": "doc1",
6            "values": [0.1, 0.2, ...],  # 1536 values
7            "metadata": {
8                "text": "Python is a programming language",
9                "category": "programming",
10                "level": 1
11            }
12        }
13    ],
14    namespace="tutorials"
15)

The notation [0.1, 0.2, ...] is shorthand: in Python the three dots are the Ellipsis object, so in real code you put a full list of 1536 numbers. A namespace splits the index into independent compartments.

A query combines a vector with a metadata filter:

1# Search
2results = index.query(
3    vector=[0.1, 0.2, ...],
4    top_k=10,
5    include_metadata=True,
6    namespace="tutorials",
7    filter={
8        "category": {"$eq": "programming"},
9        "level": {"$lte": 3}
10    }
11)
12
13# Statistics
14stats = index.describe_index_stats()
15print(f"Number of vectors: {stats['total_vector_count']}")

The $eq and $lte operators mean "equal" and "less than or equal", and describe_index_stats reports how many vectors are in the index.

Hybrid Search - Combining Vectors with BM25

Vectors are great at catching meaning, but they can miss an exact name, such as a model number. BM25 is a classic keyword search algorithm. Hybrid search combines both scores:

1from rank_bm25 import BM25Okapi
2import numpy as np
3
4class HybridSearch:
5    """Combines semantic search with keyword search."""
6
7    def __init__(self, documents: list[str], embeddings: np.ndarray):
8        self.documents = documents
9        self.embeddings = embeddings
10
11        # BM25 for keyword search
12        tokenized = [doc.lower().split() for doc in documents]
13        self.bm25 = BM25Okapi(tokenized)

The constructor splits every document into words and builds a BM25Okapi index from them, and it receives the embeddings ready-made.

The search method normalizes both scores to the 0-1 range and mixes them with the alpha weight:

1    def search(
2        self,
3        query: str,
4        query_embedding: np.ndarray,
5        alpha: float = 0.5,  # Semantic search weight
6        top_k: int = 5
7    ) -> list[tuple[str, float]]:
8        """Hybrid search with configurable weights."""
9
10        # Semantic search scores
11        semantic_scores = np.dot(self.embeddings, query_embedding)
12        semantic_scores /= np.linalg.norm(self.embeddings, axis=1)
13        semantic_scores /= np.linalg.norm(query_embedding)
14
15        # Normalize to [0, 1]
16        semantic_scores = (semantic_scores - semantic_scores.min()) / (semantic_scores.max() - semantic_scores.min())
17
18        # BM25 scores
19        bm25_scores = np.array(self.bm25.get_scores(query.lower().split()))
20        if bm25_scores.max() > 0:
21            bm25_scores = bm25_scores / bm25_scores.max()
22
23        # Hybrid score
24        hybrid_scores = alpha * semantic_scores + (1 - alpha) * bm25_scores
25
26        # Top K
27        top_indices = np.argsort(hybrid_scores)[::-1][:top_k]
28
29        return [(self.documents[i], hybrid_scores[i]) for i in top_indices]

alpha=1 is pure semantic search, alpha=0 is pure BM25. I recommend starting at 0.5 and tuning the weight on your own test questions.

Vector databases are the infrastructure of every RAG system. For a prototype take Chroma, for production Qdrant or Pinecone. In the next lesson you will learn LlamaIndex - a framework that simplifies building RAG applications and connects to these databases in a single line.

Remember: a vector database is the camp's archive of tracks, which finds the most similar prints from a single footprint.

Spotted a mistake in this lesson?

Check yourself

Answer the questions from this lesson. Pick an answer to see right away whether it is correct.

  1. 1. Which of the following is a popular vector database?

  2. 2. What is Chroma in the context of vector databases?

Hands-on tasks in the game

  • Horizontal ordering

    Arrange the elements:

  • Code editor

    Implement adding documents to Chroma

  • Horizontal ordering

    Arrange the elements:

  • Vertical ordering

    Arrange the steps for working with Qdrant in the correct order:

  • Click in order

    Arrange adding to vector store:

  • Vertical ordering

    Arrange the steps for creating a FAISS index:

Useful articles