Python course Β· Module 11: RAG and Multi-Agent Systems

LlamaIndex - RAG Framework

7 min read
In this lesson8

In the three previous lessons we wrote chunking, embeddings, search and prompt assembly ourselves. That is good tracking school, but on a real expedition nobody sews their own tent before leaving. You take proven gear and focus on the route.

LlamaIndex is a framework for building RAG applications. It simplifies the whole process - from loading documents to generating answers - and provides ready-made building blocks for every step we have written by hand so far.

LlamaIndex Basics

The whole pipeline from the first lesson fits into a few lines. The Settings object globally sets the language model and the embedding model, and SimpleDirectoryReader reads all files from a directory:

1from llama_index.core import VectorStoreIndex, SimpleDirectoryReader, Settings
2from llama_index.llms.openai import OpenAI
3from llama_index.embeddings.openai import OpenAIEmbedding
4
5# Global configuration
6Settings.llm = OpenAI(model="gpt-4o-mini", temperature=0)
7Settings.embed_model = OpenAIEmbedding(model="text-embedding-3-small")
8
9# Load documents
10documents = SimpleDirectoryReader("./docs").load_data()
11print(f"Loaded {len(documents)} documents")
12
13# Create index
14index = VectorStoreIndex.from_documents(documents)
15
16# Query engine
17query_engine = index.as_query_engine()
18
19# Query
20response = query_engine.query("What is RAG?")
21print(response)

The chain is always the same: SimpleDirectoryReader and .load_data() load the documents, VectorStoreIndex.from_documents splits them into chunks, computes embeddings and builds the index, and .as_query_engine() turns the index into a question engine. Under the hood exactly what we wrote ourselves happens: retrieve, augment, generate. Only the level you work at has changed.

Document Loaders

Documents rarely come in a single format. LlamaIndex has loaders for text, files and web pages:

1from llama_index.core import Document
2from llama_index.readers.web import SimpleWebPageReader
3from llama_index.readers.file import PDFReader, DocxReader
4
5# From text
6doc = Document(text="This is sample text to be indexed.")
7
8# From files
9pdf_reader = PDFReader()
10pdf_docs = pdf_reader.load_data(file="document.pdf")
11
12# From web page
13web_reader = SimpleWebPageReader()
14web_docs = web_reader.load_data(urls=["https://example.com"])
15
16# From multiple sources
17from llama_index.core import SimpleDirectoryReader
18
19reader = SimpleDirectoryReader(
20    input_dir="./data",
21    recursive=True,
22    required_exts=[".pdf", ".docx", ".txt", ".md"]
23)
24documents = reader.load_data()

Document is the basic data unit: text plus metadata. PDFReader and SimpleWebPageReader come from separate packages (llama-index-readers-file, llama-index-readers-web), so you have to install them. SimpleDirectoryReader with recursive=True also walks subdirectories and takes only the extensions listed in required_exts.

Node Parsing (Chunking)

In LlamaIndex a document chunk is called a node. You can order the splitting strategies from the simplest to the most advanced: cutting by characters, like our simple_chunker from the first lesson, sentence splitting, semantic splitting and the hierarchical parser. We start with sentences:

1from llama_index.core.node_parser import (
2    SentenceSplitter,
3    SemanticSplitterNodeParser,
4)
5
6# Simple sentence splitting
7sentence_parser = SentenceSplitter(
8    chunk_size=512,
9    chunk_overlap=50
10)
11nodes = sentence_parser.get_nodes_from_documents(documents)

SentenceSplitter cuts the text into chunks of up to 512 tokens while trying not to break sentences, and chunk_overlap is the overlap you already know.

Semantic splitting goes further: it compares the embeddings of neighbouring sentences and cuts where the meaning clearly changes:

1# Semantic splitting
2semantic_parser = SemanticSplitterNodeParser(
3    buffer_size=1,
4    breakpoint_percentile_threshold=95,
5    embed_model=Settings.embed_model
6)
7semantic_nodes = semantic_parser.get_nodes_from_documents(documents)

breakpoint_percentile_threshold=95 means a cut happens only at the 5% largest jumps in meaning. It is more expensive, because it needs an embedding of every sentence.

The hierarchical parser creates several levels of chunks at once, from large to small:

1# Hierarchical splitting
2from llama_index.core.node_parser import HierarchicalNodeParser
3
4hierarchical_parser = HierarchicalNodeParser.from_defaults(
5    chunk_sizes=[2048, 512, 128]
6)
7hierarchical_nodes = hierarchical_parser.get_nodes_from_documents(documents)

Small chunks (128 tokens) match well in search, while large ones (2048) give the model a wider context. The input documents do not change, only the way they are cut does.

Retriever Modes

A retriever is the scout that brings back matching chunks. You can configure it manually and add post-processing:

1from llama_index.core.retrievers import VectorIndexRetriever
2from llama_index.core.query_engine import RetrieverQueryEngine
3from llama_index.core.postprocessor import SimilarityPostprocessor
4
5# Basic retriever
6retriever = VectorIndexRetriever(
7    index=index,
8    similarity_top_k=5
9)
10
11# With postprocessing
12query_engine = RetrieverQueryEngine(
13    retriever=retriever,
14    node_postprocessors=[
15        SimilarityPostprocessor(similarity_cutoff=0.7)
16    ]
17)

similarity_top_k=5 fetches five chunks, and SimilarityPostprocessor drops the ones with similarity below 0.7, so weak hits do not clutter the prompt.

The hybrid search from the previous lesson is ready here too. Note: in current versions BM25Retriever lives in the separate llama-index-retrievers-bm25 package, not in llama_index.core:

1# Hybrid retriever
2from llama_index.retrievers.bm25 import BM25Retriever  # pip install llama-index-retrievers-bm25
3from llama_index.core.retrievers import QueryFusionRetriever
4
5vector_retriever = index.as_retriever(similarity_top_k=5)
6bm25_retriever = BM25Retriever.from_defaults(nodes=nodes, similarity_top_k=5)
7
8hybrid_retriever = QueryFusionRetriever(
9    retrievers=[vector_retriever, bm25_retriever],
10    similarity_top_k=5,
11    num_queries=1,
12)

QueryFusionRetriever merges the results of both scouts into one list. num_queries=1 means we use only the original question, without extra variants generated by the LLM.

Query Transformations

Sometimes the problem lies in the question itself. A complex question is worth breaking into smaller ones, and a short question can be rewritten to match the documents better:

1from llama_index.core.query_engine import SubQuestionQueryEngine
2from llama_index.core.tools import QueryEngineTool
3
4# Sub-question engine - breaks questions into smaller ones
5tools = [
6    QueryEngineTool.from_defaults(
7        query_engine=index.as_query_engine(),
8        name="documentation",
9        description="Contains project documentation"
10    )
11]
12
13sub_question_engine = SubQuestionQueryEngine.from_defaults(
14    query_engine_tools=tools
15)
16
17# HyDE - Hypothetical Document Embeddings
18from llama_index.core.indices.query.query_transform import HyDEQueryTransform
19from llama_index.core.query_engine import TransformQueryEngine
20
21hyde = HyDEQueryTransform(include_original=True)
22hyde_query_engine = TransformQueryEngine(
23    index.as_query_engine(),
24    query_transform=hyde
25)

SubQuestionQueryEngine breaks a question into sub-questions and sends each one to the right tool, here a single one named "documentation". HyDE first asks the model to write a hypothetical answer and searches for documents similar to it, because an answer is often closer to the documents than the question itself. include_original=True keeps the original question as well.

Chat Engine

A query engine does not remember previous questions. For a conversation we use a chat engine with memory:

1from llama_index.core.chat_engine import CondenseQuestionChatEngine
2from llama_index.core.memory import ChatMemoryBuffer
3
4# Chat memory
5memory = ChatMemoryBuffer.from_defaults(token_limit=3000)
6
7# Chat engine with history
8chat_engine = index.as_chat_engine(
9    chat_mode="condense_question",
10    memory=memory,
11    verbose=True
12)
13
14# Conversation
15response1 = chat_engine.chat("What is Python?")
16print(response1)
17
18response2 = chat_engine.chat("What are its applications?")
19print(response2)
20
21# Reset memory
22chat_engine.reset()

The condense_question mode rewrites a follow-up question such as "What are its applications?" into a standalone question using the history, and only then searches the index. reset() clears the memory. In the newest versions ChatMemoryBuffer is marked as deprecated in favour of the Memory class from the same module, but it still works.

Vector Database Integration

By default the index lives in memory. To make it survive a restart, we connect a vector database from the previous lesson. This is what Chroma looks like:

1# Chroma
2from llama_index.vector_stores.chroma import ChromaVectorStore
3import chromadb
4
5chroma_client = chromadb.PersistentClient(path="./chroma_db")
6chroma_collection = chroma_client.get_or_create_collection("llama_index")
7
8vector_store = ChromaVectorStore(chroma_collection=chroma_collection)
9index = VectorStoreIndex.from_vector_store(vector_store)

from_vector_store opens an index over a collection that already contains data. To write new documents into it, pass vector_store through a StorageContext to from_documents.

Qdrant and Pinecone connect in exactly the same way, only the store class changes:

1# Qdrant
2from llama_index.vector_stores.qdrant import QdrantVectorStore
3from qdrant_client import QdrantClient
4
5qdrant_client = QdrantClient(host="localhost", port=6333)
6vector_store = QdrantVectorStore(
7    client=qdrant_client,
8    collection_name="llama_index"
9)
10
11# Pinecone
12from llama_index.vector_stores.pinecone import PineconeVectorStore
13from pinecone import Pinecone
14
15pc = Pinecone(api_key="...")
16pinecone_index = pc.Index("llama-index")
17vector_store = PineconeVectorStore(pinecone_index=pinecone_index)

The rest of the code, meaning the query engine, retrievers and chat engine, stays unchanged, and that is the biggest advantage of this abstraction layer.

Evaluation

Finally we check whether the answers are trustworthy. LlamaIndex has evaluators that use a language model as a judge:

1from llama_index.core.evaluation import (
2    FaithfulnessEvaluator,
3    RelevancyEvaluator,
4    CorrectnessEvaluator
5)
6
7# Evaluators
8faithfulness = FaithfulnessEvaluator()
9relevancy = RelevancyEvaluator()
10
11# Evaluate response
12query = "What is RAG?"
13response = query_engine.query(query)
14
15faithfulness_result = faithfulness.evaluate_response(response=response)
16print(f"Faithfulness: {faithfulness_result.passing}")
17
18relevancy_result = relevancy.evaluate_response(query=query, response=response)
19print(f"Relevancy: {relevancy_result.passing}")

FaithfulnessEvaluator checks whether the answer follows from the retrieved context, so it catches hallucinations. RelevancyEvaluator checks whether the answer and the context match the question. The passing field is either True or False.

I recommend starting with SentenceSplitter and a simple query engine, and reaching for HyDE or hybrid search only when evaluation shows a concrete problem.

LlamaIndex is a complete framework for RAG. In the next lesson you will learn about multi-agent systems - when one agent is not enough!

Remember: LlamaIndex is ready-made expedition gear that lets you focus on the route instead of sewing the tent.

Spotted a mistake in this lesson?

Check yourself

Answer the questions from this lesson. Pick an answer to see right away whether it is correct.

  1. 1. What is LlamaIndex used for?

  2. 2. Which LlamaIndex class is used to load documents from a directory?

Hands-on tasks in the game

  • Code editor

    Implement a basic Q&A system with LlamaIndex

  • Horizontal ordering

    Arrange the basic LlamaIndex pipeline:

  • Vertical ordering

    Sort chunking strategies from simplest to most advanced:

Useful articles