Python course Β· Module 12: Final Project
AI Testing Strategy
In this lesson5
A ranger who only checks that the Land Rover starts in the car park still does not know whether it will cross the river. Testing AI applications has the same problem in a double dose. Ordinary code behaves deterministically: the same function with the same input gives the same result. A language model can answer the same question in two different ways. That is why we have to verify not only the code, but also the quality of the answers.
The Test Pyramid for AI
The classic test pyramid says: most of your tests should be fast unit tests at the base, fewer integration tests in the middle, and the fewest slow end-to-end tests at the top. In an AI project we add a layer that evaluates the model's answers. The proportions below are a rough target, not a rule:
1 β±β²
2 β± β²
3 β± E2Eβ² β 10% - End-to-end tests
4 β±βββββββ²
5 β± β²
6 β±Integrationβ² β 20% - Integration tests
7 β±βββββββββββββ²
8 β± β²
9 β± AI Evaluation β² β 30% - AI evaluation
10 β±βββββββββββββββββββ²
11 β± β²
12 β± Unit β² β 40% - Unit tests
13 β±βββββββββββββββββββββββββ²The base is the widest because unit tests are the fastest and cheapest. AI evaluation sits above them: it is slower and often costs money (every question is a model call). From cheapest to most expensive: unit, integration, E2E, and finally performance tests under load.
Unit Tests
We write unit tests in pytest. A fixture is a function marked with @pytest.fixture that prepares data or resources for a test, and pytest injects it by parameter name. Mock from unittest.mock creates a fake object, and spec= makes sure the fake only has the methods of the real class.
1import pytest
2from unittest.mock import Mock, AsyncMock
3
4class TestDocumentService:
5 @pytest.fixture
6 def mock_repo(self):
7 return Mock(spec=DocumentRepository)
8
9 @pytest.fixture
10 def mock_embedder(self):
11 embedder = Mock()
12 embedder.embed.return_value = [0.1] * 1536
13 return embedder
14
15 @pytest.fixture
16 def service(self, mock_repo, mock_embedder):
17 return DocumentService(mock_repo, mock_embedder)The service fixture itself uses two other fixtures, so pytest builds the whole dependency tree for you. The embedder returns a fixed vector of 1536 numbers without any API call. The rest of the same class holds the actual tests:
1 def test_chunk_document(self, service):
2 doc = Document(content="A" * 1000)
3 chunks = service.chunk_document(doc, chunk_size=200)
4
5 assert len(chunks) == 5
6 assert all(len(c.content) <= 200 for c in chunks)
7
8 @pytest.mark.asyncio
9 async def test_save_document(self, service, mock_repo):
10 doc = Document(title="Test", content="Content")
11 mock_repo.save = AsyncMock(return_value=doc)
12
13 result = await service.save(doc)
14
15 mock_repo.save.assert_called_once()
16 assert result.title == "Test"The chunking test checks pure logic: 1000 characters in pieces of 200 gives 5 chunks. The save test is asynchronous, so it carries the @pytest.mark.asyncio marker from the pytest-asyncio plugin and uses AsyncMock, which returns a value after await. An assertion is simply assert result == expected or, as here, assert result.title == "Test". To replace a function in a module (for example an external API call) you use unittest.mock.patch, either as a decorator or as with patch("module.function") as fake:.
Integration Tests
An integration test runs several parts together. Testcontainers starts a real Qdrant in a Docker container for the duration of the tests, and httpx.AsyncClient sends requests straight to the ASGI application, without a network.
1import pytest
2import pytest_asyncio
3from httpx import ASGITransport, AsyncClient
4from testcontainers.community.qdrant import QdrantContainer
5
6@pytest.fixture(scope="module")
7def qdrant_container():
8 with QdrantContainer() as qdrant:
9 yield qdrant
10
11@pytest_asyncio.fixture
12async def app_client(qdrant_container):
13 app = create_app(qdrant_url=f"http://{qdrant_container.rest_host_address}")
14 transport = ASGITransport(app=app)
15 async with AsyncClient(transport=transport, base_url="http://test") as client:
16 yield clientTwo things have changed in the libraries and are worth knowing. Since httpx 0.28 there is no app= parameter anymore, you pass the application through ASGITransport. In pytest-asyncio, an asynchronous fixture is marked with @pytest_asyncio.fixture. You install the Qdrant container as testcontainers[qdrant], and older versions imported it from testcontainers.qdrant. Now the test itself:
1@pytest.mark.asyncio
2async def test_upload_and_query(app_client):
3 # Upload document
4 response = await app_client.post(
5 "/documents",
6 json={"title": "Test", "content": "Python is great"}
7 )
8 assert response.status_code == 201
9
10 # Query
11 response = await app_client.post(
12 "/query",
13 json={"question": "What is Python?"}
14 )
15 assert response.status_code == 200
16 assert "Python" in response.json()["answer"]The test walks through the whole flow: upload, then query. Notice that the last assertion depends on the model's answer, so it can be flaky. Checks like that are better moved to evaluation.
AI Evaluation
Evaluation measures answer quality on a prepared set of questions. Each case contains a question and the keywords we expect in the answer:
1from dataclasses import dataclass
2
3@dataclass
4class EvalCase:
5 question: str
6 expected_keywords: list[str]
7 context_required: bool = True
8
9eval_dataset = [
10 EvalCase(
11 question="What is RAG?",
12 expected_keywords=["retrieval", "generation", "augmented"],
13 context_required=True
14 ),
15 EvalCase(
16 question="How does embedding work?",
17 expected_keywords=["vector", "semantic", "representation"],
18 context_required=True
19 )
20]The dataset itself is plain data. The scoring function computes what share of the keywords appeared in the answer and checks whether the system returned sources:
1async def evaluate_rag_system(query_fn, dataset: list[EvalCase]) -> dict:
2 results = {"total": len(dataset), "passed": 0, "failed": []}
3
4 for case in dataset:
5 response = await query_fn(case.question)
6
7 # Check keywords
8 answer_lower = response.answer.lower()
9 keywords_found = sum(1 for k in case.expected_keywords if k in answer_lower)
10 keyword_score = keywords_found / len(case.expected_keywords)
11
12 # Check sources
13 has_sources = len(response.sources) > 0 if case.context_required else True
14
15 if keyword_score >= 0.5 and has_sources:
16 results["passed"] += 1
17 else:
18 results["failed"].append({
19 "question": case.question,
20 "keyword_score": keyword_score,
21 "has_sources": has_sources
22 })
23
24 results["accuracy"] = results["passed"] / results["total"]
25 return resultsThe result is an accuracy value plus a list of failures to review. Keywords are a simple, rough method. More serious projects also score how faithful the answer is to its sources, but the principle stays: a fixed set of questions, a repeatable measurement.
CI/CD Pipeline
Tests make sense when they run on their own on every push. GitHub Actions runs the steps on a clean machine, with Qdrant running alongside as a service:
1# .github/workflows/test.yml
2name: Test Suite
3
4on: [push, pull_request]
5
6jobs:
7 test:
8 runs-on: ubuntu-latest
9 services:
10 qdrant:
11 image: qdrant/qdrant
12 ports:
13 - 6333:6333
14
15 steps:
16 - uses: actions/checkout@v4
17 - uses: actions/setup-python@v5
18 with:
19 python-version: "3.12"
20
21 - name: Install dependencies
22 run: pip install -r requirements.txt
23
24 - name: Run unit tests
25 run: pytest tests/unit -v
26
27 - name: Run integration tests
28 run: pytest tests/integration -v
29 env:
30 QDRANT_URL: http://localhost:6333
31
32 - name: Run AI evaluation
33 run: python scripts/evaluate.py
34 env:
35 OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}The API key comes from secrets, never from the repository. Code coverage is measured by the pytest-cov plugin: pytest --cov=app lists which lines were never tested. You run a single file with pytest tests/unit/test_service.py, and many pipelines also build an image after the tests with docker build -t myapp:latest ., which we will come back to in the deployment lesson. My advice: run the AI evaluation separately from the unit tests, because it is slower and costs money. In the next lesson we will work on documentation.
Remember: tests are trial crossings of the river before you lead the whole expedition through it.
Spotted a mistake in this lesson?
Check yourself
Answer the questions from this lesson. Pick an answer to see right away whether it is correct.
1. What is at the base of the test pyramid?
2. What are pytest fixtures used for?
These are 2 of 3 questions for this lesson. Solve the rest in the game.
Hands-on tasks in the game
- Vertical ordering
Arrange the types of tests from fastest to slowest:
- Code editor
Write a unit test for a data processing function using pytest
- Click in order
Click the elements to build a pytest assert statement:
- Horizontal ordering
Arrange the Docker build command in the correct order:
- Code editor
Create a test using Mock to simulate an external API dependency
- Vertical ordering
Arrange the AI test pyramid layers from the base to the top: