Python course Β· Module 12: Final Project

AI Testing Strategy

6 min read
In this lesson5

A ranger who only checks that the Land Rover starts in the car park still does not know whether it will cross the river. Testing AI applications has the same problem in a double dose. Ordinary code behaves deterministically: the same function with the same input gives the same result. A language model can answer the same question in two different ways. That is why we have to verify not only the code, but also the quality of the answers.

The Test Pyramid for AI

The classic test pyramid says: most of your tests should be fast unit tests at the base, fewer integration tests in the middle, and the fewest slow end-to-end tests at the top. In an AI project we add a layer that evaluates the model's answers. The proportions below are a rough target, not a rule:

1                    β•±β•²
2                   β•±  β•²
3                  β•± E2Eβ•²           ← 10% - End-to-end tests
4                 ╱──────╲
5                β•±        β•²
6               β•±Integrationβ•²       ← 20% - Integration tests
7              ╱────────────╲
8             β•±              β•²
9            β•± AI Evaluation  β•²     ← 30% - AI evaluation
10           ╱──────────────────╲
11          β•±                    β•²
12         β•±        Unit          β•²  ← 40% - Unit tests
13        ╱────────────────────────╲

The base is the widest because unit tests are the fastest and cheapest. AI evaluation sits above them: it is slower and often costs money (every question is a model call). From cheapest to most expensive: unit, integration, E2E, and finally performance tests under load.

Unit Tests

We write unit tests in pytest. A fixture is a function marked with @pytest.fixture that prepares data or resources for a test, and pytest injects it by parameter name. Mock from unittest.mock creates a fake object, and spec= makes sure the fake only has the methods of the real class.

1import pytest
2from unittest.mock import Mock, AsyncMock
3
4class TestDocumentService:
5    @pytest.fixture
6    def mock_repo(self):
7        return Mock(spec=DocumentRepository)
8
9    @pytest.fixture
10    def mock_embedder(self):
11        embedder = Mock()
12        embedder.embed.return_value = [0.1] * 1536
13        return embedder
14
15    @pytest.fixture
16    def service(self, mock_repo, mock_embedder):
17        return DocumentService(mock_repo, mock_embedder)

The service fixture itself uses two other fixtures, so pytest builds the whole dependency tree for you. The embedder returns a fixed vector of 1536 numbers without any API call. The rest of the same class holds the actual tests:

1    def test_chunk_document(self, service):
2        doc = Document(content="A" * 1000)
3        chunks = service.chunk_document(doc, chunk_size=200)
4
5        assert len(chunks) == 5
6        assert all(len(c.content) <= 200 for c in chunks)
7
8    @pytest.mark.asyncio
9    async def test_save_document(self, service, mock_repo):
10        doc = Document(title="Test", content="Content")
11        mock_repo.save = AsyncMock(return_value=doc)
12
13        result = await service.save(doc)
14
15        mock_repo.save.assert_called_once()
16        assert result.title == "Test"

The chunking test checks pure logic: 1000 characters in pieces of 200 gives 5 chunks. The save test is asynchronous, so it carries the @pytest.mark.asyncio marker from the pytest-asyncio plugin and uses AsyncMock, which returns a value after await. An assertion is simply assert result == expected or, as here, assert result.title == "Test". To replace a function in a module (for example an external API call) you use unittest.mock.patch, either as a decorator or as with patch("module.function") as fake:.

Integration Tests

An integration test runs several parts together. Testcontainers starts a real Qdrant in a Docker container for the duration of the tests, and httpx.AsyncClient sends requests straight to the ASGI application, without a network.

1import pytest
2import pytest_asyncio
3from httpx import ASGITransport, AsyncClient
4from testcontainers.community.qdrant import QdrantContainer
5
6@pytest.fixture(scope="module")
7def qdrant_container():
8    with QdrantContainer() as qdrant:
9        yield qdrant
10
11@pytest_asyncio.fixture
12async def app_client(qdrant_container):
13    app = create_app(qdrant_url=f"http://{qdrant_container.rest_host_address}")
14    transport = ASGITransport(app=app)
15    async with AsyncClient(transport=transport, base_url="http://test") as client:
16        yield client

Two things have changed in the libraries and are worth knowing. Since httpx 0.28 there is no app= parameter anymore, you pass the application through ASGITransport. In pytest-asyncio, an asynchronous fixture is marked with @pytest_asyncio.fixture. You install the Qdrant container as testcontainers[qdrant], and older versions imported it from testcontainers.qdrant. Now the test itself:

1@pytest.mark.asyncio
2async def test_upload_and_query(app_client):
3    # Upload document
4    response = await app_client.post(
5        "/documents",
6        json={"title": "Test", "content": "Python is great"}
7    )
8    assert response.status_code == 201
9
10    # Query
11    response = await app_client.post(
12        "/query",
13        json={"question": "What is Python?"}
14    )
15    assert response.status_code == 200
16    assert "Python" in response.json()["answer"]

The test walks through the whole flow: upload, then query. Notice that the last assertion depends on the model's answer, so it can be flaky. Checks like that are better moved to evaluation.

AI Evaluation

Evaluation measures answer quality on a prepared set of questions. Each case contains a question and the keywords we expect in the answer:

1from dataclasses import dataclass
2
3@dataclass
4class EvalCase:
5    question: str
6    expected_keywords: list[str]
7    context_required: bool = True
8
9eval_dataset = [
10    EvalCase(
11        question="What is RAG?",
12        expected_keywords=["retrieval", "generation", "augmented"],
13        context_required=True
14    ),
15    EvalCase(
16        question="How does embedding work?",
17        expected_keywords=["vector", "semantic", "representation"],
18        context_required=True
19    )
20]

The dataset itself is plain data. The scoring function computes what share of the keywords appeared in the answer and checks whether the system returned sources:

1async def evaluate_rag_system(query_fn, dataset: list[EvalCase]) -> dict:
2    results = {"total": len(dataset), "passed": 0, "failed": []}
3
4    for case in dataset:
5        response = await query_fn(case.question)
6
7        # Check keywords
8        answer_lower = response.answer.lower()
9        keywords_found = sum(1 for k in case.expected_keywords if k in answer_lower)
10        keyword_score = keywords_found / len(case.expected_keywords)
11
12        # Check sources
13        has_sources = len(response.sources) > 0 if case.context_required else True
14
15        if keyword_score >= 0.5 and has_sources:
16            results["passed"] += 1
17        else:
18            results["failed"].append({
19                "question": case.question,
20                "keyword_score": keyword_score,
21                "has_sources": has_sources
22            })
23
24    results["accuracy"] = results["passed"] / results["total"]
25    return results

The result is an accuracy value plus a list of failures to review. Keywords are a simple, rough method. More serious projects also score how faithful the answer is to its sources, but the principle stays: a fixed set of questions, a repeatable measurement.

CI/CD Pipeline

Tests make sense when they run on their own on every push. GitHub Actions runs the steps on a clean machine, with Qdrant running alongside as a service:

1# .github/workflows/test.yml
2name: Test Suite
3
4on: [push, pull_request]
5
6jobs:
7  test:
8    runs-on: ubuntu-latest
9    services:
10      qdrant:
11        image: qdrant/qdrant
12        ports:
13          - 6333:6333
14
15    steps:
16      - uses: actions/checkout@v4
17      - uses: actions/setup-python@v5
18        with:
19          python-version: "3.12"
20
21      - name: Install dependencies
22        run: pip install -r requirements.txt
23
24      - name: Run unit tests
25        run: pytest tests/unit -v
26
27      - name: Run integration tests
28        run: pytest tests/integration -v
29        env:
30          QDRANT_URL: http://localhost:6333
31
32      - name: Run AI evaluation
33        run: python scripts/evaluate.py
34        env:
35          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}

The API key comes from secrets, never from the repository. Code coverage is measured by the pytest-cov plugin: pytest --cov=app lists which lines were never tested. You run a single file with pytest tests/unit/test_service.py, and many pipelines also build an image after the tests with docker build -t myapp:latest ., which we will come back to in the deployment lesson. My advice: run the AI evaluation separately from the unit tests, because it is slower and costs money. In the next lesson we will work on documentation.

Remember: tests are trial crossings of the river before you lead the whole expedition through it.

Spotted a mistake in this lesson?

Check yourself

Answer the questions from this lesson. Pick an answer to see right away whether it is correct.

  1. 1. What is at the base of the test pyramid?

  2. 2. What are pytest fixtures used for?

These are 2 of 3 questions for this lesson. Solve the rest in the game.

Hands-on tasks in the game

  • Vertical ordering

    Arrange the types of tests from fastest to slowest:

  • Code editor

    Write a unit test for a data processing function using pytest

  • Click in order

    Click the elements to build a pytest assert statement:

  • Horizontal ordering

    Arrange the Docker build command in the correct order:

  • Code editor

    Create a test using Mock to simulate an external API dependency

  • Vertical ordering

    Arrange the AI test pyramid layers from the base to the top:

Useful articles