Python course Β· Module 12: Final Project

Planning an AI Project

6 min read
In this lesson3

Welcome to the finale of Python Safari! Picture an expedition that heads into the savanna with no map, no supply list and no goal. Three days in, someone asks "why are we actually here?" and nobody can answer. Software projects work exactly the same way: they rarely fail because of bad code, they fail because nobody wrote down which problem they solve and when they can be called finished. In this module you will combine all the skills you have gained into one complete project, and we start with the expedition map.

What You Will Learn

  • how to capture a project's problem and success metrics in Python code instead of keeping them in your head
  • how to write user stories with acceptance criteria
  • how to estimate tasks in story points and arrange them into sprints
  • how to justify your choice of tech stack

Project Planning Methodology

1. Defining the Problem

Before we write the first line of the application, let's name the kind of project. For a closed list of variants, Python has the Enum type from the enum module: every member of the enumeration has a name and a value, and a typo in a name fails immediately instead of slipping through silently.

1from dataclasses import dataclass
2from enum import Enum
3
4class ProjectType(Enum):
5    WEB_APP = "web_application"
6    API = "api_service"
7    ML_PIPELINE = "ml_pipeline"
8    RAG_SYSTEM = "rag_system"
9    AUTOMATION = "automation"

So far this is just a dictionary of expedition types, nothing happens yet. Now let's describe the expedition itself. The @dataclass decorator from the dataclasses module generates __init__, __repr__ and comparison methods from the fields with type annotations, so a data class fits in a few lines.

1@dataclass
2class ProjectDefinition:
3    """Project definition."""
4    name: str
5    problem_statement: str
6    target_users: list[str]
7    success_metrics: list[str]
8    project_type: ProjectType
9    tech_stack: list[str]

Annotations such as list[str] are documentation and a hint for your editor and for tools like mypy. Python does not enforce them at runtime, so the dataclass will not reject a wrong type. Let's look at a filled-in project definition:

1# Example
2capstone_project = ProjectDefinition(
3    name="AI Document Assistant",
4    problem_statement="Users waste time searching for information in documents",
5    target_users=["analysts", "lawyers", "researchers"],
6    success_metrics=[
7        "70% reduction in search time",
8        "Answer accuracy > 90%",
9        "Response time < 2s"
10    ],
11    project_type=ProjectType.RAG_SYSTEM,
12    tech_stack=["Python", "FastAPI", "LlamaIndex", "Qdrant", "React"]
13)

The success metrics matter most here. "Response time < 2s" can be measured, "a fast app" cannot. These are targets we set for ourselves, not ready-made data: next to each metric, write down right away how you will measure it.

2. User Stories

A user story describes a feature from the user's perspective, in a fixed format: "as a [who] I want [what] so that [why]". Acceptance criteria say when the story is done. In code we will capture this with another dataclass:

1@dataclass
2class UserStory:
3    """User story in Agile format."""
4    as_a: str
5    i_want: str
6    so_that: str
7    acceptance_criteria: list[str]
8    priority: int  # 1-5, 1 = highest

The priority field is a plain number, and the comment sets the agreement that 1 means the highest priority. Now two concrete stories, one for an analyst and one for an administrator:

1stories = [
2    UserStory(
3        as_a="analyst",
4        i_want="ask questions about documents in natural language",
5        so_that="I can quickly find the information I need",
6        acceptance_criteria=[
7            "System understands questions in English",
8            "Response includes sources",
9            "Response time < 3s"
10        ],
11        priority=1
12    ),
13    UserStory(
14        as_a="administrator",
15        i_want="easily add new documents",
16        so_that="the knowledge base stays up to date",
17        acceptance_criteria=[
18            "Upload via drag & drop",
19            "Support for PDF, DOCX, TXT",
20            "Automatic indexing"
21        ],
22        priority=2
23    )
24]

Notice that the stories say nothing about technology. The analyst does not know what a vector database is, and does not need to. Criteria such as "Answer includes sources" will later become tests, so write them in a way that can be checked.

3. Estimation and Sprint Planning

In Scrum, work is split into sprints, fixed-length periods of one month or less. In practice, teams most often choose two weeks. Tasks are estimated in story points, a measure of relative complexity rather than hours. A popular scale is the Fibonacci sequence (1, 2, 3, 5, 8, 13), because the growing gaps between numbers remind you that big tasks are less and less predictable.

1from datetime import datetime, timedelta
2
3@dataclass
4class Task:
5    name: str
6    story_points: int  # Fibonacci: 1, 2, 3, 5, 8, 13
7    dependencies: list[str]
8    assigned_to: str = ""

Task also remembers its dependencies, the names of tasks that must be finished first. The sprint gets two properties computed on the fly thanks to @property: the end date and the sum of points.

1@dataclass
2class Sprint:
3    number: int
4    start_date: datetime
5    tasks: list[Task]
6    velocity: int = 20  # Story points per sprint
7
8    @property
9    def end_date(self) -> datetime:
10        return self.start_date + timedelta(days=14)
11
12    @property
13    def total_points(self) -> int:
14        return sum(t.story_points for t in self.tasks)

velocity is the number of points the team actually completes in one sprint. We assume 20, but you will only learn the real value after a few sprints. Let's see the plan for the first two stages of the expedition:

1# Project plan
2sprints = [
3    Sprint(1, datetime(2024, 1, 1), [
4        Task("Project setup", 2, []),
5        Task("CI/CD configuration", 3, ["Project setup"]),
6        Task("Basic API", 5, ["Project setup"]),
7        Task("Vector DB integration", 5, ["Basic API"]),
8    ]),
9    Sprint(2, datetime(2024, 1, 15), [
10        Task("RAG Pipeline", 8, ["Vector DB integration"]),
11        Task("Chat interface", 5, ["Basic API"]),
12        Task("Document upload", 5, ["RAG Pipeline"]),
13    ])
14]

The order follows the dependencies: setup first, then the API, and only then RAG. The first sprint has 15 points, the second 18, so both fit within the assumed velocity and leave room for surprises.

Choosing the Technology Stack

The last part of the map is the technology decisions. Write down the reasons next to every choice, because in six months someone (maybe you) will ask why Qdrant of all things.

1tech_decisions = {
2    "backend": {
3        "choice": "FastAPI",
4        "reasons": ["Async", "Type hints", "Auto docs", "Performance"]
5    },
6    "vector_db": {
7        "choice": "Qdrant",
8        "reasons": ["Self-hosted", "Filtering", "Rust performance"]
9    },
10    "llm": {
11        "choice": "OpenAI API",
12        "reasons": ["Quality", "Function calling", "Reliable"]
13    },
14    "rag_framework": {
15        "choice": "LlamaIndex",
16        "reasons": ["Flexibility", "Integrations", "Community"]
17    }
18}

For the language model we deliberately wrote down the provider, not a specific model. Model names change every few months and older versions get retired, so pick the current model from the provider's documentation on the day the project starts. My advice: keep this list of decisions in the repository next to the code, not in your notes, because then it lives together with the project.

The full project lifecycle is planning, implementation, testing, deployment and monitoring, in that order. You install dependencies from a requirements.txt file with pip install -r requirements.txt, and you pin a specific package version with the name==version form, for example pip install numpy==1.26.4. In the next lesson we will design the system architecture that turns this plan into a code structure.

Remember: a good expedition plan is half the success, because a map drawn before departure saves weeks of wandering across the savanna.

Spotted a mistake in this lesson?

Check yourself

Answer the questions from this lesson. Pick an answer to see right away whether it is correct.

  1. 1. What is a User Story in Agile methodology?

  2. 2. How long does a typical Sprint last in Scrum?

These are 2 of 3 questions for this lesson. Solve the rest in the game.

Hands-on tasks in the game

  • Vertical ordering

    Arrange the stages of an AI project lifecycle in the correct order:

  • Code editor

    Create a project definition file with the basic structure of an AI project

  • Horizontal ordering

    Arrange the pip install command for a specific package version in the correct order:

  • Click in order

    Click the elements to build a Python dataclass definition in the correct order:

Useful articles