Python course Β· Module 10: AI and Language Models
Large Language Models - Introduction
In this lesson5
Welcome to Module 10! Darwin here with Large Language Models (LLM) - the most powerful AI technology!
Throughout Module 9 we trained models that did one thing: they recognized a species or predicted a population. Every new task meant a new dataset and new training. Large language models work differently - a single model answers questions, writes code and translates texts, and you describe the task to it in plain language.
Safari analogy: An LLM is like a wise safari guide who has read every book about nature and can answer any question, write stories and help with tasks! It has one flaw, though: when it does not know something, it can tell a made-up story with complete confidence.
What are Large Language Models?
LLMs are AI models trained on huge amounts of text. They can understand and generate human language.
Here is a list of the things we most often put them to work on:
1# An LLM can:
2# - Answer questions
3# - Write code
4# - Translate languages
5# - Analyze text
6# - Help with programmingThis is not code to run, just a cheat sheet in comments. The model performs all these tasks with the same method - the only thing that changes is the instruction you send it, the prompt.
Popular LLM models
The market changes every few months, so treat the names below as a snapshot from the expedition, not a map for ever. You can always check the current list in the provider's documentation.
OpenAI GPT
- GPT-5, GPT-5 mini
- The most popular commercial models
- Great for a wide variety of tasks
GPT-5 is a reasoning model, and OpenAI is already releasing newer generations.
Anthropic Claude
- Claude Haiku 4.5, Sonnet 5, Opus 5.5 (earlier ones include Sonnet 4.6 and Opus 4.8)
- Safe and helpful AI
- Long context (from 200k to 1M tokens, depending on the model)
Open Source
- Llama (Meta)
- Mistral
- Falcon
More precisely, these are open-weight models: you can download their weights and run them on your own hardware, but the license, as in Llama's case, may set conditions.
Basic concepts
Tokens
A model does not read letters or whole words, but tokens - chunks of text that a tokenizer splits it into:
1# Tokens are "chunks" of text
2# "Hello world" = ["Hello", " world"] = 2 tokens
3
4# Token limits (context window, examples):
5# - GPT-5: 400k tokens
6# - Claude: 200k-1M tokens (depending on the model)
7# - Llama 3.1: 128k tokensNotice that the space belongs to the " world" token. Text in Polish usually uses more tokens than the same text in English, and in the API you pay per token, so this is a real cost.
Temperature
At every step the model picks the next token from many candidates. Temperature decides how boldly it reaches for the less likely ones. To make a call we need a client from the openai library, which reads the API key from the OPENAI_API_KEY environment variable:
1# Controls the "creativity" of responses
2# temperature=0 - nearly deterministic, highly repeatable responses
3# temperature=1 - more creative and varied
4
5from openai import OpenAI
6
7client = OpenAI()
8
9response = client.chat.completions.create(
10 model="gpt-4.1-mini",
11 messages=[{"role": "user", "content": "Write a poem"}],
12 temperature=0.7 # Moderate creativity
13)messages is a list of messages with a role and content, and you read the reply text through response.choices[0].message.content. Even with temperature=0 nobody guarantees identical answers, they are only highly repeatable. OpenAI's range is 0-2, Anthropic's is 0-1. An important trap: reasoning models such as GPT-5 reject any temperature other than the default 1, which is why the example uses gpt-4.1-mini. For data extraction and code I recommend a low temperature, for brainstorming a higher one.
Context Window
The context window is the limit of tokens the model takes in at once - your question, the conversation history and its answer:
1# How much text the model "remembers" in one conversation
2# GPT-5: 400k tokens
3# Claude: 200k-1M tokens (entire books!)
4
5# Longer context = better conversation memory"Memory" is a figure of speech here. Between API calls the model remembers nothing - it is your application that sends the whole conversation so far every time. When it exceeds the window, the oldest messages have to be shortened or dropped.
How do LLMs work?
Under the hood every answer is produced in four steps:
1# Simplified model of operation:
2# 1. Tokenization - text -> tokens
3# 2. Embedding - tokens -> numerical vectors
4# 3. Transformer - attention processing
5# 4. Generation - predicting the next token
6
7# LLMs do NOT understand text the way humans do
8# They are very good "prediction machines"The attention mechanism lets the model weigh, at every token, which earlier words matter. Step 4 repeats token by token until the answer is finished. This is where hallucinations come from: the model produces text that is probable, not verified, so always check important facts.
LLM applications
In programming
- Code generation
- Debugging
- Code review
- Documentation
In business
- Customer service chatbots
- Document analysis
- Report automation
In creative work
- Writing texts
- Brainstorming
- Translations
In the next lesson we will connect to the API and send our first real request, and in the following ones we will teach the model to use tools.
Remember: an LLM is a well-read guide who predicts the next word - listen to it gladly, but check the map yourself.
Spotted a mistake in this lesson?
Check yourself
Answer the questions from this lesson. Pick an answer to see right away whether it is correct.
1. What are Large Language Models?
2. What does the 'temperature' parameter control in an LLM?
These are 2 of 3 questions for this lesson. Solve the rest in the game.
Hands-on tasks in the game
- Vertical ordering
Arrange the steps of how an LLM works in the correct order:
- Vertical ordering
Arrange the steps in the correct order: