Python course Β· Module 9: Machine Learning
Machine Learning - What is ML?
In this lesson4
Welcome to Module 9! Darwin here with Machine Learning!
Imagine you have to write a program that recognizes a lion in a camera-trap photo. You start with the rule "has a mane", but lionesses have no mane. You add fur color, but in a night-time infrared shot everything is grey. A week later you have a hundred conditions and the program still gets it wrong. Machine learning turns this process around: instead of writing rules, you show the computer examples.
Safari analogy: Machine Learning is like training safari animals - the algorithm learns from examples, just as a lion learns to hunt by watching its mother. The more examples, the better the predictions!
What is Machine Learning?
Machine Learning (ML) is a field of AI in which algorithms learn patterns from data without explicitly programmed rules.
The difference is easiest to see side by side. First a function with a hand-written rule, then a description of the ML approach:
1# Traditional programming
2def is_lion(animal):
3 if animal.has_mane and animal.color == 'golden':
4 return True
5 return False
6
7# Machine Learning
8# The model learns to recognize lions from THOUSANDS of images
9# without defining rules - it discovers them on its own!In the first part the programmer decides what makes a lion. In the second the model derives the rules itself from thousands of labeled photos. Notice what does not change: you still need a programmer, only their work shifts from writing conditions to preparing good data.
Types of machine learning
ML is divided into three main types, depending on what guidance the algorithm receives.
1. Supervised Learning
The model learns from labeled data - it knows the correct answers.
We store the data as two things: X holds the features, the numbers that describe each animal, and y holds the labels, the correct answers:
1# Training data with labels
2X_train = [
3 [180, 200, 4], # weight, length, legs
4 [5000, 400, 4],
5 [50, 150, 4],
6]
7y_train = ['lion', 'elephant', 'cheetah'] # labels
8
9# The model learns the mapping: features -> label
10# Then predicts for new dataEach row of X_train matches the label at the same position in y_train. Three examples are of course far too few - a real model needs hundreds or thousands of rows, but the shape of the data stays the same.
Applications:
- Classification: spam/not spam, animal species
- Regression: predicting prices, populations
Classification predicts a category, regression predicts a number. This distinction comes back in the next lesson, when we build our first real classifier.
2. Unsupervised Learning
The model discovers patterns in unlabeled data on its own.
This time there is no column with answers. There are only observations, and the algorithm looks for similarities among them:
1# Data without labels
2animals = [
3 [180, 200, 'savanna'],
4 [5000, 400, 'forest'],
5 [50, 150, 'savanna'],
6 [3000, 300, 'forest'],
7]
8
9# The model discovers groups (clusters) on its own
10# e.g. "savanna animals" vs "forest animals"The model does not know what a "savanna" is, it only notices that some rows are closer to each other than others. A practical note: most algorithms, KMeans for example, accept only numbers, so text such as 'savanna' has to be encoded as numbers first.
Applications:
- Clustering: grouping customers
- Dimensionality reduction: data compression
- Anomaly detection: fraud detection
3. Reinforcement Learning
An agent learns through trial and error - it receives rewards for good decisions.
Here there is no ready dataset. There is an agent, an environment, a list of actions and a reward:
1# Agent (Safari robot) in an environment
2# Actions: go_right, go_left, stop
3# Reward: +10 for finding an animal, -1 for each step
4
5# The agent learns the optimal exploration strategyThe penalty for each step pushes the agent toward short routes, and the big reward for finding an animal sets the goal. After thousands of attempts the agent picks the actions that bring the highest total reward.
Applications:
- Games (AlphaGo, Atari games)
- Robotics
- Recommendations
Basic ML concepts
Before we run our first model, let's gather the vocabulary we will use throughout the module:
1# Dataset
2X = [[1, 2], [3, 4], [5, 6]] # Features
3y = [0, 1, 0] # Labels
4
5# Training set - data for learning (70-80%)
6# Validation set - data for tuning (10-15%)
7# Test set - data for final evaluation (10-20%)
8
9# Model - an algorithm that learns from data
10# Prediction - forecasting for new data
11# Accuracy - % of correct predictionsThe most important rule hides in the data split: the model must not see the test set during training. Otherwise it is like a tracker taking an exam on the very tracks they practised on - the score looks great, but says nothing about new terrain.
Machine Learning workflow
Every ML project goes through the same stages. I will show them with the scikit-learn library (you import it as sklearn). First the data:
1# 1. Collect data
2data = load_safari_data()
3
4# 2. Prepare data (EDA, cleaning)
5X = data.drop('species', axis=1)
6y = data['species']load_safari_data() is only a placeholder name - in practice you would load the data with pd.read_csv, for example. data.drop('species', axis=1) returns the table without the species column, that is only the features, and data['species'] holds the labels. The original data table does not change.
Now the split into a training set and a test set:
1# 3. Split into train/test
2from sklearn.model_selection import train_test_split
3X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)test_size=0.2 sets 20% of the rows aside for testing. The split is random, so each run gives a slightly different result - add random_state=42 if you want reproducibility. I recommend always doing that while learning.
We choose a model and train it with the fit method:
1# 4. Select and train a model
2from sklearn.ensemble import RandomForestClassifier
3model = RandomForestClassifier()
4model.fit(X_train, y_train)A Random Forest is a forest of decision trees that vote on the answer. It is a very good choice to start with, because it works sensibly without tuning parameters. All scikit-learn models share the same fit method, so swapping the model is a one-line change.
Finally, evaluation and prediction:
1# 5. Evaluate the model
2accuracy = model.score(X_test, y_test)
3print(f"Accuracy: {accuracy:.2%}")
4
5# 6. Use the model for predictions
6predictions = model.predict(new_data)For a classifier, score returns the accuracy, the share of correct answers on the test set, and predict takes new rows in the same format as X. Accuracy alone can be misleading on imbalanced data - we will talk about better metrics in the coming lessons.
Next lesson: Supervised Learning in practice.
Remember: an ML model is a young tracker - it learns from past tracks, and you always test it on tracks it has never seen.
Spotted a mistake in this lesson?
Check yourself
Answer the questions from this lesson. Pick an answer to see right away whether it is correct.
1. What is Machine Learning?
2. What characterizes Supervised Learning?
These are 2 of 3 questions for this lesson. Solve the rest in the game.
Hands-on tasks in the game
- Vertical ordering
Arrange the Machine Learning workflow steps in the correct order:
- Vertical ordering
Arrange the steps in order:
- Code editor
Use train_test_split from sklearn to split X and y with test_size=0.2.
- Click in order
Click the buttons to import train_test_split: