By the end of this lesson, you will understand the three main types of machine learning, the standard workflow for building models, and why splitting data into training and testing sets is critical for evaluating performance.
What it is
Machine Learning (ML) is a subset of artificial intelligence where computers learn patterns from data rather than following explicit instructions. The core mental model is "learning by example." Instead of writing rules likeif temperature > 30 then predict_hot, you provide examples of temperatures and whether it was hot, letting the algorithm discover the threshold itself.
The three primary types are:
- Supervised Learning: Predicting a known output (label) from input features (e.g., predicting house prices).
- Unsupervised Learning: Finding hidden structures in unlabeled data (e.g., customer segmentation).
- Reinforcement Learning: Learning optimal actions through trial and error with rewards (e.g., game playing).
Why it matters
- Automation: Handles complex tasks like fraud detection or recommendation engines at scale.
- Pattern Discovery: Reveals insights humans might miss in large datasets.
- Adaptability: Models can improve as new data becomes available.
- Prediction: Enables forecasting future outcomes based on historical trends.
Syntax or steps
The standard ML workflow follows these steps:- Data Collection: Gather relevant data.
- Preprocessing: Clean data, handle missing values, and encode categorical variables.
- Splitting: Divide data into Training and Testing sets.
- Model Selection: Choose an algorithm (e.g., Linear Regression, Decision Tree).
- Training: Fit the model to the training data.
- Evaluation: Test the model on unseen test data.
Example
Here is a minimal Python example usingscikit-learn to demonstrate supervised learning with a train-test split.
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error
import numpy as np
# 1. Create synthetic data
X = np.array([[1], [2], [3], [4], [5]]) # Features
y = np.array([2, 4, 6, 8, 10]) # Labels (y = 2x)
# 2. Split data: 80% training, 20% testing
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# 3. Initialize and train the model
model = LinearRegression()
model.fit(X_train, y_train)
# 4. Make predictions on test set
predictions = model.predict(X_test)
# 5. Evaluate performance
mse = mean_squared_error(y_test, predictions)
print(f"Test MSE: {mse:.4f}")
print(f"Coefficient: {model.coef_[0]:.2f}")
Part-by-part explanation:
train_test_split: Randomly divides data so the model never sees the test data during training.LinearRegression(): Creates an empty model object..fit(): Trains the model by finding the best line that fitsX_trainandy_train..predict(): Uses the learned pattern to guess values forX_test.mean_squared_error: Calculates how far off the predictions were from actual values.
Common mistakes
- Data Leakage: Preprocessing the entire dataset before splitting. Always split first, then preprocess each part separately.
- Overfitting: A model performs perfectly on training data but poorly on test data. Fix by simplifying the model or adding more data.
- Ignoring Scale: Some algorithms require normalized features. Use
StandardScalerif needed. - Small Test Sets: If the test set is too small, evaluation metrics become unreliable. Ensure sufficient samples remain after splitting.
When to use it
Compare Supervised vs. Unsupervised learning based on your goal.| Scenario | Type | Reason |
|---|---|---|
| Predicting sales revenue | Supervised | You have historical labeled data (past sales). |
| Grouping similar users | Unsupervised | No predefined groups exist; find natural clusters. |
| Detecting anomalies | Unsupervised | Normal behavior defines the baseline; outliers are rare. |
Practice
Guided Exercise: Modify the code above to changetest_size to 0.5. Observe how the Mean Squared Error changes. Why might it increase?
Challenge: Add a second feature to X (e.g., [1, 10], [2, 20]...) and retrain. Does the coefficient change? Hint: Check if the relationship remains linear.