By the end of this lesson, you will understand how linear regression models relationships between variables and be able to implement simple, multiple, and polynomial versions using Python.
What it is
Linear regression is a statistical method used to model the relationship between a dependent variable (target) and one or more independent variables (features). The core mental model is fitting a straight line (or hyperplane in higher dimensions) that minimizes the distance between predicted values and actual data points. Key related terms include coefficients (weights assigned to features), intercept (bias term), and residuals (errors).
Why it matters
- Interpretability: Coefficients directly show how much the target changes per unit change in a feature.
- Baseline Performance: It provides a quick benchmark for more complex models.
- Feature Selection: Helps identify which variables have significant predictive power.
- Foundation: Many advanced algorithms are extensions or regularizations of linear regression.
Syntax or steps
- Data Preparation: Split data into training and testing sets.
- Model Initialization: Create an instance of the regressor (e.g.,
LinearRegression). - Fitting: Train the model on the training data using
.fit(). - Prediction: Generate predictions on test data using
.predict(). - Evaluation: Calculate metrics like Mean Squared Error (MSE) or R-squared.
Example
import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error
# 1. Simple Linear Regression
X_simple = np.array([[1], [2], [3], [4]])
y_simple = np.array([2, 4, 6, 8])
model_simple = LinearRegression().fit(X_simple, y_simple)
print("Simple Coef:", model_simple.coef_) # Output: [2.]
# 2. Multiple Linear Regression
X_multi = np.array([[1, 1], [1, 2], [2, 2], [2, 3]])
y_multi = np.array([2, 3, 5, 7])
model_multi = LinearRegression().fit(X_multi, y_multi)
print("Multi Coefs:", model_multi.coef_) # Output: [1. 1.]
# 3. Polynomial Regression (Degree 2)
from sklearn.preprocessing import PolynomialFeatures
poly_reg = PolynomialFeatures(degree=2)
X_poly = poly_reg.fit_transform(X_simple)
model_poly = LinearRegression().fit(X_poly, y_simple)
print("Poly Intercept:", model_poly.intercept_)
This code demonstrates three variations. In simple regression, we fit a line to single-feature data. In multiple regression, we use two features to predict the target. For polynomial regression, we transform the original feature into squared terms before applying linear regression, allowing the model to fit curves.
Common mistakes
- Ignoring Multicollinearity: Highly correlated features can make coefficients unstable. Use correlation matrices to check.
- Assuming Linearity: If residuals show patterns, the relationship may not be linear. Consider polynomial or non-linear models.
- Not Scaling Features: While standard linear regression doesn't require scaling for accuracy, regularization techniques (like Ridge/Lasso) do. Always scale if using them.
- Overfitting with High Degrees: In polynomial regression, high degrees fit noise rather than signal. Keep degree low (2-3) unless justified.
When to use it
| Scenario | Recommended Approach |
|---|---|
| Relationship appears straight-line | Simple/Multiple Linear Regression |
| Relationship curves but stays smooth | Polynomial Regression |
| Complex, non-linear interactions | Tree-based models (Random Forest) |
| Need interpretability & speed | Linear Regression |
Practice
Guided Exercise: Modify the multiple regression example to add a third feature column to X_multi and adjust y_multi accordingly. Observe how the coefficient array length increases.
Challenge: Implement polynomial regression of degree 3 on the simple dataset. Compare the MSE against degree 1 and degree 2. Hint: Use mean_squared_error(y_simple, model.predict(poly_reg.fit_transform(X_simple))).
Quick check
Q: Why might you choose polynomial regression over simple linear regression?
A: When the relationship between the independent and dependent variables is curved rather than straight, polynomial regression can capture these non-linear trends by adding powers of the input features.
Summary
Linear regression is a versatile tool for modeling continuous outcomes. By extending from simple to multiple and polynomial forms, you can handle various data structures while maintaining interpretability. Always validate assumptions about linearity and feature independence to ensure robust results.