By the end of this lesson, you will understand how to apply linear regression for predicting continuous values and logistic regression for binary classification, including their core assumptions and implementation steps.
What it is
Regression analysis is a statistical method used to estimate relationships among variables. In data science, it primarily serves two purposes: prediction and inference. Linear Regression models the relationship between a dependent variable and one or more independent variables by fitting a linear equation to observed data. It predicts continuous outcomes (e.g., house prices). Logistic Regression, despite its name, is a classification algorithm. It uses the logistic function (sigmoid) to model the probability of a binary outcome (e.g., spam vs. not spam). Key related terms include coefficients, intercepts, residuals, and decision boundaries.Why it matters
- Interpretability: Coefficients directly indicate the impact of each feature on the target variable.
- Baseline Performance: These algorithms often serve as strong baselines before trying complex models like neural networks.
- Efficiency: They are computationally inexpensive and fast to train, even on large datasets.
- Versatility: Linear regression handles numerical prediction, while logistic regression solves fundamental classification problems.
Syntax or steps
The general workflow for both methods involves: 1. Data Preparation: Clean missing values and scale features if necessary. 2. Model Selection: Choose linear for continuous targets, logistic for binary categorical targets. 3. Fitting: The algorithm minimizes a loss function (Mean Squared Error for linear; Log Loss for logistic). 4. Evaluation: Use metrics like R-squared/RMSE for linear, and Accuracy/AUC-ROC for logistic.Example
import numpy as np
from sklearn.linear_model import LinearRegression, LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error, accuracy_score
# 1. Generate synthetic data
np.random.seed(42)
X = np.random.rand(100, 1) * 10 # Feature: 0-10
y_linear = 2.5 * X.flatten() + np.random.normal(0, 2, 100) # Continuous target
y_logistic = (X.flatten() > 5).astype(int) # Binary target
# Split data
X_train, X_test, y_lin_train, y_lin_test = train_test_split(X, y_linear, test_size=0.2, random_state=42)
_, _, y_log_train, y_log_test = train_test_split(X, y_logistic, test_size=0.2, random_state=42)
# 2. Linear Regression
lin_model = LinearRegression()
lin_model.fit(X_train, y_lin_train)
pred_lin = lin_model.predict(X_test)
print(f"Linear MSE: {mean_squared_error(y_lin_test, pred_lin):.2f}")
# 3. Logistic Regression
log_model = LogisticRegression()
log_model.fit(X_train, y_log_train)
pred_log = log_model.predict(X_test)
print(f"Logistic Accuracy: {accuracy_score(y_log_test, pred_log):.2f}")
Explanation: We create simple data where `y_linear` depends on `X` with noise, and `y_logistic` is 1 if `X` exceeds 5. We split the data into training and testing sets. For linear regression, we fit the model and calculate Mean Squared Error (MSE). For logistic regression, we fit the classifier and check accuracy. Note that `sklearn` expects 2D arrays for features (`X`), hence the shape `(100, 1)`.
Common mistakes
- Ignoring Assumptions: Linear regression assumes linearity and homoscedasticity. If residuals show patterns, the model is misspecified.
- Using Linear for Classification: Predicting probabilities outside [0,1] or treating classes as numbers leads to poor performance. Always use Logistic Regression for binary tasks.
- Feature Scaling Neglect: While tree-based models don't need scaling, gradient-descent-based regressions (like Logistic) converge faster when features are standardized.
- Overfitting with High Dimensions: Adding too many irrelevant features increases variance. Use regularization (L1/L2) in logistic regression to penalize complexity.
When to use it
| Scenario | Recommended Model | Reason |
|---|---|---|
| Predicting temperature | Linear Regression | Target is continuous and likely has a linear trend. |
| Detecting fraud (Yes/No) | Logistic Regression | Target is binary; outputs probability. |
| Complex non-linear patterns | Neither (Use Trees/NN) | Both assume linear decision boundaries or relationships. |