By the end of this lesson, you will understand how logistic regression predicts probabilities for binary classification using the sigmoid function and implement a basic model in Python.
What it is
Logistic Regression is a statistical method used for binary classification. Despite its name, it is not used for predicting continuous values like linear regression; instead, it predicts the probability that an observation belongs to one of two classes (e.g., spam vs. not spam). The core mechanism is the sigmoid function, which maps any real-valued number into a value between 0 and 1. This output represents the probability of the positive class. If the probability exceeds a threshold (usually 0.5), the instance is classified as positive; otherwise, it is negative. Related terms include decision boundary, odds ratio, and maximum likelihood estimation.Why it matters
- Interpretability: Coefficients indicate the direction and magnitude of each feature's impact on the log-odds of the outcome.
- Efficiency: It trains quickly even with large datasets compared to complex models like neural networks.
- Baseline Performance: It serves as a strong benchmark for more sophisticated algorithms.
- Probability Output: Unlike some classifiers that only give hard labels, logistic regression provides confidence scores.
Syntax or steps
The mathematical foundation involves calculating a linear combination of features ($z = w^Tx + b$) and passing it through the sigmoid function: $\sigma(z) = \frac{1}{1 + e^{-z}}$. In practice, libraries handle the optimization. The general workflow is:- Prepare data (features $X$ and target $y$).
- Initialize the model.
- Fit the model to training data.
- Predict probabilities or classes on new data.
Example
Here is a minimal example usingscikit-learn to classify iris flowers based on petal width.
from sklearn.datasets import load_iris
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
import numpy as np
# Load data and select binary classification subset (Setosa vs Versicolor)
data = load_iris()
X = data.data[:, [3]] # Use only petal width
y = data.target
mask = y != 2 # Exclude Virginica
X, y = X[mask], y[mask]
# Split into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
# Initialize and train the model
model = LogisticRegression()
model.fit(X_train, y_train)
# Predict probabilities for the first 5 test samples
probabilities = model.predict_proba(X_test[:5])
print("Predicted Probabilities:")
print(np.round(probabilities, 3))
# Predict class labels
predictions = model.predict(X_test[:5])
print("Predicted Classes:", predictions)
Explanation: We isolate petal width as our single feature. After splitting the data, LogisticRegression() finds the optimal weights. predict_proba returns the probability for each class (column 0 is Setosa, column 1 is Versicolor). predict applies the 0.5 threshold to return the final class label.
Common mistakes
- Forgetting Feature Scaling: Logistic regression uses gradient descent; unscaled features can cause slow convergence. Always standardize inputs.
- Misinterpreting Coefficients: Coefficients represent change in log-odds, not direct probability change. Exponentiate them to get odds ratios.
- Assuming Linearity: If the relationship between features and log-odds is non-linear, basic logistic regression will underfit. Consider polynomial features.
- Ignoring Imbalanced Data: Default settings may bias towards the majority class. Use
class_weight='balanced'if necessary.
When to use it
Compare logistic regression with Linear Discriminant Analysis (LDA).| Feature | Logistic Regression | LDA |
|---|---|---|
| Assumptions | Linear decision boundary | Gaussian distribution of features per class |
| Data Size | Works well with small to large data | Better with small data if assumptions hold |
| Robustness | More robust to outliers | Sensitive to outliers due to mean/variance reliance |
Practice
Guided Exercise: Modify the code above to print the model coefficients (model.coef_) and intercept (model.intercept_). Interpret what a positive coefficient means for petal width regarding the Versicolor class.
Challenge: Add a second feature (sepal length) to the model. Does the accuracy improve? Hint: Check
model.score(X_test, y_test).