Back to Data Science Notes
Topic #66

Logistic Regression

By the end of this lesson, you will understand how logistic regression predicts probabilities for binary classification using the sigmoid function and implement a basic model in Python.

What it is

Logistic Regression is a statistical method used for binary classification. Despite its name, it is not used for predicting continuous values like linear regression; instead, it predicts the probability that an observation belongs to one of two classes (e.g., spam vs. not spam). The core mechanism is the sigmoid function, which maps any real-valued number into a value between 0 and 1. This output represents the probability of the positive class. If the probability exceeds a threshold (usually 0.5), the instance is classified as positive; otherwise, it is negative. Related terms include decision boundary, odds ratio, and maximum likelihood estimation.

Why it matters

  • Interpretability: Coefficients indicate the direction and magnitude of each feature's impact on the log-odds of the outcome.
  • Efficiency: It trains quickly even with large datasets compared to complex models like neural networks.
  • Baseline Performance: It serves as a strong benchmark for more sophisticated algorithms.
  • Probability Output: Unlike some classifiers that only give hard labels, logistic regression provides confidence scores.

Syntax or steps

The mathematical foundation involves calculating a linear combination of features ($z = w^Tx + b$) and passing it through the sigmoid function: $\sigma(z) = \frac{1}{1 + e^{-z}}$. In practice, libraries handle the optimization. The general workflow is:
  1. Prepare data (features $X$ and target $y$).
  2. Initialize the model.
  3. Fit the model to training data.
  4. Predict probabilities or classes on new data.

Example

Here is a minimal example using scikit-learn to classify iris flowers based on petal width.
from sklearn.datasets import load_iris
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
import numpy as np

# Load data and select binary classification subset (Setosa vs Versicolor)
data = load_iris()
X = data.data[:, [3]]  # Use only petal width
y = data.target
mask = y != 2          # Exclude Virginica
X, y = X[mask], y[mask]

# Split into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)

# Initialize and train the model
model = LogisticRegression()
model.fit(X_train, y_train)

# Predict probabilities for the first 5 test samples
probabilities = model.predict_proba(X_test[:5])
print("Predicted Probabilities:")
print(np.round(probabilities, 3))

# Predict class labels
predictions = model.predict(X_test[:5])
print("Predicted Classes:", predictions)
Explanation: We isolate petal width as our single feature. After splitting the data, LogisticRegression() finds the optimal weights. predict_proba returns the probability for each class (column 0 is Setosa, column 1 is Versicolor). predict applies the 0.5 threshold to return the final class label.

Common mistakes

  • Forgetting Feature Scaling: Logistic regression uses gradient descent; unscaled features can cause slow convergence. Always standardize inputs.
  • Misinterpreting Coefficients: Coefficients represent change in log-odds, not direct probability change. Exponentiate them to get odds ratios.
  • Assuming Linearity: If the relationship between features and log-odds is non-linear, basic logistic regression will underfit. Consider polynomial features.
  • Ignoring Imbalanced Data: Default settings may bias towards the majority class. Use class_weight='balanced' if necessary.

When to use it

Compare logistic regression with Linear Discriminant Analysis (LDA).
Feature Logistic Regression LDA
Assumptions Linear decision boundary Gaussian distribution of features per class
Data Size Works well with small to large data Better with small data if assumptions hold
Robustness More robust to outliers Sensitive to outliers due to mean/variance reliance
Use Logistic Regression when you need interpretable probabilities and suspect a linear relationship between features and log-odds. Use LDA when data is small and normally distributed within classes.

Practice

Guided Exercise: Modify the code above to print the model coefficients (model.coef_) and intercept (model.intercept_). Interpret what a positive coefficient means for petal width regarding the Versicolor class.
Challenge: Add a second feature (sepal length) to the model. Does the accuracy improve? Hint: Check model.score(X_test, y_test).

Quick check

Question: Why do we use the sigmoid function instead of a step function for classification? Answer: The sigmoid function is differentiable, allowing us to use gradient-based optimization methods to find the best parameters. A step function has zero derivative almost everywhere, making optimization impossible.

Summary

Logistic regression transforms linear combinations of features into probabilities via the sigmoid function, offering a balance of speed, interpretability, and performance for binary tasks. Mastering it requires understanding both the mathematical mapping of log-odds and the practical necessity of feature scaling.

Want to go beyond the notes?

Join Coding Now Tech Institute's Data Science course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Logistic Regression – FAQs

Quick answers about learning Logistic Regression in Data Science.

This free note from Coding Now Tech Institute explains Logistic Regression in Data Science — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Data Science topic on Coding Now Tech Institute, including Logistic Regression, is 100% free with no signup required.
With focused practice, most students grasp Logistic Regression in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now