Back to Data Science Notes
Topic #61

Regression Analysis

By the end of this lesson, you will understand how to apply linear regression for predicting continuous values and logistic regression for binary classification, including their core assumptions and implementation steps.

What it is

Regression analysis is a statistical method used to estimate relationships among variables. In data science, it primarily serves two purposes: prediction and inference. Linear Regression models the relationship between a dependent variable and one or more independent variables by fitting a linear equation to observed data. It predicts continuous outcomes (e.g., house prices). Logistic Regression, despite its name, is a classification algorithm. It uses the logistic function (sigmoid) to model the probability of a binary outcome (e.g., spam vs. not spam). Key related terms include coefficients, intercepts, residuals, and decision boundaries.

Why it matters

  • Interpretability: Coefficients directly indicate the impact of each feature on the target variable.
  • Baseline Performance: These algorithms often serve as strong baselines before trying complex models like neural networks.
  • Efficiency: They are computationally inexpensive and fast to train, even on large datasets.
  • Versatility: Linear regression handles numerical prediction, while logistic regression solves fundamental classification problems.

Syntax or steps

The general workflow for both methods involves: 1. Data Preparation: Clean missing values and scale features if necessary. 2. Model Selection: Choose linear for continuous targets, logistic for binary categorical targets. 3. Fitting: The algorithm minimizes a loss function (Mean Squared Error for linear; Log Loss for logistic). 4. Evaluation: Use metrics like R-squared/RMSE for linear, and Accuracy/AUC-ROC for logistic.

Example

import numpy as np
from sklearn.linear_model import LinearRegression, LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error, accuracy_score

# 1. Generate synthetic data
np.random.seed(42)
X = np.random.rand(100, 1) * 10  # Feature: 0-10
y_linear = 2.5 * X.flatten() + np.random.normal(0, 2, 100) # Continuous target
y_logistic = (X.flatten() > 5).astype(int) # Binary target

# Split data
X_train, X_test, y_lin_train, y_lin_test = train_test_split(X, y_linear, test_size=0.2, random_state=42)
_, _, y_log_train, y_log_test = train_test_split(X, y_logistic, test_size=0.2, random_state=42)

# 2. Linear Regression
lin_model = LinearRegression()
lin_model.fit(X_train, y_lin_train)
pred_lin = lin_model.predict(X_test)
print(f"Linear MSE: {mean_squared_error(y_lin_test, pred_lin):.2f}")

# 3. Logistic Regression
log_model = LogisticRegression()
log_model.fit(X_train, y_log_train)
pred_log = log_model.predict(X_test)
print(f"Logistic Accuracy: {accuracy_score(y_log_test, pred_log):.2f}")
Explanation: We create simple data where `y_linear` depends on `X` with noise, and `y_logistic` is 1 if `X` exceeds 5. We split the data into training and testing sets. For linear regression, we fit the model and calculate Mean Squared Error (MSE). For logistic regression, we fit the classifier and check accuracy. Note that `sklearn` expects 2D arrays for features (`X`), hence the shape `(100, 1)`.

Common mistakes

  • Ignoring Assumptions: Linear regression assumes linearity and homoscedasticity. If residuals show patterns, the model is misspecified.
  • Using Linear for Classification: Predicting probabilities outside [0,1] or treating classes as numbers leads to poor performance. Always use Logistic Regression for binary tasks.
  • Feature Scaling Neglect: While tree-based models don't need scaling, gradient-descent-based regressions (like Logistic) converge faster when features are standardized.
  • Overfitting with High Dimensions: Adding too many irrelevant features increases variance. Use regularization (L1/L2) in logistic regression to penalize complexity.

When to use it

ScenarioRecommended ModelReason
Predicting temperatureLinear RegressionTarget is continuous and likely has a linear trend.
Detecting fraud (Yes/No)Logistic RegressionTarget is binary; outputs probability.
Complex non-linear patternsNeither (Use Trees/NN)Both assume linear decision boundaries or relationships.

Practice

Guided Exercise: Modify the code above to predict `y_linear` using two features instead of one. Observe how the coefficient changes. Challenge: Implement L2 regularization in the Logistic Regression model by setting the `C` parameter to 0.1. Compare the accuracy with the default `C=1.0`. Hint: Lower `C` means stronger regularization.

Quick check

Question: Why is Logistic Regression considered a classification algorithm despite having "regression" in its name? Answer: It regresses on the log-odds (linear combination of inputs) but transforms the output via the sigmoid function to produce probabilities for classification.

Summary

Linear and logistic regressions are foundational tools for modeling continuous and binary outcomes, respectively. Mastery of these algorithms provides insight into data relationships and establishes a baseline for evaluating more complex machine learning models.

Want to go beyond the notes?

Join Coding Now Tech Institute's Data Science course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Regression Analysis – FAQs

Quick answers about learning Regression Analysis in Data Science.

This free note from Coding Now Tech Institute explains Regression Analysis in Data Science — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Data Science topic on Coding Now Tech Institute, including Regression Analysis, is 100% free with no signup required.
With focused practice, most students grasp Regression Analysis in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now