Back to Data Science Notes
Topic #60

Correlation & Covariance

By the end of this lesson, you will be able to calculate and interpret covariance and correlation coefficients to quantify the linear relationship between two variables.

What it is

Covariance measures how much two random variables change together. If large values of one variable tend to occur with large values of another, the covariance is positive. If they move in opposite directions, it is negative. However, covariance is scale-dependent, making it hard to compare across different datasets.

Correlation (specifically Pearson’s r) standardizes covariance by dividing it by the product of the standard deviations of both variables. This results in a dimensionless value between -1 and +1, representing the strength and direction of the linear relationship.

Related terms: Linear dependence, scatter plot, standard deviation, normalization.

Why it matters

  • Feature Selection: Helps identify which input variables are most relevant to the target variable in machine learning models.
  • Redundancy Check: Detects highly correlated features that may cause multicollinearity issues in regression models.
  • Pattern Discovery: Reveals hidden relationships in exploratory data analysis (e.g., stock prices vs. interest rates).
  • Standardized Comparison: Allows comparison of relationship strengths across variables with different units (e.g., height vs. weight).

Syntax or steps

To compute these metrics in Python using pandas:

  1. Load your data into a DataFrame.
  2. Use .cov() for the covariance matrix.
  3. Use .corr() for the correlation matrix.
  4. Interpret the diagonal as variance (for cov) or 1.0 (for corr), and off-diagonals as pairwise relationships.

Example

import pandas as pd
import numpy as np

# Create sample data: Hours Studied vs. Exam Score
data = {
    'Hours': [1, 2, 3, 4, 5],
    'Score': [50, 60, 70, 80, 90]
}
df = pd.DataFrame(data)

# Calculate Covariance
cov_matrix = df.cov()
print("Covariance Matrix:")
print(cov_matrix)

# Calculate Correlation
corr_matrix = df.corr()
print("\nCorrelation Matrix:")
print(corr_matrix)

Explanation:

  • df.cov() returns a matrix where the entry at row 'Hours', column 'Score' represents the covariance between study hours and exam scores.
  • df.corr() returns the Pearson correlation coefficient. In this perfectly linear example, the correlation is exactly 1.0.
  • The diagonal elements of the covariance matrix represent the variance of each individual column.

Common mistakes

  • Assuming Causation: A high correlation does not imply that one variable causes the other. Always check for confounding variables.
  • Ignoring Non-Linearity: Pearson correlation only detects linear relationships. Variables might have a strong curved relationship (e.g., quadratic) but show near-zero correlation.
  • Misinterpreting Scale: Comparing raw covariance values across different datasets is meaningless because covariance depends on the magnitude of the data. Use correlation instead.
  • Outliers: Both covariance and Pearson correlation are sensitive to outliers. A single extreme point can drastically skew the result.

When to use it

Metric Best Used For Limitations
Covariance Understanding joint variability within a specific dataset; calculating portfolio risk in finance. Not comparable across different scales; difficult to interpret magnitude.
Correlation Comparing relationship strength between different pairs of variables; feature selection. Only captures linear associations; assumes normal distribution for significance testing.

Practice

Guided Exercise: Modify the example above to include a third column called 'Sleep_Hours' with values [8, 7, 6, 5, 4]. Calculate the correlation matrix and observe how 'Sleep_Hours' correlates with 'Score'.

Challenge: Generate two random arrays using np.random.randn(100). Plot them using a scatter plot. Then, shuffle one array randomly and recalculate the correlation. What happens to the value?

Hint: The shuffled correlation should approach zero, indicating no linear relationship.

Quick check

Question: If Variable A has a correlation of 0.8 with Variable B, and Variable B has a correlation of 0.8 with Variable C, what is the correlation between A and C?

Answer: It cannot be determined solely from these values. While often high, the correlation between A and C could range significantly depending on the underlying structure of the data. Transitivity does not strictly apply to correlation coefficients.

Summary

Covariance tells you if two variables move together, while correlation tells you how strongly they move together in a linear fashion, normalized between -1 and 1. Use correlation for comparing relationships across different scales, but always visualize data to ensure the relationship is actually linear.

Want to go beyond the notes?

Join Coding Now Tech Institute's Data Science course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Correlation & Covariance – FAQs

Quick answers about learning Correlation & Covariance in Data Science.

This free note from Coding Now Tech Institute explains Correlation & Covariance in Data Science — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Data Science topic on Coding Now Tech Institute, including Correlation & Covariance, is 100% free with no signup required.
With focused practice, most students grasp Correlation & Covariance in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now