By the end of this lesson, you will be able to calculate and interpret covariance and correlation coefficients to quantify the linear relationship between two variables.
What it is
Covariance measures how much two random variables change together. If large values of one variable tend to occur with large values of another, the covariance is positive. If they move in opposite directions, it is negative. However, covariance is scale-dependent, making it hard to compare across different datasets.
Correlation (specifically Pearson’s r) standardizes covariance by dividing it by the product of the standard deviations of both variables. This results in a dimensionless value between -1 and +1, representing the strength and direction of the linear relationship.
Related terms: Linear dependence, scatter plot, standard deviation, normalization.
Why it matters
- Feature Selection: Helps identify which input variables are most relevant to the target variable in machine learning models.
- Redundancy Check: Detects highly correlated features that may cause multicollinearity issues in regression models.
- Pattern Discovery: Reveals hidden relationships in exploratory data analysis (e.g., stock prices vs. interest rates).
- Standardized Comparison: Allows comparison of relationship strengths across variables with different units (e.g., height vs. weight).
Syntax or steps
To compute these metrics in Python using pandas:
- Load your data into a DataFrame.
- Use
.cov()for the covariance matrix. - Use
.corr()for the correlation matrix. - Interpret the diagonal as variance (for cov) or 1.0 (for corr), and off-diagonals as pairwise relationships.
Example
import pandas as pd
import numpy as np
# Create sample data: Hours Studied vs. Exam Score
data = {
'Hours': [1, 2, 3, 4, 5],
'Score': [50, 60, 70, 80, 90]
}
df = pd.DataFrame(data)
# Calculate Covariance
cov_matrix = df.cov()
print("Covariance Matrix:")
print(cov_matrix)
# Calculate Correlation
corr_matrix = df.corr()
print("\nCorrelation Matrix:")
print(corr_matrix)
Explanation:
df.cov()returns a matrix where the entry at row 'Hours', column 'Score' represents the covariance between study hours and exam scores.df.corr()returns the Pearson correlation coefficient. In this perfectly linear example, the correlation is exactly 1.0.- The diagonal elements of the covariance matrix represent the variance of each individual column.
Common mistakes
- Assuming Causation: A high correlation does not imply that one variable causes the other. Always check for confounding variables.
- Ignoring Non-Linearity: Pearson correlation only detects linear relationships. Variables might have a strong curved relationship (e.g., quadratic) but show near-zero correlation.
- Misinterpreting Scale: Comparing raw covariance values across different datasets is meaningless because covariance depends on the magnitude of the data. Use correlation instead.
- Outliers: Both covariance and Pearson correlation are sensitive to outliers. A single extreme point can drastically skew the result.
When to use it
| Metric | Best Used For | Limitations |
|---|---|---|
| Covariance | Understanding joint variability within a specific dataset; calculating portfolio risk in finance. | Not comparable across different scales; difficult to interpret magnitude. |
| Correlation | Comparing relationship strength between different pairs of variables; feature selection. | Only captures linear associations; assumes normal distribution for significance testing. |
Practice
Guided Exercise: Modify the example above to include a third column called 'Sleep_Hours' with values [8, 7, 6, 5, 4]. Calculate the correlation matrix and observe how 'Sleep_Hours' correlates with 'Score'.
Challenge: Generate two random arrays using np.random.randn(100). Plot them using a scatter plot. Then, shuffle one array randomly and recalculate the correlation. What happens to the value?
Hint: The shuffled correlation should approach zero, indicating no linear relationship.
Quick check
Question: If Variable A has a correlation of 0.8 with Variable B, and Variable B has a correlation of 0.8 with Variable C, what is the correlation between A and C?
Answer: It cannot be determined solely from these values. While often high, the correlation between A and C could range significantly depending on the underlying structure of the data. Transitivity does not strictly apply to correlation coefficients.
Summary
Covariance tells you if two variables move together, while correlation tells you how strongly they move together in a linear fashion, normalized between -1 and 1. Use correlation for comparing relationships across different scales, but always visualize data to ensure the relationship is actually linear.