By the end of this lesson, you will be able to calculate and interpret correlation matrices in Python using pandas to identify linear relationships between numerical columns.
What it is
A correlation measures the strength and direction of a linear relationship between two variables. In data analysis, we often want to know if changes in one column predict changes in another. The most common metric is Pearson's r, which ranges from -1 (perfect negative correlation) to +1 (perfect positive correlation), with 0 indicating no linear relationship.
In Python, the pandas library provides the corr() method on DataFrames to compute these values efficiently. This creates a "correlation matrix," a square table showing the correlation coefficient for every pair of numeric columns.
Why it matters
- Feature Selection: Helps identify redundant features in machine learning models (e.g., if two columns correlate at 0.95, you might drop one).
- Hypothesis Testing: Quickly validates intuitive assumptions about how variables interact (e.g., does house size correlate with price?).
- Data Cleaning: Reveals unexpected relationships that might indicate data entry errors or hidden patterns.
- Visualization Prep: Correlation matrices are often used as input for heatmaps to visualize complex datasets.
Syntax or steps
The basic syntax requires a DataFrame containing only numeric data. Non-numeric columns are automatically ignored by default.
- Import
pandas. - Create or load a DataFrame.
- Call
df.corr()to generate the matrix. - Optionally, specify the method:
method='pearson'(default),'kendall', or'spearman'.
Example
import pandas as pd
# Create a sample dataset
data = {
'Hours_Studied': [2, 4, 6, 8, 10],
'Exam_Score': [50, 60, 75, 85, 95],
'Sleep_Hours': [8, 7, 6, 5, 4]
}
df = pd.DataFrame(data)
# Calculate correlation matrix
correlations = df.corr()
print(correlations)
Output Explanation:
| Hours_Studied | Exam_Score | Sleep_Hours | |
|---|---|---|---|
| Hours_Studied | 1.000000 | 0.983871 | -0.983871 |
| Exam_Score | 0.983871 | 1.000000 | -0.983871 |
| Sleep_Hours | -0.983871 | -0.983871 | 1.000000 |
Here, Hours_Studied and Exam_Score have a strong positive correlation (~0.98). Conversely, Sleep_Hours has a strong negative correlation with both, suggesting that in this specific synthetic dataset, more studying correlates with less sleep.
Common mistakes
- Including non-numeric data: If your DataFrame contains strings or categories,
corr()ignores them silently. Ensure you select only numeric columns first usingdf.select_dtypes(include='number'). - Confusing correlation with causation: A high correlation does not mean one variable causes the other. Always investigate external factors.
- Ignoring outliers: Pearson correlation is sensitive to outliers. A single extreme value can skew the result significantly. Consider using Spearman correlation (
method='spearman') for rank-based robustness. - Assuming linearity: Correlation only measures linear relationships. Two variables might have a strong curved (non-linear) relationship but show a correlation near zero.
When to use it
Use df.corr() when exploring numerical relationships in tabular data. Compare it with simple plotting:
| Method | Best For | Limitations |
|---|---|---|
df.corr() |
Quantifying strength/direction across many variables quickly. | Only detects linear trends; hard to read raw numbers without visualization. |
sns.pairplot() |
Visualizing distributions and spotting non-linear patterns. | Computationally expensive for large datasets; subjective interpretation. |
Practice
Guided Exercise: Load the built-in Iris dataset using pd.read_csv('https://raw.githubusercontent.com/mwaskom/seaborn-data/master/iris.csv'). Filter for only numeric columns and print the correlation matrix. Identify which feature has the strongest positive correlation with petal_length.
Challenge: Modify the code to use method='spearman'. Does the ranking of correlations change? Why might this happen?
Quick check
Q: What does a correlation coefficient of -0.85 indicate?
A: It indicates a strong negative linear relationship: as one variable increases, the other tends to decrease significantly.
Summary
Pandas' corr() method is an essential tool for quantifying linear associations between numerical variables. While powerful for initial exploration, always remember that correlation does not imply causation and may miss non-linear patterns, so combine it with visualizations for a complete picture.