Back to Python Notes
Topic #265

Correlations

By the end of this lesson, you will be able to calculate and interpret correlation matrices in Python using pandas to identify linear relationships between numerical columns.

What it is

A correlation measures the strength and direction of a linear relationship between two variables. In data analysis, we often want to know if changes in one column predict changes in another. The most common metric is Pearson's r, which ranges from -1 (perfect negative correlation) to +1 (perfect positive correlation), with 0 indicating no linear relationship.

In Python, the pandas library provides the corr() method on DataFrames to compute these values efficiently. This creates a "correlation matrix," a square table showing the correlation coefficient for every pair of numeric columns.

Why it matters

  • Feature Selection: Helps identify redundant features in machine learning models (e.g., if two columns correlate at 0.95, you might drop one).
  • Hypothesis Testing: Quickly validates intuitive assumptions about how variables interact (e.g., does house size correlate with price?).
  • Data Cleaning: Reveals unexpected relationships that might indicate data entry errors or hidden patterns.
  • Visualization Prep: Correlation matrices are often used as input for heatmaps to visualize complex datasets.

Syntax or steps

The basic syntax requires a DataFrame containing only numeric data. Non-numeric columns are automatically ignored by default.

  1. Import pandas.
  2. Create or load a DataFrame.
  3. Call df.corr() to generate the matrix.
  4. Optionally, specify the method: method='pearson' (default), 'kendall', or 'spearman'.

Example

import pandas as pd

# Create a sample dataset
data = {
    'Hours_Studied': [2, 4, 6, 8, 10],
    'Exam_Score': [50, 60, 75, 85, 95],
    'Sleep_Hours': [8, 7, 6, 5, 4]
}
df = pd.DataFrame(data)

# Calculate correlation matrix
correlations = df.corr()

print(correlations)

Output Explanation:

Hours_Studied Exam_Score Sleep_Hours
Hours_Studied 1.000000 0.983871 -0.983871
Exam_Score 0.983871 1.000000 -0.983871
Sleep_Hours -0.983871 -0.983871 1.000000

Here, Hours_Studied and Exam_Score have a strong positive correlation (~0.98). Conversely, Sleep_Hours has a strong negative correlation with both, suggesting that in this specific synthetic dataset, more studying correlates with less sleep.

Common mistakes

  • Including non-numeric data: If your DataFrame contains strings or categories, corr() ignores them silently. Ensure you select only numeric columns first using df.select_dtypes(include='number').
  • Confusing correlation with causation: A high correlation does not mean one variable causes the other. Always investigate external factors.
  • Ignoring outliers: Pearson correlation is sensitive to outliers. A single extreme value can skew the result significantly. Consider using Spearman correlation (method='spearman') for rank-based robustness.
  • Assuming linearity: Correlation only measures linear relationships. Two variables might have a strong curved (non-linear) relationship but show a correlation near zero.

When to use it

Use df.corr() when exploring numerical relationships in tabular data. Compare it with simple plotting:

Method Best For Limitations
df.corr() Quantifying strength/direction across many variables quickly. Only detects linear trends; hard to read raw numbers without visualization.
sns.pairplot() Visualizing distributions and spotting non-linear patterns. Computationally expensive for large datasets; subjective interpretation.

Practice

Guided Exercise: Load the built-in Iris dataset using pd.read_csv('https://raw.githubusercontent.com/mwaskom/seaborn-data/master/iris.csv'). Filter for only numeric columns and print the correlation matrix. Identify which feature has the strongest positive correlation with petal_length.

Challenge: Modify the code to use method='spearman'. Does the ranking of correlations change? Why might this happen?

Quick check

Q: What does a correlation coefficient of -0.85 indicate?

A: It indicates a strong negative linear relationship: as one variable increases, the other tends to decrease significantly.

Summary

Pandas' corr() method is an essential tool for quantifying linear associations between numerical variables. While powerful for initial exploration, always remember that correlation does not imply causation and may miss non-linear patterns, so combine it with visualizations for a complete picture.

Want to go beyond the notes?

Join Coding Now Tech Institute's Python course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Correlations – FAQs

Quick answers about learning Correlations in Python.

This free note from Coding Now Tech Institute explains Correlations in Python — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Python topic on Coding Now Tech Institute, including Correlations, is 100% free with no signup required.
With focused practice, most students grasp Correlations in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now