Understand how skewness and kurtosis describe the shape of data distributions, and learn to identify when a normal distribution assumption is valid.
What it is
Distribution shape metrics quantify how data deviates from a perfect bell curve. Skewness measures asymmetry: positive skew means a long right tail (mean > median), negative skew means a long left tail (mean < median). Kurtosis measures tail heaviness relative to a normal distribution. High kurtosis (leptokurtic) indicates heavy tails and outliers; low kurtosis (platykurtic) indicates light tails. The Normal Distribution has zero skewness and excess kurtosis of zero, serving as the baseline for many statistical tests.
Why it matters
- Model Assumptions: Many algorithms (linear regression, t-tests) assume normality; violations lead to unreliable p-values.
- Risk Management: Financial returns often exhibit high kurtosis ("fat tails"), underestimating extreme loss probabilities if ignored.
- Data Cleaning: Extreme skewness may indicate errors or require transformation (e.g., log transform) before modeling.
- Feature Engineering: Understanding shape helps select appropriate scaling methods (StandardScaler vs. RobustScaler).
Syntax or steps
To calculate these metrics in Python using scipy.stats:
- Import necessary libraries (
numpy,scipy.stats). - Create or load your dataset array.
- Use
stats.skew()to compute skewness. - Use
stats.kurtosis()to compute excess kurtosis (note: SciPy returns excess kurtosis by default, where Normal = 0). - Interpret values: Skew near 0 is symmetric; Kurtosis near 0 is mesokurtic (normal-like).
Example
import numpy as np
from scipy import stats
# Generate synthetic data
np.random.seed(42)
normal_data = np.random.normal(loc=0, scale=1, size=1000)
skewed_data = np.random.exponential(scale=2, size=1000) # Right-skewed
# Calculate metrics
print("Normal Distribution:")
print(f"Skewness: {stats.skew(normal_data):.4f}")
print(f"Excess Kurtosis: {stats.kurtosis(normal_data):.4f}")
print("\nExponential Distribution (Right-Skewed):")
print(f"Skewness: {stats.skew(skewed_data):.4f}")
print(f"Excess Kurtosis: {stats.kurtosis(skewed_data):.4f}")
Explanation: The code generates two datasets. The first follows a standard normal distribution. The second uses an exponential distribution, which is inherently right-skewed with heavier tails than normal. We use stats.skew() and stats.kurtosis() to quantify these differences. Note that kurtosis() in SciPy calculates excess kurtosis (Kurtosis - 3), so a value near 0 confirms normality.
Common mistakes
- Confusing Sample vs. Population: Small sample sizes produce unstable skew/kurtosis estimates. Use bootstrapping or larger samples for reliability.
- Ignoring Outliers: A single extreme outlier can drastically inflate skewness and kurtosis, misleadingly suggesting non-normality. Check boxplots first.
- Misinterpreting Kurtosis: Remember that "high kurtosis" refers to tail weight, not peak height. A sharp peak with thin tails is possible but rare in real-world data compared to fat-tailed distributions.
- Assuming Symmetry Equals Normality: A distribution can be symmetric (skew=0) but have heavy tails (high kurtosis), violating normal assumptions despite looking like a bell curve visually.
When to use it
| Metric | Best For | Alternative/Complement |
|---|---|---|
| Skewness/Kurtosis | Quick numerical summary of shape deviations. | Q-Q Plots (visual confirmation) |
| Shapiro-Wilk Test | Formal hypothesis testing for normality (small n). | Anderson-Darling (large n) |
| Log Transformation | Correcting strong positive skew. | Box-Cox (automated power selection) |
Practice
Guided Exercise: Load the 'tips' dataset from Seaborn. Calculate skewness and kurtosis for the 'total_bill' column. Does the result suggest a normal distribution?
Challenge: Apply a natural log transformation to 'total_bill'. Recalculate skewness. Did it move closer to zero? Why might this help linear regression?
Hint: Restaurant bills are typically right-skewed (most people spend little, few spend much). Log transforms compress large values, reducing skew.
Quick check
Question: If a dataset has a skewness of 2.5 and excess kurtosis of 8.0, what does this imply about its shape compared to a normal distribution?
Answer: It implies the data is strongly right-skewed (long tail on the right) and has very heavy tails (more outliers/extreme values) than a normal distribution.
Summary
Skewness and kurtosis provide essential numerical insights into data shape, revealing asymmetries and tail risks that mean and standard deviation miss. Always validate these metrics with visual tools like histograms or Q-Q plots before assuming normality in statistical modeling.