By the end of this lesson, you will be able to calculate and interpret key descriptive statistics—mean, median, mode, variance, and standard deviation—to summarize data distributions effectively.
What it is
Descriptive statistics are numerical measures that summarize the main features of a collection of quantitative data. They provide a snapshot of central tendency (where the data clusters) and dispersion (how spread out the data is). The mean is the arithmetic average, the median is the middle value when sorted, and the mode is the most frequent value. Variance measures the average squared deviation from the mean, while standard deviation is the square root of variance, expressed in the same units as the original data. Related terms include range, interquartile range (IQR), and skewness.Why it matters
- Data Cleaning: Identifying outliers or errors by checking if values fall outside expected ranges defined by mean and standard deviation.
- Feature Engineering: Normalizing data for machine learning models often requires knowing the mean and standard deviation of training features.
- Reporting: Providing concise summaries of large datasets for stakeholders who cannot review raw rows.
- Distribution Analysis: Comparing mean vs. median helps detect skewness; if mean > median, the distribution is likely right-skewed.
Syntax or steps
To compute these metrics manually: 1. Sort the data for median and mode identification. 2. Calculate the sum of all values divided by count $n$ for the mean ($\mu$). 3. Find the difference between each value and the mean, square them, sum them, and divide by $n-1$ (sample variance) or $n$ (population variance). 4. Take the square root of the variance for standard deviation. In Python, using the `statistics` module or `numpy`, these operations are abstracted into single function calls.Example
import numpy as np
# Sample dataset: exam scores
scores = np.array([85, 90, 78, 92, 88, 76, 90, 85])
# Central Tendency
mean_val = np.mean(scores)
median_val = np.median(scores)
# Note: NumPy doesn't have a direct mode function, so we use scipy or manual counting
unique, counts = np.unique(scores, return_counts=True)
mode_val = unique[np.argmax(counts)]
# Dispersion
variance_val = np.var(scores, ddof=1) # ddof=1 for sample variance
std_dev_val = np.std(scores, ddof=1)
print(f"Mean: {mean_val:.2f}")
print(f"Median: {median_val:.2f}")
print(f"Mode: {mode_val}")
print(f"Variance: {variance_val:.2f}")
print(f"Std Dev: {std_dev_val:.2f}")
Explanation: We define an array of scores. `np.mean()` calculates the average. `np.median()` finds the middle value. For mode, we identify unique values and their counts, selecting the one with the highest frequency. `np.var()` and `np.std()` calculate dispersion; setting `ddof=1` ensures we calculate sample statistics (unbiased estimator) rather than population statistics, which is standard in data science when working with subsets of larger populations.
Common mistakes
- Ignoring Outliers: Using the mean on skewed data with extreme outliers gives a misleading center. Use the median instead.
- Confusing Population vs. Sample Variance: Dividing by $N$ instead of $N-1$ underestimates variability in samples. Always check your library's default behavior.
- Assuming Symmetry: Standard deviation only describes spread accurately if the data is roughly normal. For highly skewed data, IQR is often more robust.
- Misinterpreting Mode: In continuous data, exact modes rarely exist. Binning data first is necessary to find a meaningful mode.
When to use it
Compare descriptive statistics with inferential statistics. Descriptive summarizes what you have; inferential predicts what you don't.| Scenario | Use Descriptive Stats | Use Inferential Stats |
|---|---|---|
| Summarizing past sales data | Yes (Mean/Median) | No |
| Predicting future trends | No | Yes (Regression/Tests) |
| Checking data quality | Yes (Outlier detection) | No |