Back to Data Science Notes
Topic #54

Descriptive Statistics

By the end of this lesson, you will be able to calculate and interpret key descriptive statistics—mean, median, mode, variance, and standard deviation—to summarize data distributions effectively.

What it is

Descriptive statistics are numerical measures that summarize the main features of a collection of quantitative data. They provide a snapshot of central tendency (where the data clusters) and dispersion (how spread out the data is). The mean is the arithmetic average, the median is the middle value when sorted, and the mode is the most frequent value. Variance measures the average squared deviation from the mean, while standard deviation is the square root of variance, expressed in the same units as the original data. Related terms include range, interquartile range (IQR), and skewness.

Why it matters

  • Data Cleaning: Identifying outliers or errors by checking if values fall outside expected ranges defined by mean and standard deviation.
  • Feature Engineering: Normalizing data for machine learning models often requires knowing the mean and standard deviation of training features.
  • Reporting: Providing concise summaries of large datasets for stakeholders who cannot review raw rows.
  • Distribution Analysis: Comparing mean vs. median helps detect skewness; if mean > median, the distribution is likely right-skewed.

Syntax or steps

To compute these metrics manually: 1. Sort the data for median and mode identification. 2. Calculate the sum of all values divided by count $n$ for the mean ($\mu$). 3. Find the difference between each value and the mean, square them, sum them, and divide by $n-1$ (sample variance) or $n$ (population variance). 4. Take the square root of the variance for standard deviation. In Python, using the `statistics` module or `numpy`, these operations are abstracted into single function calls.

Example

import numpy as np

# Sample dataset: exam scores
scores = np.array([85, 90, 78, 92, 88, 76, 90, 85])

# Central Tendency
mean_val = np.mean(scores)
median_val = np.median(scores)
# Note: NumPy doesn't have a direct mode function, so we use scipy or manual counting
unique, counts = np.unique(scores, return_counts=True)
mode_val = unique[np.argmax(counts)]

# Dispersion
variance_val = np.var(scores, ddof=1) # ddof=1 for sample variance
std_dev_val = np.std(scores, ddof=1)

print(f"Mean: {mean_val:.2f}")
print(f"Median: {median_val:.2f}")
print(f"Mode: {mode_val}")
print(f"Variance: {variance_val:.2f}")
print(f"Std Dev: {std_dev_val:.2f}")
Explanation: We define an array of scores. `np.mean()` calculates the average. `np.median()` finds the middle value. For mode, we identify unique values and their counts, selecting the one with the highest frequency. `np.var()` and `np.std()` calculate dispersion; setting `ddof=1` ensures we calculate sample statistics (unbiased estimator) rather than population statistics, which is standard in data science when working with subsets of larger populations.

Common mistakes

  • Ignoring Outliers: Using the mean on skewed data with extreme outliers gives a misleading center. Use the median instead.
  • Confusing Population vs. Sample Variance: Dividing by $N$ instead of $N-1$ underestimates variability in samples. Always check your library's default behavior.
  • Assuming Symmetry: Standard deviation only describes spread accurately if the data is roughly normal. For highly skewed data, IQR is often more robust.
  • Misinterpreting Mode: In continuous data, exact modes rarely exist. Binning data first is necessary to find a meaningful mode.

When to use it

Compare descriptive statistics with inferential statistics. Descriptive summarizes what you have; inferential predicts what you don't.
ScenarioUse Descriptive StatsUse Inferential Stats
Summarizing past sales dataYes (Mean/Median)No
Predicting future trendsNoYes (Regression/Tests)
Checking data qualityYes (Outlier detection)No

Practice

Guided Exercise: Given the list `[10, 12, 12, 15, 20]`, calculate the mean and median manually. Verify your results using Python. Challenge: Add an outlier `100` to the list above. Recalculate mean and median. Observe how the mean changes drastically while the median remains stable. This demonstrates robustness.

Quick check

Question: If a dataset has a mean significantly higher than its median, what does this suggest about the distribution? Answer: It suggests the distribution is right-skewed (positively skewed), likely due to high-value outliers pulling the mean up.

Summary

Descriptive statistics provide essential tools for summarizing data through central tendency and dispersion measures. Choosing between mean/median or variance/IQR depends on data distribution and the presence of outliers. Mastering these basics allows for effective data cleaning, reporting, and preparation for advanced modeling.

Want to go beyond the notes?

Join Coding Now Tech Institute's Data Science course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Descriptive Statistics – FAQs

Quick answers about learning Descriptive Statistics in Data Science.

This free note from Coding Now Tech Institute explains Descriptive Statistics in Data Science — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Data Science topic on Coding Now Tech Institute, including Descriptive Statistics, is 100% free with no signup required.
With focused practice, most students grasp Descriptive Statistics in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now