By the end of this lesson, you will be able to calculate the mean, median, and mode of a dataset using Python's built-in statistics module.
What it is
Measures of central tendency are single values that attempt to describe a set of data by identifying the central position within that set. The three most common measures are:
- Mean: The arithmetic average (sum of all values divided by the count).
- Median: The middle value when the data is sorted in ascending order.
- Mode: The value that appears most frequently in the dataset.
Python provides these calculations natively through the statistics module, which handles edge cases like empty datasets or non-numeric types more robustly than manual calculation.
Why it matters
- Data Analysis: Quickly summarize large datasets without plotting graphs.
- Outlier Detection: Comparing the mean and median helps identify skewed distributions caused by outliers.
- Reporting: Provides standard metrics for business intelligence and scientific reports.
- Simplicity: Avoids writing complex sorting or counting logic manually.
Syntax or steps
To use these functions, you must first import the module. Each function accepts an iterable of numbers (list, tuple, etc.).
- Import the module:
import statistics - Prepare your data as a list of numbers.
- Call the specific function:
statistics.mean(data),statistics.median(data), orstatistics.mode(data).
Example
import statistics
# Sample dataset
data = [10, 20, 20, 30, 40, 50]
# Calculate Mean
mean_val = statistics.mean(data)
# Calculate Median
median_val = statistics.median(data)
# Calculate Mode
mode_val = statistics.mode(data)
print(f"Data: {data}")
print(f"Mean: {mean_val}")
print(f"Median: {median_val}")
print(f"Mode: {mode_val}")
Explanation:
import statistics: Loads the standard library module.statistics.mean(data): Sums $10+20+20+30+40+50=170$ and divides by 6, resulting in approximately $28.33$.statistics.median(data): Sorts the data (already sorted here). Since there are 6 items (even number), it averages the two middle values ($20$ and $30$), resulting in $25.0$.statistics.mode(data): Identifies $20$ as the most frequent value.
Common mistakes
- Empty Lists: Passing an empty list raises a
StatisticsError. Always check if the list has length > 0 before calculating. - Mixed Types: The
statisticsmodule expects numeric types. Strings or None values will cause errors. Ensure data cleaning occurs first. - Multiple Modes: In older Python versions,
statistics.mode()raised an error if multiple modes existed. In Python 3.8+, it returns the first encountered mode. Usestatistics.multimode()if you need all modes. - Integer Division Confusion: Remember that
mean()always returns a float, even if the result is a whole number.
When to use it
While NumPy offers faster performance for massive arrays, the statistics module is ideal for small-to-medium lists where readability and zero dependencies are preferred.
| Feature | statistics module | NumPy |
|---|---|---|
| Setup | Built-in (no install) | Requires installation |
| Performance | Good for small lists | Optimized for large arrays |
| Complexity | Simple syntax | More verbose setup |
Practice
Guided Exercise: Create a list of temperatures: [72, 75, 72, 78, 75, 72]. Calculate the mean, median, and mode. Verify that the mode is 72.
Challenge: Write a script that takes a user-input string of comma-separated numbers, converts them to floats, and prints the median. Handle the case where the input is empty.
Hint: Use input().split(',') and float() inside a list comprehension.
Quick check
Question: What does statistics.median([1, 2, 3, 4]) return?
Answer: It returns 2.5. Because there is an even number of elements, it averages the two middle values (2 and 3).
Summary
The statistics module provides reliable, readable methods for calculating mean, median, and mode. Understanding the difference between these measures allows you to interpret data distribution accurately, particularly when dealing with outliers or skewed datasets.