Learn how to visualize the distribution of numerical data using histograms in Python with Matplotlib.
What it is
A histogram is a graphical representation of the distribution of numerical data. It divides the range of values into intervals, called bins, and counts how many observations fall into each bin. The height of each bar represents the frequency (count) or density of data points within that interval. Unlike a bar chart, which compares categorical data, a histogram shows the shape, center, and spread of continuous data. Key related terms include frequency, bin width, and distribution.
Why it matters
- Identify patterns: Quickly spot skewness, outliers, or multiple peaks (modes) in your data.
- Check assumptions: Verify if data follows a normal distribution before applying statistical tests.
- Compare groups: Overlay histograms to compare distributions between different categories.
- Data cleaning: Detect unexpected gaps or extreme values that may indicate errors.
Syntax or steps
The primary function for creating histograms in Matplotlib is plt.hist(). The basic syntax requires an array-like object containing the data. You can control the granularity of the visualization by adjusting the number of bins.
import matplotlib.pyplot as plt
# Basic usage
plt.hist(data_array, bins=number_of_bins)
plt.show()
Example
This example generates random normally distributed data and plots its histogram.
import numpy as np
import matplotlib.pyplot as plt
# Generate 1000 random numbers from a normal distribution
data = np.random.normal(loc=0, scale=1, size=1000)
# Create the histogram
plt.figure(figsize=(8, 6))
plt.hist(data, bins=30, color='skyblue', edgecolor='black')
# Add labels and title
plt.title('Distribution of Random Data')
plt.xlabel('Value')
plt.ylabel('Frequency')
# Display the plot
plt.grid(axis='y', alpha=0.75)
plt.show()
Explanation: We use np.random.normal to create synthetic data centered at 0 with a standard deviation of 1. The plt.hist function takes this data and splits it into 30 equal-width bins. The edgecolor parameter adds black borders to bars for clarity, while grid helps read frequencies against the y-axis.
Common mistakes
- Too few or too many bins: Using too few bins hides details; too many creates noise. Start with 10-20 bins and adjust based on data size.
- Ignoring normalization: If comparing datasets of different sizes, set
density=Trueto show probability density rather than raw counts. - Missing axis labels: Always label axes to ensure the plot is interpretable without context.
- Using categorical data: Histograms are for continuous numerical data. Use bar charts for discrete categories.
When to use it
Histograms are best for understanding the underlying distribution of a single variable. Compare them with other visualization tools below.
| Visualization | Best For | Limitation |
|---|---|---|
| Histogram | Detailed shape of distribution | Bin choice affects appearance |
| Box Plot | Comparing medians and outliers | Hides distribution shape |
| KDE Plot | Smooth density estimation | Can over-smooth real features |
Practice
Guided Exercise: Modify the example above to change bins=30 to bins=5 and then bins=100. Observe how the shape changes.
Challenge: Generate two sets of random data: one normal and one uniform. Plot both histograms on the same figure using alpha=0.5 for transparency. Hint: Call plt.hist() twice before plt.show().
Quick check
Question: What does the density=True parameter do in plt.hist()?
Answer: It normalizes the histogram so that the total area under the curve equals 1, representing a probability density function instead of raw counts.
Summary
Histograms are essential for exploring the distribution of continuous data. By carefully selecting bin sizes and labeling axes, you can reveal critical insights about data shape, spread, and anomalies that summary statistics alone might miss.