Back to Data Science Notes
Topic #58

Sampling & the Central Limit Theorem

Understand how the Central Limit Theorem (CLT) allows us to make reliable inferences about a population using sample means, even when the underlying data is not normally distributed.

What it is

The Central Limit Theorem states that if you take sufficiently large random samples from any population with a finite mean and variance, the distribution of the sample means will approximate a normal distribution. This holds true regardless of the shape of the original population distribution (e.g., skewed, uniform, or bimodal).

Key terms include:

  • Population Mean ($\mu$): The average of all values in the entire dataset.
  • Sample Mean ($\bar{x}$): The average of a subset of data drawn from the population.
  • Standard Error: The standard deviation of the sampling distribution of the sample mean.

Why it matters

  • Hypothesis Testing: It justifies the use of z-tests and t-tests for comparing group averages.
  • Confidence Intervals: It allows us to calculate error bars and confidence intervals around estimates.
  • Model Robustness: Many machine learning algorithms assume residuals are normally distributed; CLT helps validate this assumption at scale.
  • Data Efficiency: We do not need to analyze every single data point to understand the central tendency of a massive dataset.

Syntax or steps

To observe the CLT in action, follow these conceptual steps:

  1. Define a non-normal population distribution (e.g., exponential or uniform).
  2. Select a sample size $n$ (typically $n \geq 30$ is sufficient for approximation).
  3. Draw multiple random samples of size $n$ from the population.
  4. Calculate the mean of each sample.
  5. Plot the histogram of these sample means.

Example

This Python example uses NumPy to generate an exponentially distributed population (highly skewed) and demonstrates that the distribution of sample means becomes normal.

import numpy as np
import matplotlib.pyplot as plt

# 1. Create a skewed population (Exponential distribution)
np.random.seed(42)
population = np.random.exponential(scale=2.0, size=100000)

# 2. Define parameters
num_samples = 1000
sample_size = 50

# 3. Generate sample means
sample_means = []
for _ in range(num_samples):
    sample = np.random.choice(population, size=sample_size, replace=False)
    sample_means.append(np.mean(sample))

# 4. Visualize
plt.figure(figsize=(10, 5))

# Plot Population Distribution
plt.subplot(1, 2, 1)
plt.hist(population, bins=50, color='orange', alpha=0.7)
plt.title('Original Population (Skewed)')
plt.xlabel('Value')

# Plot Sampling Distribution of Means
plt.subplot(1, 2, 2)
plt.hist(sample_means, bins=30, color='blue', alpha=0.7)
plt.title(f'Distribution of Sample Means (n={sample_size})')
plt.xlabel('Mean Value')

plt.tight_layout()
plt.show()

Explanation:

  • np.random.exponential creates a right-skewed dataset, which is clearly not normal.
  • We loop 1,000 times, drawing a sample of 50 items each time and calculating its mean.
  • The second plot shows that while the original data was skewed, the means form a bell-shaped curve centered near the population mean.

Common mistakes

  • Assuming small samples work: If $n$ is too small (e.g., $n=5$) and the population is heavily skewed, the sample means may still look skewed. Always check $n \geq 30$.
  • Confusing population SD with Standard Error: The spread of the sample means is smaller than the spread of the raw data. Use $\sigma / \sqrt{n}$ for the standard error.
  • Ignoring independence: Samples must be independent. If you sample with replacement from a tiny population, or if data points are correlated (time-series), the CLT assumptions may fail.

When to use it

Compare CLT-based inference with bootstrapping.

Method Best Used When Limitation
Central Limit Theorem You have a moderate-to-large sample size ($n > 30$) and want fast analytical calculations. Relies on the assumption that the sampling distribution is approximately normal.
Bootstrapping Sample sizes are small, or the data distribution is complex/unknown. Computationally expensive; requires resampling thousands of times.

Practice

Guided Exercise: Modify the code above to change sample_size to 5. Observe the shape of the "Distribution of Sample Means." Does it look normal? Now change it to 100. What happens?

Challenge: Calculate the theoretical standard error using the formula $\text{SE} = \frac{\sigma}{\sqrt{n}}$. Compare this value to the actual standard deviation of your sample_means list. They should be very close.

Quick check

Question: If I draw samples of size 10 from a uniform distribution, will the distribution of the sample means be perfectly normal?

Answer: No, it will be approximately normal but likely still show some flatness or irregularity because $n=10$ is often too small for heavy convergence, especially compared to $n=30+$. However, it will be much closer to normal than the original uniform distribution.

Summary

The Central Limit Theorem bridges the gap between messy real-world data and clean statistical models. By ensuring we use adequate sample sizes, we can rely on normal distribution properties to estimate population parameters accurately, making it a cornerstone of inferential statistics.

Want to go beyond the notes?

Join Coding Now Tech Institute's Data Science course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Sampling & the Central Limit Theorem – FAQs

Quick answers about learning Sampling & the Central Limit Theorem in Data Science.

This free note from Coding Now Tech Institute explains Sampling & the Central Limit Theorem in Data Science — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Data Science topic on Coding Now Tech Institute, including Sampling & the Central Limit Theorem, is 100% free with no signup required.
With focused practice, most students grasp Sampling & the Central Limit Theorem in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now