By the end of this lesson, you will be able to select and execute appropriate hypothesis tests (t-test, chi-square, ANOVA) in Python to determine if observed data differences are statistically significant.
What it is
Hypothesis testing is a statistical method used to make decisions about population parameters based on sample data. It involves two competing hypotheses: the null hypothesis ($H_0$), which assumes no effect or difference, and the alternative hypothesis ($H_1$), which suggests an effect exists. Thep-value quantifies the probability of observing your data (or more extreme) if $H_0$ were true. Common tests include the t-test for comparing means between two groups, ANOVA for comparing means across three or more groups, and chi-square for analyzing relationships between categorical variables.
Why it matters
- Decision Making: Provides a rigorous framework to accept or reject business or scientific assumptions.
- A/B Testing: Essential for determining if a new website feature actually improves conversion rates compared to the old one.
- Quality Control: Helps identify if manufacturing defects vary significantly by shift or machine.
- Research Validation: Ensures findings are not due to random chance before publishing results.
Syntax or steps
The general workflow for any hypothesis test is:- State $H_0$ and $H_1$ clearly.
- Choose the significance level ($\alpha$), typically 0.05.
- Select the correct test based on data type (continuous vs. categorical) and distribution assumptions.
- Calculate the test statistic and p-value using software.
- Compare the p-value to $\alpha$: if $p < \alpha$, reject $H_0$.
Example
This example uses Python'sscipy.stats to perform a t-test, ANOVA, and chi-square test on synthetic data.
import numpy as np
from scipy import stats
# Generate synthetic data
np.random.seed(42)
group_a = np.random.normal(loc=50, scale=10, size=100)
group_b = np.random.normal(loc=55, scale=10, size=100)
group_c = np.random.normal(loc=50, scale=10, size=100)
# 1. Independent T-Test (Comparing Group A vs B)
t_stat, p_val_ttest = stats.ttest_ind(group_a, group_b)
print(f"T-Test P-value: {p_val_ttest:.4f}")
# 2. One-way ANOVA (Comparing Group A, B, and C)
f_stat, p_val_anova = stats.f_oneway(group_a, group_b, group_c)
print(f"ANOVA P-value: {p_val_anova:.4f}")
# 3. Chi-Square Test (Categorical Data)
observed = np.array([[10, 20], [30, 40]]) # Example contingency table
chi2, p_val_chi, dof, expected = stats.chi2_contingency(observed)
print(f"Chi-Square P-value: {p_val_chi:.4f}")
Explanation:
stats.ttest_indcalculates the t-statistic and p-value for two independent samples. If the p-value is low, the means of Group A and B are likely different.stats.f_onewayperforms ANOVA. It checks if at least one group mean differs from the others among multiple groups.stats.chi2_contingencytests independence between rows and columns in a contingency table. It compares observed frequencies against expected frequencies under the null hypothesis.
Common mistakes
- P-hacking: Running multiple tests until a significant result appears. Fix by correcting alpha levels (e.g., Bonferroni correction).
- Ignoring Assumptions: Using a t-test on non-normal data with small sample sizes. Fix by checking normality or using non-parametric tests like Mann-Whitney U.
- Confusing Statistical Significance with Practical Importance: A tiny p-value doesn't mean the effect size is large. Always report effect sizes (e.g., Cohen's d).
- Wrong Test Selection: Using ANOVA for two groups (use t-test) or t-test for categorical data (use chi-square).
When to use it
| Scenario | Data Type | Recommended Test |
|---|---|---|
| Compare means of 2 groups | Continuous | T-test |
| Compare means of 3+ groups | Continuous | ANOVA |
| Check association between categories | Categorical | Chi-Square |
| Compare distributions without normality | Ordinal/Continuous | Mann-Whitney / Kruskal-Wallis |
Practice
Guided Exercise: Modify the code above to comparegroup_a and group_c using a t-test. Note that both have a mean of 50. What do you expect the p-value to be?
Challenge: Create a new dataset where
group_d has a mean of 60. Perform an ANOVA on group_a, group_b, and group_d. Interpret the result if the p-value is less than 0.05.
Hint: In the guided exercise, since the means are identical, the p-value should be high (close to 1), indicating no significant difference.
Quick check
Question: If your p-value is 0.03 and your significance level ($\alpha$) is 0.05, what is your conclusion regarding the null hypothesis?Answer: Since $0.03 < 0.05$, you reject the null hypothesis. There is sufficient evidence to support the alternative hypothesis.