Two-proportion Z-test
Proportions (counts) · Two independent samples · reference distribution: N(0, 1)
When to use it
Compare the success rate of two independent groups, such as the conversion of two versions of a web page.
Null hypothesis
The two proportions are equal: p₁ = p₂.
Assumptions
- Independent groups
- At least 5 to 10 successes and failures in each group
Test statistic
z = \dfrac{\hat{p}_1 - \hat{p}_2}{\sqrt{\hat{p}(1 - \hat{p})\left(\frac{1}{n_1} + \frac{1}{n_2}\right)}}Effect size
Cohen’s h; the difference in proportions, the relative risk and the odds ratio are also reported.
h = 2\arcsin\sqrt{\hat{p}_1} - 2\arcsin\sqrt{\hat{p}_2}How to report it
45% vs. 30%, z = 2.19, p = .029, h = 0.31
In R and Python
R
prop.test(c(45, 30), c(100, 100), correct = FALSE)
Python
from statsmodels.stats.proportion import proportions_ztest
proportions_ztest([45, 30], [100, 100])
In Python, stats is scipy.stats and np is numpy.
Variants and alternatives
- Chi-square of independence, equivalent in a 2 × 2 table
- Fisher’s exact test, with small counts
Where it sits in the catalog
Proportions (counts). For categorical answers, such as yes or no, hit or miss: the data are how often each category appears, and the test compares proportions.
Two independent samples. Two groups of different individuals, unrelated to each other, such as treatment and control.
In the decision tree
- What do you want to do? Compare groups, or one group with a reference value
- What kind of response did you measure? Categorical: yes or no, or categories
- How many groups or measurements? Two independent groups
- Are all expected counts at least 5? Yes