Tukey’s HSD test
Means (parametric) · Multiple comparisons · reference distribution: q (studentized range)
When to use it
After a significant ANOVA, find out which pairs of groups differ, controlling the error of the whole set of comparisons.
Null hypothesis
For each pair: μᵢ = μⱼ.
Assumptions
- The same as the ANOVA
- Groups of similar size (Tukey-Kramer adjusts for unequal ones)
Test statistic
q = \dfrac{\bar{x}_i - \bar{x}_j}{\sqrt{MS_{\text{within}} / n}}Effect size
Cohen’s d for each pair, with the standard deviation estimated by the ANOVA.
d_{ij} = \dfrac{\bar{x}_i - \bar{x}_j}{\sqrt{MS_{\text{within}}}}How to report it
B − A = 4.2, 95% CI [1.1, 7.3], adjusted p = .006
In R and Python
TukeyHSD(aov(y ~ group, data = df))
from statsmodels.stats.multicomp import pairwise_tukeyhsd
pairwise_tukeyhsd(df["y"], df["group"])
In Python, stats is scipy.stats and np is numpy.
Variants and alternatives
- Games-Howell, with unequal variances
- Dunnett, comparing every group against a control
- Bonferroni, more conservative and general
Where it sits in the catalog
Means (parametric). Parametric tests assume a model for the data, usually the normal distribution, and compare its parameters, such as the mean. When the assumption holds, they are the most powerful.
Multiple comparisons. Post hoc tests: after rejecting that all groups are equal, they point out which pairs differ, controlling the error of the whole set of comparisons.
In the decision tree
It has no path of its own in the tree, but appears alongside other tests:
- As a post hoc test, if H₀ is rejected, after One-way ANOVA.