One-way ANOVA
Means (parametric) · Three or more independent groups · reference distribution: F(k−1, N−k)
When to use it
Compare the means of three or more independent groups at once, without inflating the type I error with several t-tests.
Null hypothesis
All means are equal: μ₁ = μ₂ = … = μₖ.
Assumptions
- Independent groups
- Normal residuals
- Equal variances (check with Levene)
Test statistic
F = \dfrac{MS_{\text{between}}}{MS_{\text{within}}}Effect size
η²: the share of the total variation explained by the groups. ω² is less biased. Benchmarks: 0.01, 0.06 and 0.14.
\eta^2 = \dfrac{SS_{\text{between}}}{SS_{\text{total}}}How to report it
F(2, 42) = 4.87, p = .013, η² = .19
In R and Python
summary(aov(y ~ group, data = df))
stats.f_oneway(a, b, c)
In Python, stats is scipy.stats and np is numpy.
Variants and alternatives
- Welch’s ANOVA, when the variances differ
- Kruskal-Wallis, without assuming normality
- After rejecting H₀: Tukey HSD
- With two factors and their interaction: Two-way ANOVA
Where it sits in the catalog
Means (parametric). Parametric tests assume a model for the data, usually the normal distribution, and compare its parameters, such as the mean. When the assumption holds, they are the most powerful.
Three or more independent groups. Several groups of different individuals, tested at once: one test per pair would inflate the type I error.
In the decision tree
- What do you want to do? Compare groups, or one group with a reference value
- What kind of response did you measure? Numeric: a measurement, such as weight, time or score
- How many groups or measurements? Three or more groups
- Are the groups made of different individuals or the same individuals? Different individuals (independent)
- Are the residuals approximately normal? Yes