Independent samples t-test
Means (parametric) · Two independent samples · reference distribution: t(df)
When to use it
Compare the means of two different groups, such as treatment and control.
Null hypothesis
The two population means are equal: μ₁ = μ₂.
Assumptions
- Groups independent of each other
- Normality in each group, or large samples
Test statistic
t = \dfrac{\bar{x}_1 - \bar{x}_2}{\sqrt{s_1^2/n_1 + s_2^2/n_2}}Effect size
Cohen’s d with the pooled standard deviation; Hedges’ g corrects its bias in small samples.
d = \dfrac{\bar{x}_1 - \bar{x}_2}{s_p}, \qquad s_p = \sqrt{\dfrac{(n_1 - 1)s_1^2 + (n_2 - 1)s_2^2}{n_1 + n_2 - 2}}How to report it
t(37.8) = 2.15, p = .038, d = 0.68 (Welch)
In R and Python
R
t.test(y ~ group, data = df) # Welch is the default
Python
stats.ttest_ind(a, b, equal_var=False) # Welch
In Python, stats is scipy.stats and np is numpy.
Variants and alternatives
- Welch (above): does not assume equal variances and is the recommended default
- Classic Student, with pooled variance, if the variances are equal
- Mann-Whitney U, without assuming normality
- TOST (two one-sided t-tests), to show equivalence within a margin instead of a difference
Where it sits in the catalog
Means (parametric). Parametric tests assume a model for the data, usually the normal distribution, and compare its parameters, such as the mean. When the assumption holds, they are the most powerful.
Two independent samples. Two groups of different individuals, unrelated to each other, such as treatment and control.
In the decision tree
- What do you want to do? Compare groups, or one group with a reference value
- What kind of response did you measure? Numeric: a measurement, such as weight, time or score
- How many groups or measurements? Two groups
- Are the groups made of different individuals or the same individuals? Different individuals (independent)
- Are the data approximately normal in each group, or are the samples large? Yes