Português

Independent samples t-test

Means (parametric) · Two independent samples · reference distribution: t(df)

When to use it

Compare the means of two different groups, such as treatment and control.

Null hypothesis

The two population means are equal: μ₁ = μ₂.

Assumptions

Test statistic

t = \dfrac{\bar{x}_1 - \bar{x}_2}{\sqrt{s_1^2/n_1 + s_2^2/n_2}}

Effect size

Cohen’s d with the pooled standard deviation; Hedges’ g corrects its bias in small samples.

d = \dfrac{\bar{x}_1 - \bar{x}_2}{s_p}, \qquad s_p = \sqrt{\dfrac{(n_1 - 1)s_1^2 + (n_2 - 1)s_2^2}{n_1 + n_2 - 2}}

How to report it

t(37.8) = 2.15, p = .038, d = 0.68 (Welch)

In R and Python

R
t.test(y ~ group, data = df)  # Welch is the default
Python
stats.ttest_ind(a, b, equal_var=False)  # Welch

In Python, stats is scipy.stats and np is numpy.

Variants and alternatives

Where it sits in the catalog

Means (parametric). Parametric tests assume a model for the data, usually the normal distribution, and compare its parameters, such as the mean. When the assumption holds, they are the most powerful.

Two independent samples. Two groups of different individuals, unrelated to each other, such as treatment and control.

In the decision tree

  1. What do you want to do? Compare groups, or one group with a reference value
  2. What kind of response did you measure? Numeric: a measurement, such as weight, time or score
  3. How many groups or measurements? Two groups
  4. Are the groups made of different individuals or the same individuals? Different individuals (independent)
  5. Are the data approximately normal in each group, or are the samples large? Yes

Open the decision tree