Cohen’s kappa
Association and correlation · reference distribution: N(0, 1) approx.
When to use it
Measure the agreement between two raters who classify the same items into categories, discounting the agreement expected by chance.
Null hypothesis
The agreement is only that of chance: κ = 0.
Assumptions
- The same items classified by both raters
- Fixed, mutually exclusive categories
- Ratings independent of each other
Test statistic
\kappa = \dfrac{p_o - p_e}{1 - p_e}Effect size
κ itself. By the Landis and Koch benchmark, above 0.6 is substantial agreement and above 0.8, almost perfect.
How to report it
κ = .68, 95% CI [.52, .84]: substantial agreement
In R and Python
R
library(irr)
kappa2(df[, c("rater1", "rater2")])
Python
from sklearn.metrics import cohen_kappa_score
cohen_kappa_score(rater1, rater2)
In Python, stats is scipy.stats and np is numpy.
Variants and alternatives
- Weighted kappa, for ordinal categories
- Fleiss’ kappa, for more than two raters
Where it sits in the catalog
Association and correlation. Instead of comparing groups, they measure whether two variables move together, and how strongly.
In the decision tree
- What do you want to do? Measure the relationship between two variables
- What kind of variables are they? Two raters classifying the same items