Hosmer-Lemeshow test
Models and regression · reference distribution: χ²(g−2)
When to use it
Check whether the probabilities predicted by a logistic regression match the observed frequencies, grouping the cases by risk band.
Null hypothesis
The model fits well: observed and predicted frequencies agree.
Assumptions
- Large sample
- The result depends on the number of groups g, usually 10
Test statistic
HL = \sum_{j=1}^{g} \dfrac{(O_j - E_j)^2}{E_j\,(1 - E_j / n_j)}How to report it
χ²(8) = 6.3, p = .61: no evidence of poor fit
In R and Python
R
library(ResourceSelection)
hoslem.test(model$y, fitted(model), g = 10)
Python
# not in statsmodels: group by deciles of p_hat
g = pd.qcut(p_hat, 10, labels=False)
t = pd.DataFrame({"y": y, "p": p_hat, "g": g}).groupby("g").agg(o=("y", "sum"), e=("p", "sum"), n=("y", "size"))
HL = ((t.o - t.e)**2 / (t.e * (1 - t.e / t.n))).sum()
p = stats.chi2.sf(HL, 10 - 2)
In Python, stats is scipy.stats and np is numpy.
Variants and alternatives
- Likelihood ratio, to compare nested models
Where it sits in the catalog
Models and regression. Tests run inside a fitted model: whether a coefficient matters, whether the model explains anything and whether the residuals meet the assumptions.
In the decision tree
- What do you want to do? Test a regression model
- What do you want to test? The fit of a logistic regression