Durbin-Watson test
Models and regression · reference distribution: d (tabulated)
When to use it
Check for first-order autocorrelation in the residuals of a regression, common with data over time.
Null hypothesis
The residuals have no autocorrelation.
Assumptions
- Regression with an intercept
- No lagged response among the explanatory variables
Test statistic
d = \dfrac{\sum_{t=2}^{n} (e_t - e_{t-1})^2}{\sum_{t=1}^{n} e_t^2}How to report it
d = 1.32, p = .004
In R and Python
R
library(lmtest)
dwtest(model)
Python
from statsmodels.stats.stattools import durbin_watson
durbin_watson(model.resid) # the statistic only, no p-value
In Python, stats is scipy.stats and np is numpy.
Variants and alternatives
- Breusch-Godfrey, for higher lags
- Ljung-Box, on the residuals of time series
Where it sits in the catalog
Models and regression. Tests run inside a fitted model: whether a coefficient matters, whether the model explains anything and whether the residuals meet the assumptions.
In the decision tree
- What do you want to do? Test a regression model
- What do you want to test? The residuals
- What problem do you suspect? Autocorrelation, in data over time