states that \(H_0\) is rejected for small \(\lambda\), states Wilks' theorem with \(r\) the number of restrictions, and asserts that the normal-mean likelihood ratio test “reduces exactly to the one-sample \(t\) test”. None of that is repeated.
This unit does the three things that treatment does not. It carries out the reduction rather than asserting it; it measures how far the \(\chi^{2}\) approximation is from the truth at a realistic sample size; and it puts the likelihood ratio beside its two asymptotic equals, the Wald and Rao score statistics, on data where all three disagree.
Given. \(X_1, \dots, X_n\) independent \(N(\mu, \sigma^{2})\) with both parameters unknown. Test \(H_0: \mu = \mu_0\) against \(H_1: \mu \ne \mu_0\).
Step 1 — the unrestricted maximum. The maximum likelihood estimates are \(\hat\mu = \bar x\) and \(\hat\sigma^{2} = \frac{1}{n}\sum(x_i - \bar x)^{2}\), and substituting them,
\[ \sup_{\Theta} L = \left(2\pi\hat\sigma^{2}\right)^{-n/2}e^{-n/2}. \]The exponent collapses to \(-n/2\) because \(\sum(x_i-\bar x)^{2} = n\hat\sigma^{2}\) at the maximum — that cancellation is what makes the whole calculation short.
Step 2 — the restricted maximum. With \(\mu\) held at \(\mu_0\), the variance estimate becomes \(\hat\sigma_0^{2} = \frac{1}{n}\sum(x_i - \mu_0)^{2}\), and by the same cancellation
\[ \sup_{\Theta_0} L = \left(2\pi\hat\sigma_0^{2}\right)^{-n/2}e^{-n/2}. \]Step 3 — the ratio. Everything except the two variance estimates cancels:
\[ \lambda = \left(\frac{\hat\sigma_0^{2}}{\hat\sigma^{2}}\right)^{-n/2} = \left(\frac{\sum(x_i-\mu_0)^{2}}{\sum(x_i-\bar x)^{2}}\right)^{-n/2}. \]Step 4 — the identity that finishes it. Expand about \(\bar x\):
\[ \sum(x_i-\mu_0)^{2} = \sum(x_i-\bar x)^{2} + n(\bar x - \mu_0)^{2}, \]the cross term vanishing because \(\sum(x_i - \bar x) = 0\). Writing \(s^{2} = \frac{1}{n-1}\sum(x_i-\bar x)^{2}\) and \(t = \frac{\bar x - \mu_0}{s/\sqrt n}\),
\[ \frac{\sum(x_i-\mu_0)^{2}}{\sum(x_i-\bar x)^{2}} = 1 + \frac{n(\bar x-\mu_0)^{2}}{(n-1)s^{2}} = 1 + \frac{t^{2}}{n-1}, \] \[ \boxed{\;\lambda = \left(1 + \frac{t^{2}}{n-1}\right)^{-n/2}, \qquad -2\ln\lambda = n\ln\!\left(1 + \frac{t^{2}}{n-1}\right).\;} \]Step 5 — read it. \(\lambda\) is a strictly decreasing function of \(t^{2}\). So “reject for small \(\lambda\)” and “reject for large \(|t|\)” are the same rule, and the likelihood ratio test is the two-sided \(t\) test — not approximately, identically, at every sample size.
The same three steps give the other familiar tests. Restricting \(\sigma^{2} = \sigma_0^{2}\) instead produces a \(\lambda\) monotone in \((n-1)s^{2}/\sigma_0^{2}\), which is the \(\chi^{2}\) test; restricting two normal means to be equal produces the two-sample \(t\). The familiar tests are not a collection of separate inventions — they are one construction applied three times.
Step 5 above gives an unusual opportunity: a case where the exact null distribution of \(-2\ln\lambda\) is known, so Wilks' approximation can be checked rather than trusted.
Given. \(n = 10\) and an observed \(t = 2.5\).
Step 1 — the statistic.
\[ -2\ln\lambda = 10\ln\!\left(1 + \frac{2.5^{2}}{9}\right) = 10\ln\!\left(1 + \frac{25}{36}\right) = 10\ln\frac{61}{36} = 5.273549. \]Keep the bracket as the fraction \(61/36\): rounding it to \(1.694444\) first and taking the logarithm of that gives \(5.273547\), wrong in the last digit.
Step 2 — the two \(p\) values. \(H_0\) imposes one restriction, so Wilks gives \(\chi^{2}_1\):
\[ P\!\left(\chi^{2}_{1} > 5.273549\right) = 0.021652, \]against the exact answer from the \(t\) distribution on \(9\) degrees of freedom,
\[ P\!\left(|t_{9}| > 2.5\right) = 0.033862. \]Step 3 — the size of the error. The asymptotic \(p\) value is too small by \(0.012210\); the exact value is \(1.5639\) times it.
Interpretation, and it matters. At the conventional \(5\%\) threshold the two answers fall on opposite sides of nothing — both reject — but the asymptotic calculation reports a result about half again as significant as the truth. Wilks' theorem errs in the anti-conservative direction: it understates \(p\), overstates significance, and does so by a factor that only disappears as \(n \to \infty\). At \(n = 10\) the exact test exists and should be used; the approximation earns its place only where no exact distribution is available, which is most of the time.
Let \(\ell(\theta)\) be the log-likelihood, \(U(\theta) = \ell'(\theta)\) the score, and \(I(\theta)\) the Fisher information (built in Estimation Theory, Unit 1, and not rebuilt here). To test \(H_0: \theta = \theta_0\):
| Test | Statistic | Evaluated at | Needs |
|---|---|---|---|
| Likelihood ratio | \(W_{LR} = 2\left[\ell(\hat\theta) - \ell(\theta_0)\right]\) | both | the maximum, and the null value |
| Wald | \(W_{W} = \left(\hat\theta - \theta_0\right)^{2} I(\hat\theta)\) | the estimate | fitting only the full model |
| Rao score | \(W_{S} = U(\theta_0)^{2} / I(\theta_0)\) | the null | no fitting at all |
All three converge in distribution to \(\chi^{2}_{r}\) under \(H_0\), and the reason is a two-term Taylor expansion: \(\ell\) is locally quadratic near its maximum, and for an exactly quadratic log-likelihood the three are algebraically identical. They differ only by the curvature that a real log-likelihood has and a parabola does not.
Which to use is a practical question, not a theoretical one. The score test needs no estimate of \(\theta\) at all, which is why it is the one used when fitting is expensive or when the maximum lies on a boundary. Wald needs only the fitted model, which is why every regression table prints it. The likelihood ratio needs both fits and is generally the best behaved of the three — and it is invariant under reparameterisation. So is the score test when it uses the expected information; Wald is the one that is not.
That last point is the decisive one. Testing \(\theta = 1\) and testing \(\log\theta = 0\) are the same hypothesis. The likelihood ratio (and the score test) gives the same answer for both; Wald does not, because \(\hat\theta - \theta_0\) and \(\log\hat\theta - \log\theta_0\) are different distances. A test whose conclusion depends on how the parameter was written down is a test with a defect.
Given. \(X \sim \text{Bin}(20, p)\), observed \(x = 6\), testing \(H_0: p = 0.5\) against \(p \ne 0.5\). The estimate is \(\hat p = 6/20 = 0.300000\).
Step 1 — likelihood ratio.
\[ W_{LR} = 2\left[x\ln\frac{\hat p}{p_0} + (n-x)\ln\frac{1-\hat p}{1-p_0}\right] = 2\left[6\ln\frac{0.3}{0.5} + 14\ln\frac{0.7}{0.5}\right] = 3.291315. \]Step 2 — Wald. The information at \(\hat p\) is \(n/[\hat p(1-\hat p)]\), so
\[ W_{W} = \frac{\left(\hat p - p_0\right)^{2}}{\hat p(1-\hat p)/n} = \frac{(0.2)^{2}}{0.21/20} = 3.809524. \]Step 3 — score. The same quantity with the information evaluated under the null:
\[ W_{S} = \frac{\left(\hat p - p_0\right)^{2}}{p_0(1-p_0)/n} = \frac{(0.2)^{2}}{0.25/20} = 3.200000. \]Step 4 — the four \(p\) values.
| Test | Statistic | \(p\) value from \(\chi^{2}_{1}\) | Verdict at \(5\%\) |
|---|---|---|---|
| Rao score | 3.200000 | 0.073638 | do not reject |
| Likelihood ratio | 3.291315 | 0.069647 | do not reject |
| Wald | 3.809524 | 0.050962 | do not reject, barely |
| Exact binomial | — | 0.115318 | do not reject |
Interpretation. The three asymptotic answers span \(0.051\) to \(0.074\) — a factor of \(1.44\) between them — on the same twenty observations, and all three are badly wrong: the exact \(p\) value is \(0.115318\), about \(1.6\) times the largest of them and more than twice the Wald value.
Two things follow, and they are the reason this example is here rather than a tidier one:
With \(\theta \in \mathbb{R}^{p}\) and \(H_0\) imposing \(r\) independent restrictions, all three statistics converge to \(\chi^{2}_{r}\). The count \(r\) is the drop in dimension, not the number of parameters: testing that three regression coefficients are all zero gives \(r = 3\) whatever else is in the model.
The regularity conditions are not decoration. Wilks' theorem fails, and fails badly, when:
Each of these is a case where the software prints a \(p\) value and the \(p\) value means nothing. The conditions are worth memorising in the negative — as the three situations in which not to believe the output.