Skip to the content

Topics Covered

Completeness Lehmann–Scheffé MLE Properties CAN & BAN Jackknife Bootstrap
On this page
  1. 1. Completeness
  2. 2. The Lehmann–Scheffé Theorem
  3. 3. Maximum Likelihood and Its Large-Sample Properties
  4. 4. The Jackknife
  5. 5. The Bootstrap
  6. Key Take-aways
Where this unit starts. Unit 1 improved an estimator by conditioning on a sufficient statistic, but stopped short of saying the result was the best. Completeness is the missing ingredient, and Lehmann–Scheffé is the theorem that supplies the conclusion. This unit also takes up the large-sample side — CAN and BAN estimators — and two resampling methods that work when no formula is available at all.

1. Completeness

DEFINITION

A family of distributions \(\{g(t;\theta)\}\) of a statistic \(T\) is complete if

\[ E_\theta\big[h(T)\big] = 0 \ \text{ for all } \theta \quad \Longrightarrow \quad h(T) = 0 \ \text{ almost surely.} \]

Read it as: no non-trivial function of \(T\) has expectation zero for every \(\theta\). The consequence is the one that matters: a complete sufficient statistic admits at most one unbiased estimator of any given parametric function. For if \(h_1(T)\) and \(h_2(T)\) were both unbiased for \(\tau(\theta)\) then \(E[h_1(T) - h_2(T)] = 0\) for all \(\theta\), forcing \(h_1 = h_2\).

EXAMPLE 2.1 — PROVING COMPLETENESS, AND SEEING IT FAIL

(a) The Poisson total is complete. \(S \sim\) Poisson\((n\lambda)\). Suppose \(E[h(S)] = 0\) for every \(\lambda > 0\). Then

\[ \sum_{s=0}^{\infty} h(s)\,\frac{e^{-n\lambda}(n\lambda)^{s}}{s!} = 0 \;\Longrightarrow\; \sum_{s=0}^{\infty} \frac{h(s)\,n^{s}}{s!}\,\lambda^{s} = 0 \quad \text{for all } \lambda > 0, \]

after multiplying through by \(e^{n\lambda} > 0\). A power series that vanishes on an interval has every coefficient zero, so \(h(s)n^{s}/s! = 0\) and hence \(h(s) = 0\) for every \(s\). \(\blacksquare\)

(b) A family that is not complete. Let \(X \sim N(\theta, 1)\) and consider the statistic \(T = X\) with \(\theta\) restricted to \(\{-1, +1\}\). Take \(h(t) = t\). Then

\[ E_{\theta = 1}\big[h(X)\big] = 1, \qquad E_{\theta = -1}\big[h(X)\big] = -1, \]

so this particular \(h\) does not have zero expectation throughout. But \(h(t) = t^{2} - 2\) does: \(E_\theta(X^{2}) = \theta^{2} + 1 = 2\) at both \(\theta = 1\) and \(\theta = -1\), so \(E_\theta[h(X)] = 0\) on the whole parameter set, although \(h(X)\) is not zero. The family is therefore not complete. The essential point is structural: completeness is a property of the whole family, and shrinking the parameter set can destroy it. A two-point parameter space carries too little variation to pin down a function on the real line.

Interpretation. Completeness is a richness condition on the family, not on the statistic alone. In the one-parameter exponential family of Distribution Theory, Unit 2, the natural sufficient statistic is complete provided the natural parameter space contains an open interval — which is the usual situation and the reason the theory works so smoothly there.

2. The Lehmann–Scheffé Theorem

STATEMENT AND PROOF

Statement. Let \(S\) be a complete sufficient statistic and let \(T^{*} = h(S)\) be unbiased for \(\tau(\theta)\). Then \(T^{*}\) is the unique UMVU estimator of \(\tau(\theta)\).

Step 1 — take any competitor. Let \(T\) be any unbiased estimator of \(\tau(\theta)\) with finite variance.

Step 2 — Rao–Blackwell it. By Unit 1, \(E(T \mid S)\) is unbiased for \(\tau(\theta)\), is a function of \(S\), and has variance no larger than \(\operatorname{Var}(T)\).

Step 3 — use completeness. Both \(E(T\mid S)\) and \(T^{*}\) are functions of \(S\) and both are unbiased for \(\tau(\theta)\). By completeness there is only one such function, so \(E(T \mid S) = T^{*}\) almost surely.

Step 4 — conclude. Combining Steps 2 and 3,

\[ \operatorname{Var}(T^{*}) = \operatorname{Var}\big(E(T\mid S)\big) \le \operatorname{Var}(T). \]

Since \(T\) was arbitrary, \(T^{*}\) is UMVU; and Step 3 shows it is the only one. \(\blacksquare\)

How it is used in practice. Two recipes, both one step:

  1. Find a complete sufficient \(S\); guess a function of it that is unbiased; done.
  2. Find any unbiased estimator at all, however crude; condition it on \(S\); done.

Neither requires a variance calculation, and neither requires the regularity conditions of Unit 1.

EXAMPLE 2.2 — THE UNIFORM, WHICH CRAMÉR–RAO CANNOT TOUCH

Given. \(X_1, \ldots, X_n\) i.i.d. \(U(0,\theta)\). Find the UMVU estimator of \(\theta\).

Step 1 — note why Unit 1 is unavailable. The support \((0,\theta)\) depends on \(\theta\), so the first regularity condition fails and neither Fisher information nor the Cramér–Rao bound is defined.

Step 2 — find the sufficient statistic by factorisation.

\[ L(\theta) = \prod_{i=1}^{n}\frac{1}{\theta}\mathbf{1}\{0 < x_i < \theta\} = \frac{1}{\theta^{n}}\,\mathbf{1}\{x_{(n)} < \theta\}\,\mathbf{1}\{x_{(1)} > 0\}, \]

which depends on the data only through \(X_{(n)} = \max_i X_i\). So \(X_{(n)}\) is sufficient.

Step 3 — its distribution. From Distribution Theory, Unit 4,

\[ F_{(n)}(y) = \left(\frac{y}{\theta}\right)^{n}, \qquad f_{(n)}(y) = \frac{n y^{n-1}}{\theta^{n}}, \qquad 0 < y < \theta. \]

Step 4 — its mean.

\[ E\left(X_{(n)}\right) = \int_0^\theta y \cdot \frac{n y^{n-1}}{\theta^{n}}\,dy = \frac{n}{\theta^{n}}\left[\frac{y^{n+1}}{n+1}\right]_0^{\theta} = \frac{n}{n+1}\,\theta. \]

So \(X_{(n)}\) is biased downwards, as it must be — it can never exceed \(\theta\).

Step 5 — correct the bias.

\[ T^{*} = \frac{n+1}{n}\,X_{(n)}, \qquad E(T^{*}) = \frac{n+1}{n}\cdot\frac{n}{n+1}\theta = \theta. \]

Step 6 — completeness. Suppose \(E[h(X_{(n)})] = 0\) for all \(\theta\):

\[ \int_0^\theta h(y)\,\frac{n y^{n-1}}{\theta^{n}}\,dy = 0 \;\Longrightarrow\; \int_0^\theta h(y)\,y^{n-1}\,dy = 0 \ \text{ for all } \theta > 0. \]

Differentiating with respect to \(\theta\) — the fundamental theorem of calculus — gives \(h(\theta)\theta^{n-1} = 0\), hence \(h(\theta) = 0\) for every \(\theta > 0\). So the family is complete.

Step 7 — apply Lehmann–Scheffé. \(X_{(n)}\) is complete sufficient and \(T^{*}\) is a function of it that is unbiased, so \(T^{*} = \frac{n+1}{n}X_{(n)}\) is the unique UMVU estimator of \(\theta\).

Step 8 — compare with the obvious alternative. \(2\bar X\) is also unbiased, since \(E(\bar X) = \theta/2\). Their variances are

\[ \operatorname{Var}\left(\frac{n+1}{n}X_{(n)}\right) = \frac{\theta^{2}}{n(n+2)}, \qquad \operatorname{Var}\left(2\bar X\right) = \frac{\theta^{2}}{3n}. \]

The ratio is \(\dfrac{\theta^{2}/[n(n+2)]}{\theta^{2}/(3n)} = \dfrac{3}{n+2}\). At \(n = 10\) that is \(3/12 = 0.25\): the UMVU estimator has a quarter the variance, and the advantage grows without limit with \(n\).

Interpretation. The estimator converges at rate \(1/n\) rather than the usual \(1/\sqrt n\) — one of the few places in statistics where that happens, and a direct consequence of the support boundary carrying information about \(\theta\). This is exactly the case that Unit 1's machinery is blind to.

3. Maximum Likelihood and Its Large-Sample Properties

THE PROPERTIES, AS STATEMENTS (WHICH IS WHAT THE SYLLABUS ASKS) \[ \begin{aligned} &\textbf{(M1) Invariance: } \widehat{g(\theta)} = g(\hat\theta) \text{ for any function } g \\ &\textbf{(M2) } \text{if a sufficient statistic exists, } \hat\theta \text{ is a function of it} \\ &\textbf{(M3) Consistency: } \hat\theta \xrightarrow{P} \theta \text{ under regularity} \\ &\textbf{(M4) Asymptotic normality: } \sqrt{n}\left(\hat\theta - \theta\right) \xrightarrow{d} N\!\left(0, \frac{1}{I(\theta)}\right) \\ &\textbf{(M5) Asymptotic efficiency: } \text{the limiting variance equals the Cramér–Rao bound} \end{aligned} \]

What is not claimed. The MLE need not be unbiased in finite samples — for \(N(\mu,\sigma^{2})\) the MLE of \(\sigma^{2}\) divides by \(n\), not \(n-1\). It need not be unique, and it need not exist. (M1) is the property that has no analogue for unbiased estimation: \(\widehat{\theta^{2}} = \hat\theta^{2}\) always, whereas the square of an unbiased estimator is essentially never unbiased for the square.

CAN AND BAN ESTIMATORS

An estimator \(T_n\) of \(\theta\) is CAN — consistent and asymptotically normal — if

\[ T_n \xrightarrow{P} \theta \qquad\text{and}\qquad \sqrt{n}\left(T_n - \theta\right) \xrightarrow{d} N\big(0, \sigma^{2}(\theta)\big) \]

for some finite \(\sigma^{2}(\theta) > 0\). It is BAN — best asymptotically normal, also called asymptotically efficient — if in addition

\[ \sigma^{2}(\theta) = \frac{1}{I(\theta)}, \]

the smallest limiting variance any CAN estimator can have.

Every BAN estimator is CAN; the converse fails. Both modes of convergence used here are those of Probability Theory, Unit 3, and the asymptotic qualifier is essential: a BAN estimator may be badly behaved at any fixed \(n\), and its optimality is a statement about the limit alone.

EXAMPLE 2.3 — A CAN ESTIMATOR THAT IS NOT BAN

Given. \(X_1, \ldots, X_n\) i.i.d. \(N(\mu, \sigma^{2})\) with \(\sigma^{2}\) known. Compare \(\bar X\) with the sample median \(M_n\).

Step 1 — the mean. \(\sqrt n(\bar X - \mu) \xrightarrow{d} N(0, \sigma^{2})\) by the central limit theorem, and \(I(\mu) = 1/\sigma^{2}\), so the limiting variance \(\sigma^{2}\) equals \(1/I(\mu)\). \(\bar X\) is BAN.

Step 2 — the median. For a sample from a density \(f\) with median \(m\),

\[ \sqrt n\left(M_n - m\right) \xrightarrow{d} N\!\left(0, \frac{1}{4 f(m)^{2}}\right). \]

For the normal, \(f(\mu) = 1/(\sigma\sqrt{2\pi})\), so the limiting variance is

\[ \frac{1}{4}\cdot\frac{1}{f(\mu)^{2}} = \frac{2\pi\sigma^{2}}{4} = \frac{\pi\sigma^{2}}{2}. \]

Step 3 — the efficiency.

\[ \frac{\text{variance of } \bar X}{\text{variance of } M_n} = \frac{\sigma^{2}}{\pi\sigma^{2}/2} = \frac{2}{\pi} = 0.636620. \]

Step 4 — read it. The median is consistent and asymptotically normal, so it is CAN; but its limiting variance is \(\pi/2 = 1.570796\) times the bound, so it is not BAN. Its asymptotic efficiency is \(63.66\%\) — equivalently, the median needs about \(157\) observations to match the mean's \(100\).

Interpretation. That \(36\%\) is the price of robustness, and it is paid only when the data really are normal. Under a Laplace distribution the ranking reverses exactly, because there the median is the MLE — the point made in Distribution Theory, Unit 1. The efficiency comparison is always relative to an assumed model.

4. The Jackknife

THE METHOD

Let \(\hat\theta = \hat\theta(x_1,\ldots,x_n)\) and let \(\hat\theta_{(-i)}\) be the same statistic computed with \(x_i\) omitted. Define

\[ \bar{\hat\theta}_{(\cdot)} = \frac{1}{n}\sum_{i=1}^{n}\hat\theta_{(-i)}, \qquad \widehat{\text{bias}} = (n-1)\left(\bar{\hat\theta}_{(\cdot)} - \hat\theta\right), \] \[ \hat\theta_{\text{jack}} = \hat\theta - \widehat{\text{bias}} = n\hat\theta - (n-1)\bar{\hat\theta}_{(\cdot)}. \]

The pseudo-values \(\tilde\theta_i = n\hat\theta - (n-1)\hat\theta_{(-i)}\) average to \(\hat\theta_{\text{jack}}\), and their sample variance gives the standard error:

\[ \widehat{SE}_{\text{jack}} = \sqrt{\frac{1}{n(n-1)}\sum_{i=1}^{n} \left(\tilde\theta_i - \bar{\tilde\theta}\right)^{2}}. \]

Why it works. If the bias has the form \(a_1/n + a_2/n^{2} + \cdots\), the construction cancels the \(a_1/n\) term exactly, leaving a bias of order \(n^{-2}\). Nothing about the distribution is assumed.

EXAMPLE 2.4 — THE JACKKNIFE RECOVERS THE DIVISOR n−1 EXACTLY

Given. The sample \(3, 5, 7, 11, 14\), and the biased variance \(\hat\theta = \frac{1}{n}\sum_i (x_i - \bar x)^{2}\).

Step 1 — the full-sample values. \(\bar x = 40/5 = 8\), and

\[ \sum_i (x_i - 8)^{2} = 25 + 9 + 1 + 9 + 36 = 80, \qquad \hat\theta = \frac{80}{5} = 16. \]

Step 2 — the five leave-one-out values.

omittedremaining samplemean\(\hat\theta_{(-i)}\)
35, 7, 11, 149.2512.187500
53, 7, 11, 148.7517.187500
73, 5, 11, 148.2519.687500
113, 5, 7, 147.2517.187500
143, 5, 7, 116.508.750000
mean15.000000

Step 3 — the bias estimate.

\[ \widehat{\text{bias}} = (5-1)(15 - 16) = 4 \times (-1) = -4. \]

Step 4 — the corrected estimate.

\[ \hat\theta_{\text{jack}} = 16 - (-4) = 20. \]

Step 5 — compare with the known answer. The unbiased variance is

\[ s^{2} = \frac{80}{5-1} = 20. \]

The jackknife reproduced it exactly, having been told nothing about the correct divisor.

Step 6 — the pseudo-values and the standard error. \(\tilde\theta_i = 5(16) - 4\hat\theta_{(-i)} = 80 - 4\hat\theta_{(-i)}\):

\[ 31.25,\quad 11.25,\quad 1.25,\quad 11.25,\quad 45.00, \]

with mean \(20\) as Step 4 requires. Their sample variance divided by \(n\) gives

\[ \widehat{SE}_{\text{jack}} = 7.925434. \]

Interpretation. The exact recovery is not a coincidence — the bias of the divisor-\(n\) variance is exactly \(-\sigma^{2}/n\), of the form the jackknife is built to remove, so the higher-order terms are absent. It also shows the method's real value: the same four lines apply to a statistic whose bias nobody has worked out, and give a standard error where no formula exists.

5. The Bootstrap

THE METHOD, AND THE THEOREM UNDER IT

From the observed sample \(x_1, \ldots, x_n\), draw \(B\) resamples of size \(n\) with replacement, compute \(\hat\theta^{*}_b\) on each, and use the spread of those \(B\) values as the sampling distribution of \(\hat\theta\):

\[ \widehat{SE}_{\text{boot}} = \sqrt{\frac{1}{B-1}\sum_{b=1}^{B} \left(\hat\theta^{*}_b - \bar{\hat\theta}^{*}\right)^{2}}, \qquad \widehat{\text{bias}} = \bar{\hat\theta}^{*} - \hat\theta. \]

Resampling from the data is resampling from the empirical distribution function \(F_n\), and the justification is the Glivenko–Cantelli lemma of Probability Theory, Unit 3: \(F_n\) converges to \(F\) uniformly, so sampling from \(F_n\) is asymptotically the same as sampling from \(F\). Without that theorem the bootstrap would be a plausible trick; with it, it is a method.

Where it fails. Where \(F_n\) is a poor stand-in for \(F\) in the relevant respect — extreme order statistics, a parameter on the boundary, very small \(n\), or heavy tails with infinite variance.

EXAMPLE 2.5 — THE BOOTSTRAP DONE COMPLETELY, WITHOUT SIMULATION

Given. The sample \(2, 5, 11\), and \(\hat\theta = \bar x\).

Step 1 — enumerate rather than simulate. With \(n = 3\) there are exactly \(3^{3} = 27\) resamples, so the bootstrap distribution can be written down in full and no randomness enters.

Step 2 — the statistic on the original sample. \(\bar x = (2 + 5 + 11)/3 = 6\).

Step 3 — the mean of the 27 bootstrap means.

\[ \bar{\hat\theta}^{*} = 6.000000, \]

exactly equal to \(\bar x\). So the estimated bias is \(0\) — correctly, since the sample mean is unbiased.

Step 4 — the bootstrap variance. Computed over all 27,

\[ \operatorname{Var}^{*}\left(\bar X^{*}\right) = 4.666667, \qquad \widehat{SE}_{\text{boot}} = \sqrt{4.666667} = 2.160247. \]

Step 5 — identify what that number is. The biased sample variance is

\[ s_n^{2} = \frac{(2-6)^{2} + (5-6)^{2} + (11-6)^{2}}{3} = \frac{16 + 1 + 25}{3} = 14, \]

and \(s_n^{2}/n = 14/3 = 4.666667\) — exactly the bootstrap variance. The bootstrap standard error of a mean is therefore \(\sqrt{s_n^{2}/n}\), with divisor \(n\), not the classical \(\sqrt{s^{2}/n} = \sqrt{21/3} = 2.645751\) with divisor \(n-1\).

Interpretation. The two differ by the factor \(\sqrt{(n-1)/n}\), which is \(\sqrt{2/3} = 0.8165\) here and negligible for realistic \(n\). More importantly, the exact agreement in Step 5 shows the bootstrap is not producing an approximation to be hoped about: for the mean it reproduces a known formula exactly. For a statistic with no known formula — a median, a trimmed mean, a ratio, a correlation — the same procedure runs unchanged, and that is the whole reason it exists.

JACKKNIFE AGAINST BOOTSTRAP
JackknifeBootstrap
resamplesexactly \(n\), deterministic\(B\) chosen by the user, random
cost\(n\) refitsusually 1,000 or more refits
reproduciblealwaysonly with a fixed seed
bias correctionremoves the \(1/n\) term exactlyestimates the bias directly
gives a distribution?no — a standard error onlyyes — percentiles, intervals, the whole shape
fails fornon-smooth statistics such as the medianextremes, boundaries, infinite variance

The jackknife's failure for the median is worth knowing: deleting one observation from an odd-sized sample moves the median by a whole gap between order statistics rather than infinitesimally, so the linear approximation the method assumes does not hold. The bootstrap handles the median without difficulty.

Key Take-aways