Unbiasedness alone settles nothing: for a normal sample both \(\bar X\) and the single observation \(X_1\) are unbiased for \(\mu\), and so is any weighted average of the observations with weights summing to 1. There are infinitely many unbiased estimators, and they differ enormously in variance.
An estimator \(T\) is UMVU (uniformly minimum variance unbiased) for \(\tau(\theta)\) if \(E_\theta(T) = \tau(\theta)\) for all \(\theta\) and \(\operatorname{Var}_\theta(T) \le \operatorname{Var}_\theta(T^{*})\) for every other unbiased \(T^{*}\) and every \(\theta\). The word uniformly is doing work: the same estimator must win at every parameter value, not just at a convenient one.
Two routes lead to it, and this course takes both:
For a family \(f(x;\theta)\) satisfying the regularity conditions below, the score is \(U(\theta) = \dfrac{\partial}{\partial\theta}\ln f(X;\theta)\), and Fisher's information is
\[ I(\theta) = E\left[U(\theta)^{2}\right] = \operatorname{Var}\big(U(\theta)\big) = -E\left[\frac{\partial^{2}}{\partial\theta^{2}}\ln f(X;\theta)\right]. \]The second equality holds because \(E[U(\theta)] = 0\); the third is obtained by differentiating that identity once more.
Why the score has mean zero.
\[ E\big[U(\theta)\big] = \int \frac{\partial \ln f}{\partial\theta} f \, dx = \int \frac{1}{f}\frac{\partial f}{\partial\theta} f \, dx = \int \frac{\partial f}{\partial\theta}\,dx = \frac{\partial}{\partial\theta}\int f \, dx = \frac{\partial}{\partial\theta}(1) = 0. \]The only questionable step is the last interchange of \(\partial/\partial\theta\) with \(\int\), and that is exactly Leibniz's rule from Mathematical Analysis, Unit 3. The "regularity conditions" quoted throughout this course are its hypotheses, no more and no less.
Additivity. For a random sample of size \(n\) the log-likelihood is a sum, so the information adds: \(I_n(\theta) = n\,I(\theta)\). Information accumulates linearly in the sample size, which is the origin of the \(1/n\) in every variance formula.
\(U(0,\theta)\) fails the first, because its support \((0,\theta)\) moves with the parameter. Everything in sections 3 and 4 is therefore unavailable for it. And yet its UMVU estimator exists and is found easily by the Lehmann–Scheffé route of Unit 2 — which is the practical argument for preferring that route.
Statement. Under the regularity conditions, for any estimator \(T\) with \(E_\theta(T) = \tau(\theta)\),
\[ \operatorname{Var}_\theta(T) \;\ge\; \frac{\left[\tau'(\theta)\right]^{2}}{n\,I(\theta)}. \]In particular, for an unbiased estimator of \(\theta\) itself, \(\operatorname{Var}(T) \ge 1/\big(nI(\theta)\big)\).
Step 1 — differentiate the unbiasedness identity. From \(\int T(x) f(x;\theta)\,dx = \tau(\theta)\), differentiating under the integral sign,
\[ \int T(x)\,\frac{\partial f}{\partial\theta}\,dx = \tau'(\theta). \]Step 2 — introduce the score. Writing \(\partial f/\partial\theta = f \cdot \partial \ln f/\partial\theta = f\,U\),
\[ E\big[T\,U\big] = \tau'(\theta). \]Step 3 — centre it. Since \(E(U) = 0\),
\[ \operatorname{Cov}(T, U) = E(TU) - E(T)E(U) = \tau'(\theta) - 0 = \tau'(\theta). \]Step 4 — apply Cauchy–Schwarz. From Probability Theory, Unit 2, \(\left[\operatorname{Cov}(T,U)\right]^{2} \le \operatorname{Var}(T)\operatorname{Var}(U)\), so
\[ \left[\tau'(\theta)\right]^{2} \le \operatorname{Var}(T) \cdot n I(\theta), \]using \(\operatorname{Var}(U) = nI(\theta)\) for a sample of size \(n\).
Step 5 — divide. Since \(nI(\theta) > 0\),
\[ \operatorname{Var}(T) \ge \frac{\left[\tau'(\theta)\right]^{2}}{n I(\theta)}. \quad \blacksquare \]When is the bound attained? Exactly when Cauchy–Schwarz holds with equality, that is when \(T - \tau(\theta)\) is proportional to \(U\):
\[ \frac{\partial \ln L}{\partial\theta} = k(\theta)\left[T - \tau(\theta)\right]. \]That happens only in the one-parameter exponential family of Distribution Theory, Unit 2, with \(T\) the natural sufficient statistic. Outside it the bound is a bound and not a target.
(a) Poisson mean. \(X_1,\ldots,X_{25}\) i.i.d. Poisson\((\lambda)\).
Step 1 — the log-density and its derivatives. \(\ln f = -\lambda + x\ln\lambda - \ln x!\), so
\[ U = \frac{\partial \ln f}{\partial\lambda} = -1 + \frac{x}{\lambda}, \qquad \frac{\partial^{2}\ln f}{\partial\lambda^{2}} = -\frac{x}{\lambda^{2}}. \]Step 2 — the information. Using the third form,
\[ I(\lambda) = -E\left(-\frac{X}{\lambda^{2}}\right) = \frac{E(X)}{\lambda^{2}} = \frac{\lambda}{\lambda^{2}} = \frac{1}{\lambda}. \]At \(\lambda = 4\): \(I(4) = 0.25\) per observation, and \(25 \times 0.25 = 6.25\) in total.
Step 3 — the bound. \(\operatorname{Var}(T) \ge \dfrac{1}{nI(\lambda)} = \dfrac{\lambda}{n} = \dfrac{4}{25} = 0.16\).
Step 4 — compare with \(\bar X\). \(\operatorname{Var}(\bar X) = \lambda/n = 0.16\). The bound is attained, so \(\bar X\) is UMVU for \(\lambda\).
Step 5 — confirm through the attainment condition. The score for the whole sample is
\[ \frac{\partial \ln L}{\partial\lambda} = -n + \frac{\sum x_i}{\lambda} = \frac{n}{\lambda}\left(\bar x - \lambda\right), \]which is of the required form \(k(\lambda)[T - \tau(\lambda)]\) with \(k(\lambda) = n/\lambda\) and \(T = \bar X\). \(\checkmark\)
(b) Normal variance. \(X_1, \ldots, X_{20}\) i.i.d. \(N(\mu, \sigma^{2})\) with \(\sigma^{2} = 9\) to be estimated and \(\mu\) unknown.
Step 1 — the bound. The information about \(\sigma^{2}\) is \(1/(2\sigma^{4})\) per observation, so
\[ \text{CRLB} = \frac{2\sigma^{4}}{n} = \frac{2 \times 81}{20} = 8.1. \]Step 2 — the variance of \(S^{2}\). From Distribution Theory, Unit 3,
\[ \operatorname{Var}(S^{2}) = \frac{2\sigma^{4}}{n-1} = \frac{162}{19} = 8.526316. \]Step 3 — compare.
\[ \frac{\operatorname{Var}(S^{2})}{\text{CRLB}} = \frac{8.526316}{8.1} = 1.052632 = \frac{n}{n-1}, \]so the efficiency is \((n-1)/n = 19/20 = 0.95\).
Interpretation. \(S^{2}\) misses the Cramér–Rao bound by \(5\%\) and is nevertheless UMVU — proved in Unit 2 by Lehmann–Scheffé. The bound is simply not attainable here. Reading "does not attain the bound" as "is not best" is the commonest error in this subject; the bound is a guarantee that nothing does better, not a promise that something does as well.
Cramér–Rao uses only the first derivative of the likelihood. The Bhattacharya system uses the first \(k\):
\[ \operatorname{Var}(T) \;\ge\; \mathbf{b}'\,J_k^{-1}\,\mathbf{b}, \qquad b_r = \frac{d^{r}\tau(\theta)}{d\theta^{r}}, \qquad \left(J_k\right)_{rs} = E\left[\frac{L^{(r)}}{L}\cdot\frac{L^{(s)}}{L}\right], \]where \(L^{(r)} = \partial^{r} L/\partial\theta^{r}\).
Three facts that fix its place. At \(k = 1\) it reduces exactly to Cramér–Rao. The bounds increase with \(k\), so each is at least as sharp as the last. And when the Cramér–Rao bound is attained, every higher Bhattacharya bound equals it — the extra derivatives add nothing.
So it is worth the work only when Cramér–Rao is known not to be attained. In practice that is rare, because Lehmann–Scheffé usually identifies the UMVU estimator outright, making any bound unnecessary. The Bhattacharya system matters as a theoretical statement about how much information successive derivatives carry.
\(T\) is sufficient for \(\theta\) if the conditional distribution of the sample given \(T\) does not involve \(\theta\). The factorisation criterion is the working test:
\[ L(\theta; \mathbf{x}) = g\big(T(\mathbf{x}); \theta\big)\, h(\mathbf{x}). \]In the exponential family of Distribution Theory, Unit 2 the factorisation is immediate and \(T\) is the natural sufficient statistic — which is why that family and this theory are always met together.
Statement. Let \(T\) be unbiased for \(\tau(\theta)\) with finite variance and let \(S\) be sufficient. Put \(T^{*} = E(T \mid S)\). Then
Proof of (2). The tower property from Probability Theory, Unit 2:
\[ E(T^{*}) = E\big[E(T \mid S)\big] = E(T) = \tau(\theta). \]Proof of (3). The variance decomposition from the same unit:
\[ \operatorname{Var}(T) = \underbrace{E\big[\operatorname{Var}(T \mid S)\big]}_{\ge 0} + \underbrace{\operatorname{Var}\big[E(T \mid S)\big]}_{= \operatorname{Var}(T^{*})} \;\ge\; \operatorname{Var}(T^{*}). \quad \blacksquare \]The whole theorem is those two identities, applied in order. The discarded term \(E[\operatorname{Var}(T\mid S)]\) is exactly the variability that conditioning removes, and it is zero only when \(T\) was already a function of \(S\).
Given. \(X_1, \ldots, X_5\) i.i.d. Poisson\((\lambda)\). Estimate
\[ \tau(\lambda) = P(X = 0) = e^{-\lambda}. \]Step 1 — a crude unbiased estimator. Let \(T = \mathbf{1}\{X_1 = 0\}\). Then \(E(T) = P(X_1 = 0) = e^{-\lambda}\), so it is unbiased — but it uses only one observation and takes only the values \(0\) and \(1\), which is plainly wasteful.
Its variance is that of a Bernoulli with \(p = e^{-\lambda}\):
\[ \operatorname{Var}(T) = p(1-p) = e^{-\lambda}\left(1 - e^{-\lambda}\right). \]At \(\lambda = 1\): \(0.367879 \times 0.632121 = 0.232544\).
Step 2 — the sufficient statistic. By factorisation, \(S = \sum_{i} X_i\) is sufficient, and \(S \sim\) Poisson\((n\lambda)\).
Step 3 — the conditional distribution. Given \(S = s\), the \(s\) events are distributed among the \(n\) observations independently and equally, so
\[ X_1 \mid S = s \ \sim\ \text{Bin}\!\left(s, \frac{1}{n}\right). \]Step 4 — take the conditional expectation.
\[ T^{*} = E(T \mid S = s) = P(X_1 = 0 \mid S = s) = \binom{s}{0}\left(\frac1n\right)^{0}\left(1 - \frac1n\right)^{s} = \left(\frac{n-1}{n}\right)^{s}. \]With \(n = 5\) the Rao–Blackwell estimator is \((0.8)^{S}\) — a function of the whole sample, and no longer restricted to two values.
Step 5 — check unbiasedness directly. With \(S \sim\) Poisson\((5)\) at \(\lambda = 1\),
\[ E\left[(0.8)^{S}\right] = \sum_{k \ge 0} (0.8)^{k}\frac{e^{-5}5^{k}}{k!} = e^{-5}e^{5 \times 0.8} = e^{-5 + 4} = e^{-1} = 0.367879, \]using the exponential series. Computing the sum numerically gives \(0.367879\). \(\checkmark\)
Step 6 — the variance, and the gain. By the same device, \(E[(0.8)^{2S}] = E[(0.64)^{S}] = e^{-5 + 3.2} = e^{-1.8} = 0.165299\), so
\[ \operatorname{Var}(T^{*}) = 0.165299 - (0.367879)^{2} = 0.165299 - 0.135335 = 0.029964. \]Step 7 — compare.
\[ \frac{\operatorname{Var}(T)}{\operatorname{Var}(T^{*})} = \frac{0.232544}{0.029964} = 7.761. \]Interpretation. Conditioning on the sufficient statistic cut the variance by a factor of nearly eight, from one line of algebra and no new data. And because \(S\) is also complete — the subject of Unit 2 — the result is not merely better but the best: \((0.8)^{S}\) is UMVU for \(e^{-\lambda}\). Note also what the Cramér–Rao route would have given here: a bound of \(\lambda e^{-2\lambda}/n = 0.027067\) at \(\lambda = 1\), which is below \(0.029964\) and therefore not attained. The bound would have left the question open; conditioning settled it.