A decision function \(d(\mathbf{X})\) maps the data to an action. A loss function \(L(\theta, a) \ge 0\) says what it costs to take action \(a\) when the true state is \(\theta\). The risk is the expected loss:
\[ R(\theta, d) = E_\theta\Big[L\big(\theta, d(\mathbf{X})\big)\Big]. \]Note that risk is a function of \(\theta\), not a number. That is the whole difficulty: two decision rules' risk curves generally cross, so neither is better everywhere, and some further principle is needed to choose.
| loss function | \(L(\theta, a)\) | optimal estimator under it |
|---|---|---|
| squared error (quadratic) | \((\theta - a)^{2}\) | the posterior mean |
| absolute error | \(|\theta - a|\) | the posterior median |
| zero-one | \(0\) if \(|a - \theta| < \varepsilon\), else \(1\) | the posterior mode |
Under squared error the risk is the mean squared error,
\[ R(\theta, d) = E_\theta\left[\left(\theta - d\right)^{2}\right] = \operatorname{Var}_\theta(d) + \left[\text{bias}_\theta(d)\right]^{2}, \]so the bias-variance decomposition is a statement about risk under one particular loss, and restricting attention to unbiased estimators — as Units 1 and 2 did — is a choice that forces the second term to zero and then minimises the first.
\(d_1\) dominates \(d_2\) if \(R(\theta, d_1) \le R(\theta, d_2)\) for every \(\theta\), with strict inequality somewhere. A rule is admissible if nothing dominates it, and inadmissible otherwise.
Admissibility only removes rules that are beaten outright; among the survivors a principle is needed:
\[ \begin{aligned} \textbf{Minimax:}\quad & \text{minimise } \sup_\theta R(\theta, d) \quad\text{— guard against the worst case} \\ \textbf{Bayes:}\quad & \text{minimise } r(\pi, d) = \int R(\theta, d)\,\pi(\theta)\,d\theta \quad\text{— average over a prior} \\ \textbf{Unbiasedness:}\quad & \text{restrict to } E_\theta(d) = \theta, \text{ then minimise } R \end{aligned} \]A warning that ought to be better known. Admissibility is weaker than it sounds. Estimating \(\theta\) by the constant \(d \equiv 7\), ignoring the data entirely, is admissible under squared error: no rule can beat it at \(\theta = 7\), where its risk is exactly zero. It is also useless. Admissibility rules out the indefensible; it does not identify the good.
Treat \(\theta\) as random with prior \(\pi(\theta)\). After observing \(\mathbf{x}\), Bayes' theorem gives the posterior
\[ \pi(\theta \mid \mathbf{x}) = \frac{L(\mathbf{x};\theta)\,\pi(\theta)} {\int L(\mathbf{x};\theta')\,\pi(\theta')\,d\theta'} \;\propto\; \text{likelihood} \times \text{prior}. \]The Bayes estimator minimises the posterior expected loss. Under quadratic loss that minimiser is the posterior mean:
\[ \hat\theta_{\text{Bayes}} = E\big(\theta \mid \mathbf{x}\big), \]because \(E[(\theta - a)^{2} \mid \mathbf{x}]\) is a quadratic in \(a\) minimised at \(a = E(\theta\mid\mathbf{x})\) — the same one-line argument that makes the mean the best constant predictor.
A prior family is conjugate if the posterior stays in the family.
| likelihood | conjugate prior | posterior |
|---|---|---|
| Binomial\((n,p)\) | Beta\((a,b)\) | Beta\((a + x,\ b + n - x)\) |
| Poisson\((\lambda)\) | Gamma\((a, b)\) | Gamma\(\left(a + \sum x_i,\ b + n\right)\) |
| \(N(\mu, \sigma^{2})\), \(\sigma^{2}\) known | \(N(\mu_0, \tau^{2})\) | normal, with a precision-weighted mean |
| Exponential\((\theta)\) | Gamma\((a,b)\) | Gamma\(\left(a + n,\ b + \sum x_i\right)\) |
The Poisson–Gamma row is the compound distribution computed in Distribution Theory, Unit 2, Example 2.6, read the other way round: what was there a mixture producing a negative binomial is here a prior producing a posterior.
Given. \(x = 7\) successes in \(n = 20\) trials, with a Beta\((2,3)\) prior on \(p\).
Step 1 — read the prior. Its mean is
\[ \frac{a}{a+b} = \frac{2}{5} = 0.400000, \]and since \(a + b = 5\), it carries about as much information as 5 earlier trials — the interpretation that makes Beta priors easy to choose honestly.
Step 2 — the posterior. By conjugacy,
\[ p \mid x \ \sim\ \text{Beta}(a + x,\ b + n - x) = \text{Beta}(2 + 7,\ 3 + 13) = \text{Beta}(9, 16). \]Step 3 — the Bayes estimate under quadratic loss.
\[ \hat p_{\text{Bayes}} = \frac{a + x}{a + b + n} = \frac{9}{25} = 0.360000. \]Step 4 — compare with the MLE. \(\hat p_{\text{MLE}} = x/n = 7/20 = 0.350000\).
Step 5 — see the shrinkage explicitly. The Bayes estimate is a weighted average of the two:
\[ \hat p_{\text{Bayes}} = w\,\hat p_{\text{MLE}} + (1 - w)\,\frac{a}{a+b}, \qquad w = \frac{n}{a + b + n} = \frac{20}{25} = 0.800000, \] \[ 0.8 \times 0.350000 + 0.2 \times 0.400000 = 0.280000 + 0.080000 = 0.360000. \checkmark \]Step 6 — the posterior spread. For Beta\((\alpha,\beta)\) the variance is \(\alpha\beta/[(\alpha+\beta)^{2}(\alpha+\beta+1)]\):
\[ \operatorname{Var}(p \mid x) = \frac{9 \times 16}{25^{2} \times 26} = \frac{144}{16250} = 0.00886154, \qquad \text{sd} = 0.094136. \]Interpretation. The estimate moves only \(0.01\) from the MLE, because with \(n = 20\) the data carry \(80\%\) of the weight. As \(n\) grows, \(w \to 1\) and the prior washes out — which is the formal content of the remark that Bayesian and frequentist answers agree in large samples. The prior matters most exactly where it should: when the data are few.
A rule \(d^{*}\) is minimax if it minimises the worst-case risk, \(\sup_\theta R(\theta, d^{*}) \le \sup_\theta R(\theta, d)\) for every \(d\).
The standard route to finding one. If a Bayes rule with respect to some prior has constant risk, it is minimax. The reasoning is short: a constant-risk Bayes rule has Bayes risk equal to that constant; no rule can have smaller Bayes risk under that prior; and a rule with smaller maximum risk would also have smaller Bayes risk — a contradiction.
Worked instance. For \(X \sim \text{Bin}(n,p)\) under squared-error loss, the Beta\(\left(\sqrt{n}/2, \sqrt{n}/2\right)\) prior gives the Bayes estimator
\[ \hat p = \frac{X + \sqrt n/2}{n + \sqrt n}, \]whose risk is the constant \(\dfrac{1}{4\left(1 + \sqrt n\right)^{2}}\) — free of \(p\). It is therefore minimax. Compare the MLE \(X/n\), whose risk \(p(1-p)/n\) varies with \(p\), reaching \(1/(4n)\) at \(p = \tfrac12\) and falling to \(0\) at the ends. At \(n = 100\) the minimax risk is \(1/(4 \times 121) = 0.002066\) against the MLE's worst case of \(0.0025\): the minimax rule is better where the MLE is worst, and worse where the MLE is best. Neither dominates, which is precisely why a principle had to be chosen.
Estimate the density \(f\) itself, assuming nothing about its form. The histogram does this crudely — it is discontinuous, and it depends on where the bin edges are placed as well as how wide they are. The methods below remove both defects.
Definition. Count the observations within \(h\) of \(x\) and scale:
\[ \hat f_n(x) = \frac{\#\{i : |X_i - x| < h\}}{2nh} = \frac{F_n(x+h) - F_n(x-h)}{2h}, \]a difference quotient of the empirical distribution function. It is a histogram whose bin is re-centred at whatever \(x\) is being asked about, which removes the bin-placement problem but not the discontinuity.
Bias. Its expectation is
\[ E\big[\hat f_n(x)\big] = \frac{F(x+h) - F(x-h)}{2h}, \]and expanding \(F\) by Taylor's theorem from Mathematical Analysis, Unit 1,
\[ \text{bias} = \frac{h^{2}}{6}f''(x) + O\left(h^{4}\right). \]Variance. The count is binomial with \(n\) trials and success probability \(F(x+h) - F(x-h) \approx 2hf(x)\), so
\[ \operatorname{Var}\big[\hat f_n(x)\big] = \frac{f(x)}{2nh} + O\!\left(\frac{1}{n}\right). \]The trade-off, and the optimal bandwidth. Bias falls with \(h\) while variance rises as \(h\) shrinks. Adding them,
\[ \text{MSE} \approx \frac{h^{4}}{36}\left[f''(x)\right]^{2} + \frac{f(x)}{2nh}, \]and differentiating with respect to \(h\) gives \(h_{\text{opt}} \propto n^{-1/5}\), whence
\[ \text{MSE}\big(h_{\text{opt}}\big) = O\!\left(n^{-4/5}\right). \]Read that rate. A parametric estimator has MSE of order \(n^{-1}\). Density estimation is strictly slower — \(n^{-4/5}\) — and that gap is the price of assuming nothing about the shape. Consistency holds provided \(h \to 0\) and \(nh \to \infty\): the window must shrink, but not faster than the sample grows, or it ends up empty.
Definition. Replace the sharp count by a smooth weight:
\[ \hat f_n(x) = \frac{1}{nh}\sum_{i=1}^{n} K\!\left(\frac{x - X_i}{h}\right), \]where the kernel \(K\) satisfies \(\int K = 1\), \(K \ge 0\), and is usually symmetric.
| kernel | \(K(u)\) | note |
|---|---|---|
| uniform | \(\tfrac12\) on \(|u| < 1\) | recovers Rosenblatt's estimator exactly |
| Gaussian | \(\dfrac{1}{\sqrt{2\pi}}e^{-u^{2}/2}\) | infinitely smooth; unbounded support |
| Epanechnikov | \(\tfrac34(1 - u^{2})\) on \(|u| < 1\) | minimises the asymptotic MISE |
The result that settles practice. The asymptotic mean integrated squared error is
\[ \text{MISE} \approx \frac{h^{4}}{4}\mu_2(K)^{2}\int\left[f''\right]^{2} + \frac{1}{nh}\int K^{2}, \]again minimised at \(h \propto n^{-1/5}\) with MISE of order \(n^{-4/5}\) — the same rate as the naive estimator. The choice of kernel barely matters; the choice of bandwidth matters enormously. Epanechnikov is optimal, but the Gaussian kernel is only about \(5\%\) less efficient, whereas halving or doubling \(h\) changes the picture completely.
Given. The sample \(2.1,\ 2.8,\ 3.3,\ 4.0,\ 5.2\), a Gaussian kernel and bandwidth \(h = 0.8\). Estimate \(f(3.0)\).
Step 1 — standardise each observation. \(u_i = (x - x_i)/h = (3.0 - x_i)/0.8\), then evaluate \(K(u) = e^{-u^{2}/2}/\sqrt{2\pi}\):
| \(x_i\) | \(u_i\) | \(K(u_i)\) |
|---|---|---|
| 2.1 | 1.1250 | 0.211877 |
| 2.8 | 0.2500 | 0.386668 |
| 3.3 | −0.3750 | 0.371855 |
| 4.0 | −1.2500 | 0.182649 |
| 5.2 | −2.7500 | 0.009094 |
| total | 1.162143 | |
Step 2 — scale by \(nh\).
\[ \hat f(3.0) = \frac{1.162143}{5 \times 0.8} = \frac{1.162143}{4} = 0.290536. \]Step 3 — compare with the naive estimator. Two observations, \(2.8\) and \(3.3\), lie within \(0.8\) of \(3.0\), so
\[ \hat f_{\text{naive}}(3.0) = \frac{2}{2 \times 5 \times 0.8} = \frac{2}{8} = 0.250000. \]Interpretation. Note the weights in Step 1. The nearest point, \(2.8\), contributes \(0.3867\); the farthest, \(5.2\), contributes \(0.0091\) — more than forty times less, but not zero. That is the whole difference from the naive estimator, which counts \(2.8\) and \(3.3\) fully, \(4.0\) not at all, and would jump discontinuously the instant \(4.0\) crossed the window edge. The kernel estimate is smooth in \(x\) and its derivative exists, which is what makes it usable for finding modes, and for the plots that accompany every fitted distribution in Distribution Theory, Practical STS-107.