Skip to the content

Topics Covered

Loss & Risk Admissibility Bayes Estimation Conjugate Priors Minimax Rules Rosenblatt's Estimator Kernel Density Estimation
On this page
  1. 1. Loss, Risk and Decision Functions
  2. 2. Bayes Estimation
  3. 3. Non-Parametric Density Estimation
  4. Key Take-aways
Where this unit starts. Units 1 to 3 judged an estimator by unbiasedness and variance. That is one criterion among many, and decision theory asks the prior question: what is it going to cost to be wrong? Estimation and testing both become special cases. The unit then turns to estimating a whole density rather than a parameter, which is where the course's last section, non-parametric estimation, sits.

1. Loss, Risk and Decision Functions

THE THREE OBJECTS

A decision function \(d(\mathbf{X})\) maps the data to an action. A loss function \(L(\theta, a) \ge 0\) says what it costs to take action \(a\) when the true state is \(\theta\). The risk is the expected loss:

\[ R(\theta, d) = E_\theta\Big[L\big(\theta, d(\mathbf{X})\big)\Big]. \]

Note that risk is a function of \(\theta\), not a number. That is the whole difficulty: two decision rules' risk curves generally cross, so neither is better everywhere, and some further principle is needed to choose.

loss function\(L(\theta, a)\)optimal estimator under it
squared error (quadratic)\((\theta - a)^{2}\)the posterior mean
absolute error\(|\theta - a|\)the posterior median
zero-one\(0\) if \(|a - \theta| < \varepsilon\), else \(1\)the posterior mode

Under squared error the risk is the mean squared error,

\[ R(\theta, d) = E_\theta\left[\left(\theta - d\right)^{2}\right] = \operatorname{Var}_\theta(d) + \left[\text{bias}_\theta(d)\right]^{2}, \]

so the bias-variance decomposition is a statement about risk under one particular loss, and restricting attention to unbiased estimators — as Units 1 and 2 did — is a choice that forces the second term to zero and then minimises the first.

ADMISSIBILITY, AND THE THREE PRINCIPLES

\(d_1\) dominates \(d_2\) if \(R(\theta, d_1) \le R(\theta, d_2)\) for every \(\theta\), with strict inequality somewhere. A rule is admissible if nothing dominates it, and inadmissible otherwise.

Admissibility only removes rules that are beaten outright; among the survivors a principle is needed:

\[ \begin{aligned} \textbf{Minimax:}\quad & \text{minimise } \sup_\theta R(\theta, d) \quad\text{— guard against the worst case} \\ \textbf{Bayes:}\quad & \text{minimise } r(\pi, d) = \int R(\theta, d)\,\pi(\theta)\,d\theta \quad\text{— average over a prior} \\ \textbf{Unbiasedness:}\quad & \text{restrict to } E_\theta(d) = \theta, \text{ then minimise } R \end{aligned} \]

A warning that ought to be better known. Admissibility is weaker than it sounds. Estimating \(\theta\) by the constant \(d \equiv 7\), ignoring the data entirely, is admissible under squared error: no rule can beat it at \(\theta = 7\), where its risk is exactly zero. It is also useless. Admissibility rules out the indefensible; it does not identify the good.

2. Bayes Estimation

PRIOR, POSTERIOR AND THE BAYES RULE

Treat \(\theta\) as random with prior \(\pi(\theta)\). After observing \(\mathbf{x}\), Bayes' theorem gives the posterior

\[ \pi(\theta \mid \mathbf{x}) = \frac{L(\mathbf{x};\theta)\,\pi(\theta)} {\int L(\mathbf{x};\theta')\,\pi(\theta')\,d\theta'} \;\propto\; \text{likelihood} \times \text{prior}. \]

The Bayes estimator minimises the posterior expected loss. Under quadratic loss that minimiser is the posterior mean:

\[ \hat\theta_{\text{Bayes}} = E\big(\theta \mid \mathbf{x}\big), \]

because \(E[(\theta - a)^{2} \mid \mathbf{x}]\) is a quadratic in \(a\) minimised at \(a = E(\theta\mid\mathbf{x})\) — the same one-line argument that makes the mean the best constant predictor.

A prior family is conjugate if the posterior stays in the family.

likelihoodconjugate priorposterior
Binomial\((n,p)\)Beta\((a,b)\)Beta\((a + x,\ b + n - x)\)
Poisson\((\lambda)\)Gamma\((a, b)\)Gamma\(\left(a + \sum x_i,\ b + n\right)\)
\(N(\mu, \sigma^{2})\), \(\sigma^{2}\) known\(N(\mu_0, \tau^{2})\)normal, with a precision-weighted mean
Exponential\((\theta)\)Gamma\((a,b)\)Gamma\(\left(a + n,\ b + \sum x_i\right)\)

The Poisson–Gamma row is the compound distribution computed in Distribution Theory, Unit 2, Example 2.6, read the other way round: what was there a mixture producing a negative binomial is here a prior producing a posterior.

EXAMPLE 4.1 — A BAYES ESTIMATE, AND WHAT THE PRIOR IS WORTH

Given. \(x = 7\) successes in \(n = 20\) trials, with a Beta\((2,3)\) prior on \(p\).

Step 1 — read the prior. Its mean is

\[ \frac{a}{a+b} = \frac{2}{5} = 0.400000, \]

and since \(a + b = 5\), it carries about as much information as 5 earlier trials — the interpretation that makes Beta priors easy to choose honestly.

Step 2 — the posterior. By conjugacy,

\[ p \mid x \ \sim\ \text{Beta}(a + x,\ b + n - x) = \text{Beta}(2 + 7,\ 3 + 13) = \text{Beta}(9, 16). \]

Step 3 — the Bayes estimate under quadratic loss.

\[ \hat p_{\text{Bayes}} = \frac{a + x}{a + b + n} = \frac{9}{25} = 0.360000. \]

Step 4 — compare with the MLE. \(\hat p_{\text{MLE}} = x/n = 7/20 = 0.350000\).

Step 5 — see the shrinkage explicitly. The Bayes estimate is a weighted average of the two:

\[ \hat p_{\text{Bayes}} = w\,\hat p_{\text{MLE}} + (1 - w)\,\frac{a}{a+b}, \qquad w = \frac{n}{a + b + n} = \frac{20}{25} = 0.800000, \] \[ 0.8 \times 0.350000 + 0.2 \times 0.400000 = 0.280000 + 0.080000 = 0.360000. \checkmark \]

Step 6 — the posterior spread. For Beta\((\alpha,\beta)\) the variance is \(\alpha\beta/[(\alpha+\beta)^{2}(\alpha+\beta+1)]\):

\[ \operatorname{Var}(p \mid x) = \frac{9 \times 16}{25^{2} \times 26} = \frac{144}{16250} = 0.00886154, \qquad \text{sd} = 0.094136. \]

Interpretation. The estimate moves only \(0.01\) from the MLE, because with \(n = 20\) the data carry \(80\%\) of the weight. As \(n\) grows, \(w \to 1\) and the prior washes out — which is the formal content of the remark that Bayesian and frequentist answers agree in large samples. The prior matters most exactly where it should: when the data are few.

MINIMAX RULES, AND THE BRIDGE FROM BAYES

A rule \(d^{*}\) is minimax if it minimises the worst-case risk, \(\sup_\theta R(\theta, d^{*}) \le \sup_\theta R(\theta, d)\) for every \(d\).

The standard route to finding one. If a Bayes rule with respect to some prior has constant risk, it is minimax. The reasoning is short: a constant-risk Bayes rule has Bayes risk equal to that constant; no rule can have smaller Bayes risk under that prior; and a rule with smaller maximum risk would also have smaller Bayes risk — a contradiction.

Worked instance. For \(X \sim \text{Bin}(n,p)\) under squared-error loss, the Beta\(\left(\sqrt{n}/2, \sqrt{n}/2\right)\) prior gives the Bayes estimator

\[ \hat p = \frac{X + \sqrt n/2}{n + \sqrt n}, \]

whose risk is the constant \(\dfrac{1}{4\left(1 + \sqrt n\right)^{2}}\) — free of \(p\). It is therefore minimax. Compare the MLE \(X/n\), whose risk \(p(1-p)/n\) varies with \(p\), reaching \(1/(4n)\) at \(p = \tfrac12\) and falling to \(0\) at the ends. At \(n = 100\) the minimax risk is \(1/(4 \times 121) = 0.002066\) against the MLE's worst case of \(0.0025\): the minimax rule is better where the MLE is worst, and worse where the MLE is best. Neither dominates, which is precisely why a principle had to be chosen.

3. Non-Parametric Density Estimation

THE PROBLEM

Estimate the density \(f\) itself, assuming nothing about its form. The histogram does this crudely — it is discontinuous, and it depends on where the bin edges are placed as well as how wide they are. The methods below remove both defects.

ROSENBLATT'S NAIVE ESTIMATOR, WITH ITS BIAS AND VARIANCE

Definition. Count the observations within \(h\) of \(x\) and scale:

\[ \hat f_n(x) = \frac{\#\{i : |X_i - x| < h\}}{2nh} = \frac{F_n(x+h) - F_n(x-h)}{2h}, \]

a difference quotient of the empirical distribution function. It is a histogram whose bin is re-centred at whatever \(x\) is being asked about, which removes the bin-placement problem but not the discontinuity.

Bias. Its expectation is

\[ E\big[\hat f_n(x)\big] = \frac{F(x+h) - F(x-h)}{2h}, \]

and expanding \(F\) by Taylor's theorem from Mathematical Analysis, Unit 1,

\[ \text{bias} = \frac{h^{2}}{6}f''(x) + O\left(h^{4}\right). \]

Variance. The count is binomial with \(n\) trials and success probability \(F(x+h) - F(x-h) \approx 2hf(x)\), so

\[ \operatorname{Var}\big[\hat f_n(x)\big] = \frac{f(x)}{2nh} + O\!\left(\frac{1}{n}\right). \]

The trade-off, and the optimal bandwidth. Bias falls with \(h\) while variance rises as \(h\) shrinks. Adding them,

\[ \text{MSE} \approx \frac{h^{4}}{36}\left[f''(x)\right]^{2} + \frac{f(x)}{2nh}, \]

and differentiating with respect to \(h\) gives \(h_{\text{opt}} \propto n^{-1/5}\), whence

\[ \text{MSE}\big(h_{\text{opt}}\big) = O\!\left(n^{-4/5}\right). \]

Read that rate. A parametric estimator has MSE of order \(n^{-1}\). Density estimation is strictly slower — \(n^{-4/5}\) — and that gap is the price of assuming nothing about the shape. Consistency holds provided \(h \to 0\) and \(nh \to \infty\): the window must shrink, but not faster than the sample grows, or it ends up empty.

KERNEL DENSITY ESTIMATORS

Definition. Replace the sharp count by a smooth weight:

\[ \hat f_n(x) = \frac{1}{nh}\sum_{i=1}^{n} K\!\left(\frac{x - X_i}{h}\right), \]

where the kernel \(K\) satisfies \(\int K = 1\), \(K \ge 0\), and is usually symmetric.

kernel\(K(u)\)note
uniform\(\tfrac12\) on \(|u| < 1\)recovers Rosenblatt's estimator exactly
Gaussian\(\dfrac{1}{\sqrt{2\pi}}e^{-u^{2}/2}\)infinitely smooth; unbounded support
Epanechnikov\(\tfrac34(1 - u^{2})\) on \(|u| < 1\)minimises the asymptotic MISE

The result that settles practice. The asymptotic mean integrated squared error is

\[ \text{MISE} \approx \frac{h^{4}}{4}\mu_2(K)^{2}\int\left[f''\right]^{2} + \frac{1}{nh}\int K^{2}, \]

again minimised at \(h \propto n^{-1/5}\) with MISE of order \(n^{-4/5}\) — the same rate as the naive estimator. The choice of kernel barely matters; the choice of bandwidth matters enormously. Epanechnikov is optimal, but the Gaussian kernel is only about \(5\%\) less efficient, whereas halving or doubling \(h\) changes the picture completely.

EXAMPLE 4.2 — A KERNEL ESTIMATE COMPUTED BY HAND

Given. The sample \(2.1,\ 2.8,\ 3.3,\ 4.0,\ 5.2\), a Gaussian kernel and bandwidth \(h = 0.8\). Estimate \(f(3.0)\).

Step 1 — standardise each observation. \(u_i = (x - x_i)/h = (3.0 - x_i)/0.8\), then evaluate \(K(u) = e^{-u^{2}/2}/\sqrt{2\pi}\):

\(x_i\)\(u_i\)\(K(u_i)\)
2.11.12500.211877
2.80.25000.386668
3.3−0.37500.371855
4.0−1.25000.182649
5.2−2.75000.009094
total1.162143

Step 2 — scale by \(nh\).

\[ \hat f(3.0) = \frac{1.162143}{5 \times 0.8} = \frac{1.162143}{4} = 0.290536. \]

Step 3 — compare with the naive estimator. Two observations, \(2.8\) and \(3.3\), lie within \(0.8\) of \(3.0\), so

\[ \hat f_{\text{naive}}(3.0) = \frac{2}{2 \times 5 \times 0.8} = \frac{2}{8} = 0.250000. \]

Interpretation. Note the weights in Step 1. The nearest point, \(2.8\), contributes \(0.3867\); the farthest, \(5.2\), contributes \(0.0091\) — more than forty times less, but not zero. That is the whole difference from the naive estimator, which counts \(2.8\) and \(3.3\) fully, \(4.0\) not at all, and would jump discontinuously the instant \(4.0\) crossed the window edge. The kernel estimate is smooth in \(x\) and its derivative exists, which is what makes it usable for finding modes, and for the plots that accompany every fitted distribution in Distribution Theory, Practical STS-107.

Key Take-aways