Let \(h(x_1, \ldots, x_m)\) be a symmetric function of \(m\) arguments, called the kernel, with
\[ E\big[h(X_1, \ldots, X_m)\big] = \theta. \]The corresponding U-statistic of degree \(m\) is the average of the kernel over all subsets of size \(m\):
\[ U_n = \binom{n}{m}^{-1} \sum_{1 \le i_1 < \cdots < i_m \le n} h\left(X_{i_1}, \ldots, X_{i_m}\right). \]Unbiasedness is automatic. Every term has expectation \(\theta\), and an average of terms each with expectation \(\theta\) has expectation \(\theta\). So a U-statistic is unbiased by construction, whatever the underlying distribution — it is distribution-free in that sense.
| parameter \(\theta\) | kernel \(h\) | \(m\) | the U-statistic is |
|---|---|---|---|
| \(\mu = E(X)\) | \(h(x) = x\) | 1 | \(\bar X\) |
| \(\sigma^{2}\) | \(h(x_1,x_2) = \tfrac12(x_1 - x_2)^{2}\) | 2 | \(S^{2}\) |
| \(P(X_1 + X_2 > 0)\) | \(h = \mathbf{1}\{x_1 + x_2 > 0\}\) | 2 | the Wilcoxon signed-rank statistic |
| Gini mean difference | \(h(x_1,x_2) = |x_1 - x_2|\) | 2 | the Gini estimator |
| Kendall's \(\tau\) | sign of concordance | 2 | Kendall's \(\tau\) |
The third and fifth rows explain why this section sits in an estimation paper at all: the rank tests of Inferential Statistics, Unit 5 are U-statistics, so one asymptotic theorem covers all of them.
Given. The sample \(3, 5, 7, 11, 14\), and the kernel \(h(x_1, x_2) = \tfrac12(x_1 - x_2)^{2}\).
Step 1 — check the kernel is unbiased for \(\sigma^{2}\). For independent \(X_1, X_2\) with the same distribution,
\[ E\left[\tfrac12 (X_1 - X_2)^{2}\right] = \tfrac12 \operatorname{Var}(X_1 - X_2) = \tfrac12\left(\sigma^{2} + \sigma^{2}\right) = \sigma^{2}, \]using \(E(X_1 - X_2) = 0\) and independence. \(\checkmark\)
Step 2 — evaluate the kernel on all \(\binom{5}{2} = 10\) pairs.
| pair | \(\tfrac12(x_i - x_j)^{2}\) | pair | \(\tfrac12(x_i - x_j)^{2}\) |
|---|---|---|---|
| (3, 5) | 2 | (5, 14) | 40.5 |
| (3, 7) | 8 | (7, 11) | 8 |
| (3, 11) | 32 | (7, 14) | 24.5 |
| (3, 14) | 60.5 | (11, 14) | 4.5 |
| (5, 7) | 2 | total 200 | |
| (5, 11) | 18 | ||
Step 3 — average.
\[ U = \frac{200}{10} = 20. \]Step 4 — compare with the ordinary sample variance. With \(\bar x = 8\), \(\sum(x_i - \bar x)^{2} = 80\), so
\[ s^{2} = \frac{80}{5 - 1} = 20. \]The two agree exactly, and this is an identity, not a coincidence of these numbers: the algebraic expansion of \(\sum_{i<j}(x_i - x_j)^{2}\) gives \(n\sum_i(x_i - \bar x)^{2}\).
Interpretation. The familiar \(S^{2}\) is the U-statistic for the kernel \(\tfrac12(x_1-x_2)^{2}\), which explains without further work why the divisor is \(n-1\) and why it is unbiased for every distribution with a finite variance, not merely for the normal. The same sample gave \(20\) by the jackknife in Unit 2, Example 2.4 — three routes, one number.
Define the first projection
\[ h_1(x) = E\big[h(x, X_2, \ldots, X_m)\big], \qquad \zeta_1 = \operatorname{Var}\big(h_1(X_1)\big). \]Then, provided \(\zeta_1 > 0\) and \(E(h^{2}) < \infty\),
\[ \sqrt{n}\left(U_n - \theta\right) \xrightarrow{d} N\!\left(0,\; m^{2}\zeta_1\right). \]Why the projection appears. \(U_n\) is an average of dependent terms — any two subsets sharing an observation are correlated — so the central limit theorem of Probability Theory, Unit 4 does not apply directly. Hoeffding's argument replaces \(U_n\) by the average of the independent quantities \(h_1(X_i)\), shows the difference is \(o_P(n^{-1/2})\), and applies the CLT to that. The factor \(m^{2}\) counts how many of the \(m\) slots each observation occupies.
What it gives. Every U-statistic is CAN in the sense of Unit 2, with a variance that can be estimated. That single theorem provides the large-sample normal approximation used by Wilcoxon, Mann–Whitney and Kendall's \(\tau\) alike.
A pivot is a function \(Q(\mathbf{X}, \theta)\) whose distribution does not depend on \(\theta\). Choosing \(a, b\) with \(P(a \le Q \le b) = 1 - \alpha\) and inverting the inequalities to isolate \(\theta\) gives a \(100(1-\alpha)\%\) confidence interval.
| parameter | pivot | its distribution |
|---|---|---|
| \(\mu\), \(\sigma\) known | \(\dfrac{\bar X - \mu}{\sigma/\sqrt n}\) | \(N(0,1)\) |
| \(\mu\), \(\sigma\) unknown | \(\dfrac{\bar X - \mu}{S/\sqrt n}\) | \(t_{n-1}\) |
| \(\sigma^{2}\) | \(\dfrac{(n-1)S^{2}}{\sigma^{2}}\) | \(\chi^{2}_{n-1}\) |
| \(\theta\) (exponential) | \(2n\theta\bar X\) | \(\chi^{2}_{2n}\) |
| \(\sigma_1^{2}/\sigma_2^{2}\) | \(\dfrac{S_1^{2}/\sigma_1^{2}}{S_2^{2}/\sigma_2^{2}}\) | \(F_{n_1-1,\,n_2-1}\) |
The exponential row is worth noting: it is an exact pivot, not a large-sample one, because \(2n\theta\bar X = \sum_{i=1}^{n} 2\theta X_i\) is a sum of \(n\) independent \(\chi^{2}_{2}\) terms (each \(2\theta X_i \sim \chi^{2}_{2}\)).
Given. \(n = 25\) observations from \(N(\mu, \sigma^{2})\), with \(\bar x = 52.4\) and \(s^{2} = 6.25\). Take \(\sigma = 8\) known for part (a).
(a) A \(95\%\) interval for \(\mu\), with \(\sigma\) known.
\[ SE = \frac{\sigma}{\sqrt n} = \frac{8}{5} = 1.6, \qquad 1.96 \times 1.6 = 3.136, \] \[ 52.4 \pm 3.136 = (49.2640,\ 55.5360). \]The multiplier is correct because \(P(|Z| > 1.96) = 0.049996 \approx 0.05\).
(b) A \(95\%\) interval for \(\sigma^{2}\). The pivot is \((n-1)S^{2}/\sigma^{2} \sim \chi^{2}_{24}\), whose \(2.5\%\) points are \(12.401\) and \(39.364\). (Check: their upper-tail areas are \(0.975002\) and \(0.025000\).) Inverting
\[ P\left(12.401 \le \frac{24 s^{2}}{\sigma^{2}} \le 39.364\right) = 0.95 \]gives, with \(24 s^{2} = 24 \times 6.25 = 150\),
\[ \left(\frac{150}{39.364},\ \frac{150}{12.401}\right) = (3.810588,\ 12.095799). \]Interpretation. The interval for \(\sigma^{2}\) is strikingly asymmetric about \(s^{2} = 6.25\): it reaches \(2.44\) below and \(5.85\) above. That is the skewness of the \(\chi^{2}\) distribution, and it is why a variance interval must never be written as an estimate plus or minus something. Its width is \(8.285\), which is \(1.33\) times \(s^{2}\) itself — twenty-five observations pin down a variance only loosely, the same point measured differently in Distribution Theory, Unit 3, Example 3.1.
Infinitely many \((a, b)\) satisfy \(P(a \le Q \le b) = 1 - \alpha\). Which gives the shortest interval?
When the pivot has a symmetric unimodal density — \(N(0,1)\) or \(t_\nu\) — the equal-tailed choice is the shortest, and the usual interval is already optimal. Nothing is gained by looking further.
When the density is skewed — \(\chi^{2}\) or \(F\) — it is not. The shortest interval is obtained by minimising the length subject to the coverage constraint, and the solution is a condition on the density \(f\) at the two endpoints rather than an equal-tail one. Which condition depends on how the interval is built from the pivot. When its length is proportional to \(b - a\) (a location-type interval), the rule is equal-ordinate: \(f(a) = f(b)\). For the variance interval \(\big(24s^{2}/b,\ 24s^{2}/a\big)\), whose length is proportional to \(1/a - 1/b\), the rule is \(a^{2}f(a) = b^{2}f(b)\).
Given. The variance interval of Example 3.2(b), \(n = 25\), \(s^{2} = 6.25\).
Step 1 — the equal-tailed width.
\[ L_{\text{eq}} = 24 s^{2}\left(\frac{1}{12.401} - \frac{1}{39.364}\right) = 12.0958 - 3.8106 = 8.2852. \]Step 2 — search over the split of the \(5\%\). Let the lower tail carry area \(a\) and the upper \(0.05 - a\). Minimising the width over \(a\) gives
\[ a = 0.0433, \qquad \text{points } (13.5230,\ 44.4837), \qquad L_{\min} = 24 s^{2}\left(\frac{1}{13.5230} - \frac{1}{44.4837}\right) = 7.7202. \]Check. At these points \(a^{2}f(a) = 3.58 = b^{2}f(b)\), the condition for a variance interval, while the ordinates themselves differ (\(f(13.523) = 0.0196\), \(f(44.484) = 0.0018\)).
Step 3 — the saving.
\[ \frac{8.2852 - 7.7202}{8.2852} = 0.0682, \quad \text{that is } 6.82\%. \]Interpretation. The optimal split is nowhere near equal: \(4.33\%\) in the lower tail against \(0.67\%\) in the upper, because the \(\chi^{2}\) density falls away slowly on the right. The gain is real but modest, and it costs a non-standard pair of critical values that no table carries. That trade-off — \(6.8\%\) of width against the convenience of a printed table — is why equal-tailed intervals remain standard practice, and the honest statement is that they are conventional rather than optimal.
Let \(\xi_p\) be the \(p\)-th quantile of a continuous distribution, so \(P(X \le \xi_p) = p\). Then each observation independently falls below \(\xi_p\) with probability \(p\), so
\[ \#\{i : X_i \le \xi_p\} \sim \text{Bin}(n, p). \]Consequently, for order statistics \(X_{(r)} < X_{(s)}\),
\[ P\left(X_{(r)} < \xi_p < X_{(s)}\right) = P\big(r \le B \le s - 1\big), \qquad B \sim \text{Bin}(n, p). \]No assumption about the distribution is made anywhere beyond continuity. The coverage is exact and comes from the binomial table, and the interval's endpoints are observed data values.
Given. \(n = 15\) observations from an unknown continuous distribution; \(p = 0.5\).
Step 1 — try \(r = 4\), \(s = 12\). The coverage is
\[ P\big(4 \le B \le 11\big), \qquad B \sim \text{Bin}(15, 0.5), \]which evaluates to \(0.964844\).
Step 2 — try the next pair in, \(r = 5\), \(s = 11\).
\[ P\big(5 \le B \le 10\big) = 0.881531. \]Step 3 — choose. For at least \(95\%\) coverage, take \(\left(X_{(4)},\ X_{(12)}\right)\) with actual confidence \(96.48\%\). The next pair in falls to \(88.15\%\), which is well short.
Interpretation. Only certain confidence levels are attainable, because the binomial is discrete — \(96.48\%\) and \(88.15\%\) are available and nothing between them is. The conservative choice is taken, and the actual level is stated rather than the nominal one. Note what has been bought: an exact interval for the median of any continuous distribution, with no normality assumption, no variance estimate and no asymptotics. The price is width, and the discreteness of the level.
A confidence interval is a statement about a parameter: \(P(L \le \mu \le U) = 1 - \alpha\).
A tolerance interval is a statement about the population. The pair \((L, U)\) is a \(\gamma\)-content, \((1-\alpha)\)-confidence tolerance interval if
\[ P\Big[\,F(U) - F(L) \ge \gamma\,\Big] = 1 - \alpha, \]that is: with confidence \(1-\alpha\), at least a proportion \(\gamma\) of the population lies between \(L\) and \(U\).
The distinction in one comparison. For a normal sample:
| form | behaviour as \(n \to \infty\) | |
|---|---|---|
| confidence interval for \(\mu\) | \(\bar x \pm t\,s/\sqrt n\) | width \(\to 0\) |
| tolerance interval | \(\bar x \pm k\,s\) | width \(\to 2 z_{(1+\gamma)/2}\,\sigma\) |
The confidence interval shrinks to a point, because a parameter is one number and enough data locates it exactly. The tolerance interval does not, because the population genuinely has spread and no amount of data removes it. Confusing the two — quoting a confidence interval for the mean as though it described where individual observations fall — is the commonest misuse of an interval estimate anywhere in applied work.
Distribution-free tolerance limits. The order statistics supply them with no model at all: for a continuous distribution,
\[ P\Big[F\left(X_{(n)}\right) - F\left(X_{(1)}\right) \ge \gamma\Big] = 1 - n\gamma^{n-1} + (n-1)\gamma^{n}, \]a formula depending only on \(n\) and \(\gamma\). The result is a corollary of the distribution of the range in Distribution Theory, Unit 4 — since \(F(X)\) is uniform, \(F(X_{(n)}) - F(X_{(1)})\) is the range of a uniform sample and has a Beta\((n-1, 2)\) distribution.