For a random vector \(\mathbf{Y} = (Y_1, \ldots, Y_n)'\) and a symmetric \(n \times n\) matrix \(A\), the scalar
\[ Q = \mathbf{Y}' A \mathbf{Y} = \sum_{i=1}^{n}\sum_{j=1}^{n} a_{ij} Y_i Y_j \]is a quadratic form in \(\mathbf{Y}\). Every sum of squares in statistics is one: the total, the between-groups and the within-groups sums of squares of analysis of variance each correspond to a particular \(A\).
Statement. Let \(\mathbf{Z} \sim N_n(\mathbf{0}, I)\) and let \(A\) be symmetric. Then
\[ \mathbf{Z}'A\mathbf{Z} \sim \chi^{2}_{r} \iff A \text{ is idempotent } (A^{2} = A), \quad\text{and then } r = \operatorname{rank}(A) = \operatorname{trace}(A). \]Why idempotency is the right condition. A symmetric idempotent matrix is a projection: its eigenvalues are all 0 or 1, because \(A^{2} = A\) forces \(\lambda^{2} = \lambda\). Diagonalising \(A = P\Lambda P'\) with \(P\) orthogonal and setting \(\mathbf{W} = P'\mathbf{Z}\) — again \(N_n(\mathbf{0}, I)\), by the argument of Unit 3, Step 3 — gives
\[ \mathbf{Z}'A\mathbf{Z} = \mathbf{W}'\Lambda\mathbf{W} = \sum_{i=1}^{n} \lambda_i W_i^{2} = \sum_{i \,:\, \lambda_i = 1} W_i^{2}, \]a sum of exactly \(r\) independent squared standard normals, which is \(\chi^{2}_{r}\). And since the eigenvalues are 0 and 1, their number equals both the rank and the trace.
Two companion results. \(\mathbf{Z}'A\mathbf{Z}\) and \(\mathbf{Z}'B\mathbf{Z}\) are independent if and only if \(AB = 0\); and \(\mathbf{Z}'A\mathbf{Z}\) is independent of the linear form \(L\mathbf{Z}\) if and only if \(LA = 0\).
Given. \(\mathbf{Z} \sim N_n(\mathbf{0}, I)\), and \(J\) the \(n \times n\) matrix of ones.
Asked. Show that \(\sum_i (Z_i - \bar Z)^{2}\) is a quadratic form with an idempotent matrix, and recover the \(\chi^{2}_{n-1}\) result of Unit 3 from the theorem.
Step 1 — write the sum of squares in matrix form. Since \(\bar Z = \tfrac{1}{n}\mathbf{1}'\mathbf{Z}\) and \(J = \mathbf{1}\mathbf{1}'\),
\[ \sum_{i=1}^{n}\left(Z_i - \bar Z\right)^{2} = \mathbf{Z}'\mathbf{Z} - n\bar Z^{2} = \mathbf{Z}'\mathbf{Z} - \frac{1}{n}\mathbf{Z}'J\mathbf{Z} = \mathbf{Z}'\left(I - \frac{J}{n}\right)\mathbf{Z}. \]So \(A = I - J/n\), which is symmetric because \(I\) and \(J\) both are.
Step 2 — check idempotency. The key fact is \(J^{2} = \mathbf{1}(\mathbf{1}'\mathbf{1})\mathbf{1}' = n\mathbf{1}\mathbf{1}' = nJ\), since \(\mathbf{1}'\mathbf{1} = n\). Therefore
\[ A^{2} = \left(I - \frac{J}{n}\right)^{2} = I - \frac{2J}{n} + \frac{J^{2}}{n^{2}} = I - \frac{2J}{n} + \frac{nJ}{n^{2}} = I - \frac{2J}{n} + \frac{J}{n} = I - \frac{J}{n} = A. \checkmark \]Step 3 — find the rank from the trace. For an idempotent matrix these are equal, and the trace is easy:
\[ \operatorname{trace}(A) = \operatorname{trace}(I) - \frac{1}{n}\operatorname{trace}(J) = n - \frac{n}{n} = n - 1, \]because \(J\) has \(n\) ones on its diagonal.
Step 4 — apply the theorem. \(\sum_i (Z_i - \bar Z)^{2} \sim \chi^{2}_{n-1}\), and returning to the original scale,
\[ \frac{(n-1)S^{2}}{\sigma^{2}} \sim \chi^{2}_{n-1}. \]Step 5 — independence, for free. The sample mean is the linear form \(L\mathbf{Z}\) with \(L = \tfrac{1}{n}\mathbf{1}'\). Then
\[ LA = \frac{1}{n}\mathbf{1}'\left(I - \frac{J}{n}\right) = \frac{1}{n}\left(\mathbf{1}' - \frac{\mathbf{1}'J}{n}\right) = \frac{1}{n}\left(\mathbf{1}' - \frac{n\mathbf{1}'}{n}\right) = \mathbf{0}, \]using \(\mathbf{1}'J = n\mathbf{1}'\). By the companion result, \(\bar Z\) and \(\sum_i(Z_i - \bar Z)^{2}\) are independent.
Interpretation. Unit 3 obtained both facts through the Helmert transformation, in six steps. Here they fall out of one trace and one matrix product. That economy is the point of the quadratic-form machinery: the same two lines handle the analysis of variance table, where Helmert's explicit construction would be hopeless.
Statement. Let \(\mathbf{Z} \sim N_n(\mathbf{0}, I)\) and suppose
\[ \mathbf{Z}'\mathbf{Z} = Q_1 + Q_2 + \cdots + Q_k, \qquad Q_j = \mathbf{Z}'A_j\mathbf{Z}, \quad \operatorname{rank}(A_j) = r_j. \]Then the following are equivalent:
What it delivers. One arithmetic check — do the degrees of freedom add up to \(n\)? — certifies that every sum of squares in a decomposition is chi-square and that they are mutually independent. That is the entire theoretical basis of the analysis of variance table of Design and Analysis of Experiments, Unit 1, where a one-way layout with \(k\) treatments and \(n\) observations splits as
\[ \underbrace{n-1}_{\text{total}} = \underbrace{(k-1)}_{\text{between}} + \underbrace{(n-k)}_{\text{within}}, \]and it is why the ratio of the two mean squares is an \(F\): independence is what the definition of \(F\) in Unit 3 requires, and Fisher–Cochran is what supplies it.
Let \(X_1, \ldots, X_n\) be i.i.d. with density \(f\) and distribution function \(F\). Arranging them in increasing order gives the order statistics
\[ X_{(1)} \le X_{(2)} \le \cdots \le X_{(n)}, \]with \(X_{(1)} = \min_i X_i\) and \(X_{(n)} = \max_i X_i\). Their joint density is
\[ f_{(1),\ldots,(n)}(y_1, \ldots, y_n) = n! \prod_{i=1}^{n} f(y_i), \qquad y_1 < y_2 < \cdots < y_n. \]Where the \(n!\) comes from. Any one ordered vector could have arisen from any of the \(n!\) permutations of the original sample, each equally likely and each carrying the same density \(\prod_i f(y_i)\). Note that the order statistics are never independent, whatever the parent: the constraint \(y_1 < \cdots < y_n\) ties them together.
Statement.
\[ f_{(r)}(x) = \frac{n!}{(r-1)!\,(n-r)!} \left[F(x)\right]^{r-1} f(x) \left[1 - F(x)\right]^{n-r}. \]Derivation, by counting. For \(X_{(r)}\) to lie in the small interval \((x, x + dx)\), three things must happen at once:
The number of ways to allocate the \(n\) labelled observations into these three groups is the multinomial coefficient \(\dfrac{n!}{(r-1)!\,1!\,(n-r)!}\). Multiplying gives the density. \(\blacksquare\)
The two ends, directly. Rather than substituting \(r = 1\) and \(r = n\), it is quicker to argue from the distribution function:
\[ F_{(n)}(x) = P\left(\text{all } X_i \le x\right) = \left[F(x)\right]^{n} \;\Rightarrow\; f_{(n)}(x) = n\left[F(x)\right]^{n-1} f(x), \] \[ F_{(1)}(x) = 1 - P\left(\text{all } X_i > x\right) = 1 - \left[1 - F(x)\right]^{n} \;\Rightarrow\; f_{(1)}(x) = n\left[1 - F(x)\right]^{n-1} f(x), \]each using independence to turn "all of them" into an \(n\)-th power.
Given. \(X_1, \ldots, X_5\) i.i.d. \(U(0,1)\), so \(f(x) = 1\) and \(F(x) = x\) on \((0,1)\).
Asked. Find the density of \(X_{(r)}\), its mean and variance, and \(P\left(X_{(5)} > 0.9\right)\).
Step 1 — substitute into the general formula. With \(n = 5\), \(F(x) = x\), \(f(x) = 1\),
\[ f_{(r)}(x) = \frac{5!}{(r-1)!\,(5-r)!} x^{r-1}(1-x)^{5-r}, \qquad 0 < x < 1. \]Step 2 — recognise the Beta. This is the Beta\((r, 6-r)\) density, since the Beta\((p,q)\) density is \(x^{p-1}(1-x)^{q-1}/B(p,q)\) and \(1/B(r, 6-r) = \dfrac{\Gamma(6)}{\Gamma(r)\Gamma(6-r)} = \dfrac{5!}{(r-1)!(5-r)!}\). So
\[ X_{(r)} \sim \text{Beta}(r,\; n - r + 1). \]Step 3 — mean and variance from the Beta formulae. For Beta\((p,q)\), the mean is \(p/(p+q)\) and the variance \(pq/[(p+q)^{2}(p+q+1)]\). Here \(p + q = 6\), so
\[ E\left(X_{(r)}\right) = \frac{r}{6}, \qquad \operatorname{Var}\left(X_{(r)}\right) = \frac{r(6-r)}{36 \times 7} = \frac{r(6-r)}{252}. \]| \(r\) | distribution | \(E\left(X_{(r)}\right)\) | \(\operatorname{Var}\left(X_{(r)}\right)\) |
|---|---|---|---|
| 1 (min) | Beta(1,5) | 0.166667 | 0.019841 |
| 2 | Beta(2,4) | 0.333333 | 0.031746 |
| 3 (median) | Beta(3,3) | 0.500000 | 0.035714 |
| 4 | Beta(4,2) | 0.666667 | 0.031746 |
| 5 (max) | Beta(5,1) | 0.833333 | 0.019841 |
Step 4 — a spot check on the density. At \(r = 3\), \(x = 0.5\),
\[ f_{(3)}(0.5) = \frac{5!}{2!\,2!}(0.5)^{2}(0.5)^{2} = 30 \times 0.25 \times 0.25 = 1.875. \]Step 5 — the tail probability for the maximum.
\[ P\left(X_{(5)} > 0.9\right) = 1 - P\left(\text{all five} \le 0.9\right) = 1 - (0.9)^{5} = 1 - 0.590490 = 0.409510. \]Interpretation. The five means divide \((0,1)\) into six equal gaps of \(1/6\) each — the order statistics of a uniform sample space themselves out evenly, which is exactly why the plotting positions \(r/(n+1)\) are used on a probability plot. The variances are largest in the middle and smallest at the ends, which is the opposite of what intuition usually suggests: the extremes are pinned against the boundaries \(0\) and \(1\), while the median is free to move.
For \(r < s\), the joint density of \(X_{(r)}\) and \(X_{(s)}\) is, by the same counting argument with four groups instead of three,
\[ f_{(r),(s)}(u,v) = \frac{n!}{(r-1)!\,(s-r-1)!\,(n-s)!} \left[F(u)\right]^{r-1} f(u) \left[F(v) - F(u)\right]^{s-r-1} f(v)\left[1-F(v)\right]^{n-s}, \]for \(u < v\). Taking \(r = 1\) and \(s = n\) leaves only the middle group:
\[ f_{(1),(n)}(u,v) = n(n-1)\left[F(v) - F(u)\right]^{n-2} f(u) f(v), \qquad u < v. \]The range is \(R = X_{(n)} - X_{(1)}\). Substituting \(v = u + r\) and integrating out \(u\),
\[ f_R(r) = n(n-1) \int_{-\infty}^{\infty} \left[F(u + r) - F(u)\right]^{n-2} f(u) f(u+r) \, du, \qquad r > 0. \]Given. \(X_1, \ldots, X_n\) i.i.d. \(U(0,1)\).
Asked. Find the density of \(R = X_{(n)} - X_{(1)}\), and evaluate its mean and variance at \(n = 5\).
Step 1 — substitute the uniform \(F\) and \(f\). With \(F(u) = u\) and \(f(u) = 1\) on \((0,1)\), the integrand's bracket becomes \((u + r) - u = r\), which does not involve \(u\) at all:
\[ f_R(r) = n(n-1) \int_{0}^{1-r} r^{\,n-2} \, du. \]Step 2 — fix the limits. Both \(u\) and \(u + r\) must lie in \((0,1)\), so \(0 < u < 1 - r\). The integral is therefore just the length \(1 - r\):
\[ f_R(r) = n(n-1)\, r^{\,n-2} (1 - r), \qquad 0 < r < 1. \]Step 3 — recognise it. This is the Beta\((n-1, 2)\) density, since \(1/B(n-1,2) = \dfrac{\Gamma(n+1)}{\Gamma(n-1)\Gamma(2)} = n(n-1)\).
Step 4 — mean and variance from the Beta formulae. With \(p = n-1\), \(q = 2\), \(p + q = n+1\),
\[ E(R) = \frac{n-1}{n+1}, \qquad \operatorname{Var}(R) = \frac{2(n-1)}{(n+1)^{2}(n+2)}. \]Step 5 — evaluate at \(n = 5\).
\[ E(R) = \frac{4}{6} = 0.666667, \qquad \operatorname{Var}(R) = \frac{2 \times 4}{36 \times 7} = \frac{8}{252} = 0.031746. \]Step 6 — a consistency check. Since \(R = X_{(5)} - X_{(1)}\), its mean must be \(E(X_{(5)}) - E(X_{(1)})\):
\[ 0.833333 - 0.166667 = 0.666667. \checkmark \]Interpretation. Five observations from \(U(0,1)\) cover on average only two thirds of the interval, not the whole of it. As \(n\) grows, \((n-1)/(n+1) \to 1\), but slowly — at \(n = 20\) the expected coverage is still only \(19/21 = 0.905\). This is why the range is a poor estimator of spread for small samples and why control charts tabulate the correction factor \(d_2 = E(R)/\sigma\) separately for each subgroup size, as in Statistical Quality Control, Unit 2.
Given. Five components with independent Exponential\((\theta = 0.1)\) lifetimes, in a system that fails when the first component fails.
Asked. Find the distribution of system life, its mean, and \(P(\text{system survives past } 4)\). Compare with the expected life of the longest-lived component.
Step 1 — the survival function of the minimum. The system survives past \(x\) only if every component does, and the components are independent, so
\[ P\left(X_{(1)} > x\right) = \prod_{i=1}^{5} P(X_i > x) = \left[e^{-0.1x}\right]^{5} = e^{-0.5x}. \]Step 2 — identify the distribution. \(e^{-0.5x}\) is the survival function of an Exponential\((0.5)\). So
\[ X_{(1)} \sim \text{Exponential}(n\theta) = \text{Exponential}(0.5), \]and in general the minimum of \(n\) i.i.d. exponentials is exponential with the rates added.
Step 3 — the mean.
\[ E\left(X_{(1)}\right) = \frac{1}{n\theta} = \frac{1}{0.5} = 2. \]Each component alone has mean life \(1/0.1 = 10\); putting five in series cuts the expected life to \(2\), a fifth.
Step 4 — the survival probability.
\[ P\left(X_{(1)} > 4\right) = e^{-0.5 \times 4} = e^{-2} = 0.135335. \]Step 5 — the maximum, for contrast. For exponentials the maximum has mean
\[ E\left(X_{(n)}\right) = \frac{1}{\theta}\sum_{i=1}^{n}\frac{1}{i} = 10 \times \left(1 + \tfrac12 + \tfrac13 + \tfrac14 + \tfrac15\right) = 10 \times 2.283333 = 22.833333. \]Interpretation. The gap is the whole story of system design. Five components in series last \(2\) on average; the same five in parallel, where the system survives until the last one fails, last \(22.83\) — more than eleven times longer, from identical parts differently arranged. Note also that the maximum grows only like \(\ln n\), since the harmonic sum does, so redundancy has sharply diminishing returns.