Suppose \(p = 10\) variables are each tested at the \(5\%\) level. Two things go wrong, and they go wrong in opposite directions.
(a) Too many false positives. If the null hypothesis is true and the variables were independent, the probability that at least one of the ten tests rejects is \(1 - (0.95)^{10} = 0.401263\). Four times in ten a completely null data set would produce a "significant" variable. Correlation between the variables changes the number but not the problem.
(b) Too few true positives. A real difference may be spread thinly across all ten variables, too small to reach significance on any one of them, yet unmistakable when they are combined. Ten separate tests cannot see it at all.
Hotelling's \(T^{2}\) answers both at once, by testing the mean vector as a single object. Its construction is the natural one: form the discrepancy vector, and measure its length in the metric set by the estimated dispersion matrix.
For a sample of \(n\) observations from \(N_p(\boldsymbol\mu, \boldsymbol\Sigma)\) with \(\boldsymbol\Sigma\) unknown, test \(H_0: \boldsymbol\mu = \boldsymbol\mu_0\) with
\[ T^{2} = n\,(\bar{\mathbf{x}} - \boldsymbol\mu_0)'\,\mathbf{S}^{-1}\, (\bar{\mathbf{x}} - \boldsymbol\mu_0), \qquad \mathbf{S} = \frac{\mathbf{A}}{n-1}. \]Under \(H_0\),
\[ \frac{n-p}{(n-1)p}\;T^{2} \;\sim\; F_{p,\;n-p}. \]Where the \(F\) comes from. By Unit 1, \(\bar{\mathbf{X}} \sim N_p(\boldsymbol\mu_0, \boldsymbol\Sigma/n)\) under \(H_0\), and \(\mathbf{A} \sim W_p(n-1, \boldsymbol\Sigma)\) independently of it. So \(T^{2}\) is of the form
\[ T^{2} = m\,\mathbf{Z}'\mathbf{A}^{-1}\mathbf{Z}, \qquad \mathbf{Z}\sim N_p(\mathbf{0}, \boldsymbol\Sigma) \text{ independent of } \mathbf{A}\sim W_p(m, \boldsymbol\Sigma), \]with \(m = n-1\) — a normal quadratic form divided by an independent Wishart. That ratio is the multivariate counterpart of \(Z^{2}/(\chi^{2}_m/m)\), which is \(t^{2}\), and it reduces to a scaled \(F\). At \(p = 1\) the constant is \((n-1)/(n-1) = 1\) and \(T^{2} = t^{2} \sim F_{1, n-1}\), so nothing has been lost.
Why not \(\chi^{2}\)? If \(\boldsymbol\Sigma\) were known, the statistic \(n(\bar{\mathbf{x}} - \boldsymbol\mu_0)'\boldsymbol\Sigma^{-1} (\bar{\mathbf{x}} - \boldsymbol\mu_0)\) would be exactly \(\chi^{2}_{p}\). Replacing \(\boldsymbol\Sigma\) by \(\mathbf{S}\) inflates the statistic, and the \(F\) distribution is what accounts for that inflation — the same role the \(t\) plays against the \(Z\).
Invariance. \(T^{2}\) is unchanged by any non-singular linear transformation \(\mathbf{x} \mapsto \mathbf{Cx} + \mathbf{d}\). Substituting \(\bar{\mathbf{x}} \mapsto \mathbf{C}\bar{\mathbf{x}} + \mathbf{d}\) and \(\mathbf{S} \mapsto \mathbf{CSC}'\) gives
\[ n\,(\mathbf{C}\mathbf{d}^{*})'(\mathbf{CSC}')^{-1}(\mathbf{C}\mathbf{d}^{*}) = n\,\mathbf{d}^{*\prime}\mathbf{C}'\mathbf{C}'^{-1}\mathbf{S}^{-1} \mathbf{C}^{-1}\mathbf{C}\mathbf{d}^{*} = n\,\mathbf{d}^{*\prime}\mathbf{S}^{-1}\mathbf{d}^{*}, \]where \(\mathbf{d}^{*} = \bar{\mathbf{x}} - \boldsymbol\mu_0\). So the test does not care what units the variables are measured in, nor whether they are first standardised. A test built from \(p\) separate \(t\) statistics has no such property.
It is the likelihood ratio test. Maximising the multivariate normal likelihood under \(H_0\) and without restriction and forming the ratio \(\Lambda = L(\hat\Omega_0)/L(\hat\Omega)\) gives \(\Lambda^{2/n} = \left(1 + T^{2}/(n-1)\right)^{-1}\), a strictly decreasing function of \(T^{2}\). Rejecting for small \(\Lambda\) is therefore identical to rejecting for large \(T^{2}\): the statistic is not an ad hoc construction but the likelihood ratio in disguise.
Given. \(n = 10\) observations on \(p = 2\) variables, with
\[ \bar{\mathbf{x}} = \begin{pmatrix}7\\12\end{pmatrix}, \qquad \mathbf{S} = \begin{pmatrix}4 & 3\\ 3 & 9\end{pmatrix}. \]Asked. Test \(H_0: \boldsymbol\mu = (5, 9)'\) at \(5\%\).
Step 1 — the discrepancy vector.
\[ \mathbf{d} = \bar{\mathbf{x}} - \boldsymbol\mu_0 = \begin{pmatrix}7 - 5\\ 12 - 9\end{pmatrix} = \begin{pmatrix}2\\3\end{pmatrix}. \]Step 2 — invert \(\mathbf{S}\). \(|\mathbf{S}| = 4(9) - 3(3) = 36 - 9 = 27\), so
\[ \mathbf{S}^{-1} = \frac{1}{27}\begin{pmatrix}9 & -3\\ -3 & 4\end{pmatrix}. \]Step 3 — the quadratic form. Multiplying out,
\[ \mathbf{d}'\mathbf{S}^{-1}\mathbf{d} = \frac{1}{27}\left[9(2)^{2} - 2(3)(2)(3) + 4(3)^{2}\right] = \frac{36 - 36 + 36}{27} = \frac{36}{27} = \frac{4}{3} = 1.333333. \]Step 4 — the statistic.
\[ T^{2} = n\,\mathbf{d}'\mathbf{S}^{-1}\mathbf{d} = 10 \times \frac{4}{3} = \frac{40}{3} = 13.333333. \]Step 5 — convert to \(F\). With \(p = 2\) and \(n = 10\),
\[ F = \frac{n-p}{(n-1)p}\,T^{2} = \frac{8}{18} \times \frac{40}{3} = \frac{320}{54} = \frac{160}{27} = 5.925926, \]on \((2, 8)\) degrees of freedom.
Step 6 — decide. The upper \(5\%\) point of \(F_{2,8}\) is \(4.4590\), and \(5.925926 > 4.4590\), so \(H_0\) is rejected; the exact \(p\)-value is \(0.026373\).
Step 7 — look at what the separate tests would have said. Variable by variable,
\[ t_1 = \frac{2}{\sqrt{4/10}} = \frac{2}{0.632456} = 3.1623, \qquad t_2 = \frac{3}{\sqrt{9/10}} = \frac{3}{0.948683} = 3.1623, \]each on 9 degrees of freedom, each significant. Here the two approaches agree.
Interpretation. Notice the arithmetic in Step 3: the cross-product term \(-36\) exactly cancelled the first term, so the whole statistic rests on the third. That cancellation is the covariance at work. Because \(x_1\) and \(x_2\) are positively correlated, a joint shift in the same direction is ordinary and counts for little; a shift in opposite directions would have been extraordinary. \(T^{2}\) measures distance in the metric of \(\mathbf{S}^{-1}\), which is exactly what "unusual, given how these variables normally move together" means.
The Mahalanobis distance between two mean vectors, in the metric of a common dispersion matrix, is
\[ D^{2} = (\boldsymbol\mu_1 - \boldsymbol\mu_2)'\boldsymbol\Sigma^{-1} (\boldsymbol\mu_1 - \boldsymbol\mu_2), \]estimated by \(D^{2} = (\bar{\mathbf{x}}_1 - \bar{\mathbf{x}}_2)'\mathbf{S}_p^{-1} (\bar{\mathbf{x}}_1 - \bar{\mathbf{x}}_2)\) with the pooled matrix
\[ \mathbf{S}_p = \frac{\mathbf{A}_1 + \mathbf{A}_2}{n_1 + n_2 - 2}, \]which is legitimate precisely because of the Wishart additivity property (ii) of Unit 2: \(\mathbf{A}_1 + \mathbf{A}_2 \sim W_p(n_1 + n_2 - 2, \boldsymbol\Sigma)\) when the two populations share \(\boldsymbol\Sigma\). That shared \(\boldsymbol\Sigma\) is an assumption, and it is the assumption this whole section stands on.
Three things make \(D^{2}\) the right notion of distance between populations, and ordinary Euclidean distance the wrong one:
The two-sample statistic and its null distribution are
\[ T^{2} = \frac{n_1n_2}{n_1+n_2}\,D^{2}, \qquad \frac{n_1+n_2-p-1}{(n_1+n_2-2)p}\,T^{2} \;\sim\; F_{p,\;n_1+n_2-p-1}. \]Given. \(n_1 = 12\), \(n_2 = 15\), \(p = 2\), with
\[ \bar{\mathbf{x}}_1 = \begin{pmatrix}10\\14\end{pmatrix}, \qquad \bar{\mathbf{x}}_2 = \begin{pmatrix}8\\11\end{pmatrix}, \qquad \mathbf{S}_p = \begin{pmatrix}4 & 2\\ 2 & 5\end{pmatrix}. \]Step 1 — the difference vector.
\[ \mathbf{d} = \bar{\mathbf{x}}_1 - \bar{\mathbf{x}}_2 = \begin{pmatrix}2\\3\end{pmatrix}. \]Step 2 — invert the pooled matrix. \(|\mathbf{S}_p| = 20 - 4 = 16\), so
\[ \mathbf{S}_p^{-1} = \frac{1}{16}\begin{pmatrix}5 & -2\\ -2 & 4\end{pmatrix}. \]Step 3 — Mahalanobis \(D^{2}\).
\[ D^{2} = \frac{1}{16}\left[5(2)^{2} - 2(2)(2)(3) + 4(3)^{2}\right] = \frac{20 - 24 + 36}{16} = \frac{32}{16} = 2, \qquad D = \sqrt{2} = 1.414214. \]Step 4 — the multiplier.
\[ \frac{n_1n_2}{n_1+n_2} = \frac{12 \times 15}{27} = \frac{180}{27} = \frac{20}{3} = 6.666667. \]Step 5 — the statistic.
\[ T^{2} = \frac{20}{3} \times 2 = \frac{40}{3} = 13.333333. \]Step 6 — convert to \(F\). With \(n_1 + n_2 - p - 1 = 27 - 3 = 24\) and \(n_1+n_2-2 = 25\),
\[ F = \frac{24}{25 \times 2} \times \frac{40}{3} = 0.48 \times 13.333333 = 6.4, \]on \((2, 24)\) degrees of freedom, with \(p\)-value \(0.005921\). The mean vectors differ.
Interpretation. \(D = 1.414\) says the two population centres are about one and a half "standard multivariate units" apart. That is a statement about the populations and does not change with sample size. \(T^{2} = 13.333\) is a statement about the evidence, and it grows with \(n_1\) and \(n_2\) through the multiplier in Step 4. Keeping those two apart is the single most useful habit in this unit: a large \(T^{2}\) with a small \(D^{2}\) means a reliably detected but practically trivial difference.
Sometimes the question is not whether \(\boldsymbol\mu\) equals a given vector, but whether its own components are equal — for instance whether \(p\) treatments applied to the same subjects give the same mean response. Write that hypothesis as
\[ H_0: \mathbf{C}\boldsymbol\mu = \mathbf{0}, \qquad \mathbf{C} = \begin{pmatrix} 1 & -1 & 0 & \cdots & 0\\ 0 & 1 & -1 & \cdots & 0\\ \vdots & & & \ddots & \vdots\\ 0 & \cdots & 0 & 1 & -1\end{pmatrix}, \]a \((p-1) \times p\) contrast matrix of rank \(q = p - 1\). Applying \(\mathbf{C}\) to every observation turns the problem into a one-sample \(T^{2}\) on the transformed data, which by the transformation rules of Unit 1 has mean \(\mathbf{C}\boldsymbol\mu\) and dispersion \(\mathbf{C}\boldsymbol\Sigma\mathbf{C}'\):
\[ T^{2} = n\,(\mathbf{C}\bar{\mathbf{x}})'(\mathbf{CSC}')^{-1}(\mathbf{C}\bar{\mathbf{x}}), \qquad \frac{n-q}{(n-1)q}\,T^{2} \sim F_{q,\;n-q}. \]Any contrast matrix of rank \(p-1\) gives the same \(T^{2}\), by the invariance property of section 2 — so the arbitrary-looking choice of \(\mathbf{C}\) costs nothing.
Given. The sample of Example 3.1: \(n = 10\), \(\bar{\mathbf{x}} = (7, 12)'\), \(\mathbf{S} = \begin{pmatrix}4 & 3\\3 & 9\end{pmatrix}\). Asked. Test \(H_0: \mu_1 = \mu_2\).
Step 1 — the contrast. With \(p = 2\), \(\mathbf{C} = (1, -1)\) and \(q = 1\).
Step 2 — transform the mean. \(\mathbf{C}\bar{\mathbf{x}} = 7 - 12 = -5\).
Step 3 — transform the dispersion.
\[ \mathbf{CSC}' = \begin{pmatrix}1 & -1\end{pmatrix} \begin{pmatrix}4 & 3\\3 & 9\end{pmatrix}\begin{pmatrix}1\\-1\end{pmatrix} = 4 - 3 - 3 + 9 = 7. \]This is just \(\operatorname{Var}(X_1 - X_2) = s_{11} + s_{22} - 2s_{12}\), which is why a positive covariance helps a paired comparison.
Step 4 — the statistic.
\[ T^{2} = 10 \times \frac{(-5)^{2}}{7} = \frac{250}{7} = 35.714286. \]Step 5 — convert. With \(q = 1\) and \(n = 10\) the constant is \((n-q)/[(n-1)q] = 9/9 = 1\), so \(F = 35.714286\) on \((1, 9)\) degrees of freedom, \(p = 0.000209\).
Step 6 — recognise it. A one-degree-of-freedom \(F\) is a squared \(t\), and here the two statistics are the same expression. The paired \(t\) on the differences \(x_{i1} - x_{i2}\) is \(t = \mathbf{C}\bar{\mathbf{x}}/\sqrt{\mathbf{CSC}'/n}\), so
\[ t^{2} = \frac{(\mathbf{C}\bar{\mathbf{x}})^{2}}{\mathbf{CSC}'/n} = n\,\frac{(\mathbf{C}\bar{\mathbf{x}})^{2}}{\mathbf{CSC}'} = T^{2} = \frac{250}{7} \quad\text{identically.} \]Numerically, \(t = -5/\sqrt{0.7} = -5/0.836660 = -5.9761\), and squaring recovers \(250/7\).
Interpretation. At \(p = 2\) the contrast \(T^{2}\) is the paired \(t\) test of Inferential Statistics, Unit 4, squared. At \(p = 5\) it is a single test of all four contrasts at once, which four separate paired \(t\) tests could not provide. That is the whole reason the contrast form is worth learning.
With \(g\) groups rather than two, the total sum-of-squares-and-products matrix splits exactly as an analysis of variance table does:
\[ \mathbf{T} = \mathbf{W} + \mathbf{B}, \]where \(\mathbf{W}\) is the within-groups matrix (a sum of \(g\) Wisharts, pooled by property (ii) of Unit 2) and \(\mathbf{B}\) the between-groups matrix. Wilks' lambda is the ratio of generalized variances
\[ \Lambda = \frac{|\mathbf{W}|}{|\mathbf{W} + \mathbf{B}|}, \qquad 0 < \Lambda \le 1, \]and small values are evidence against equal mean vectors. It is the likelihood ratio statistic, and at \(p = 1\) it reduces to \(\Lambda = SS_{\text{within}}/SS_{\text{total}} = 1 - R^{2}\), so the familiar ANOVA \(F\) is a monotone function of it.
Its distribution. Under \(H_0\), \(\Lambda\) is distributed as a product of independent beta variables. Exact \(F\) transformations exist for small \(p\) or small \(g\): for \(g = 2\) groups,
\[ \frac{1 - \Lambda}{\Lambda}\cdot\frac{n - p - 1}{p} \sim F_{p,\;n-p-1}, \qquad n = n_1 + n_2. \]For the general case Bartlett's approximation is used:
\[ -\left[n - 1 - \frac{p + g}{2}\right]\ln\Lambda \;\dot\sim\; \chi^{2}_{p(g-1)}. \]Given. The two samples of Example 3.2, for which \(T^{2} = 40/3\), \(n = 27\), \(p = 2\), \(g = 2\).
Step 1 — get \(\Lambda\) from \(T^{2}\). For two groups the two statistics are linked exactly:
\[ \Lambda = \frac{1}{1 + \dfrac{T^{2}}{n_1 + n_2 - 2}} = \frac{1}{1 + \dfrac{40/3}{25}} = \frac{1}{1 + \dfrac{8}{15}} = \frac{15}{23} = 0.652174. \]Step 2 — the exact \(F\). With \(n - p - 1 = 24\),
\[ F = \frac{1 - \Lambda}{\Lambda}\cdot\frac{24}{2} = \frac{8/23}{15/23} \times 12 = \frac{8}{15} \times 12 = \frac{96}{15} = 6.400000. \]This is the same \(6.4\) obtained from \(T^{2}\) in Example 3.2, as it must be — the two statistics are one-to-one functions of each other when \(g = 2\).
Step 3 — Bartlett's approximation, for comparison.
\[ -\left[27 - 1 - \frac{2+2}{2}\right]\ln(0.652174) = -24 \times (-0.427444) = 10.2587, \]on \(p(g-1) = 2\) degrees of freedom, giving \(p\)-value \(0.0059\) — against \(0.0059\) from the exact \(F\).
Interpretation. The approximation is excellent here, but that is not guaranteed: it is a large-sample result, and with \(g = 2\) the exact \(F\) is available and should be used. Bartlett's form earns its place when \(g > 2\) and \(p > 2\), where no exact \(F\) exists. Note also what \(\Lambda = 0.652\) means directly: the within-group generalized variance is \(65\%\) of the total, so \(35\%\) of the multivariate spread is accounted for by the group difference.
\(T^{2}\) asks whether two groups differ. Discriminant analysis asks the follow-up question: given a new observation, which group is it from? Fisher's idea was to reduce the \(p\) variables to one number, chosen so that the two groups are as far apart as possible on that one scale.
Consider \(y = \mathbf{a}'\mathbf{x}\). The two group means become \(\mathbf{a}'\bar{\mathbf{x}}_1\) and \(\mathbf{a}'\bar{\mathbf{x}}_2\), and the common variance becomes \(\mathbf{a}'\mathbf{S}_p\mathbf{a}\). Maximise the separation
\[ \frac{\left[\mathbf{a}'(\bar{\mathbf{x}}_1 - \bar{\mathbf{x}}_2)\right]^{2}} {\mathbf{a}'\mathbf{S}_p\mathbf{a}}. \]This is a Rayleigh quotient, exactly the object maximised in Linear Algebra, Unit 3, and by the Cauchy–Schwarz inequality proved there its maximum is attained at
\[ \mathbf{a} = \mathbf{S}_p^{-1}(\bar{\mathbf{x}}_1 - \bar{\mathbf{x}}_2), \]with maximum value \(D^{2}\). So the best achievable separation on any single scale is exactly the Mahalanobis distance — which is why \(D^{2}\) and not some other distance governs how well the groups can be told apart.
The classification rule. Allocate \(\mathbf{x}_0\) to population 1 if
\[ y_0 = \mathbf{a}'\mathbf{x}_0 \;\ge\; m = \tfrac12\left(\bar y_1 + \bar y_2\right) = \tfrac12\,\mathbf{a}'(\bar{\mathbf{x}}_1 + \bar{\mathbf{x}}_2), \]and to population 2 otherwise. Because \(y\) is linear in \(\mathbf{x}\), the boundary \(y = m\) is a straight line (a hyperplane when \(p > 2\)). It is straight because the two dispersion matrices were assumed equal; if they differ, the quadratic terms no longer cancel and the boundary is a conic — quadratic discriminant analysis.
How often it is wrong. Under normality, equal dispersion matrices and equal prior probabilities, the probability of misclassification in either direction is
\[ P(\text{error}) = \Phi\!\left(-\frac{D}{2}\right). \]Why. For an observation truly from population 1, \(y\) is normal with mean \(\bar y_1\) and variance \(\mathbf{a}'\mathbf{S}_p\mathbf{a} = D^{2}\); the midpoint \(m\) lies \(D^{2}/2\) below that mean, which is \(D/2\) standard deviations. The error is \(P(y < m) = \Phi(-D/2)\), and the same holds by symmetry for population 2.
Unequal priors and unequal costs. If the prior probabilities are \(\pi_1\) and \(\pi_2\) and the costs of the two errors are \(c(2\mid1)\) and \(c(1\mid2)\), the Bayes rule shifts the cut-off to
\[ y_0 \;\ge\; m + \ln\!\left(\frac{\pi_2\,c(1\mid2)}{\pi_1\,c(2\mid1)}\right), \]which is the decision-theoretic rule of Estimation Theory, Unit 4 applied here: minimise expected loss, not error count.
Given. The two groups of Example 3.2. Asked. The discriminant function, the cut-off, the error rate, and the classification of \(\mathbf{x}_0 = (9, 12)'\).
Step 1 — the coefficient vector.
\[ \mathbf{a} = \mathbf{S}_p^{-1}\mathbf{d} = \frac{1}{16}\begin{pmatrix}5 & -2\\ -2 & 4\end{pmatrix}\begin{pmatrix}2\\3\end{pmatrix} = \frac{1}{16}\begin{pmatrix}10 - 6\\ -4 + 12\end{pmatrix} = \frac{1}{16}\begin{pmatrix}4\\8\end{pmatrix} = \begin{pmatrix}0.25\\0.50\end{pmatrix}. \]So \(y = 0.25\,x_1 + 0.50\,x_2\).
Step 2 — the two group scores.
\[ \bar y_1 = 0.25(10) + 0.50(14) = 2.5 + 7.0 = 9.5, \qquad \bar y_2 = 0.25(8) + 0.50(11) = 2.0 + 5.5 = 7.5. \]Step 3 — a check. The gap \(\bar y_1 - \bar y_2 = 2.0\) must equal \(D^{2}\), because \(\mathbf{a}'\mathbf{d} = \mathbf{d}'\mathbf{S}_p^{-1}\mathbf{d} = D^{2}\). Example 3.2 gave \(D^{2} = 2\). \(\checkmark\)
Step 4 — the cut-off.
\[ m = \tfrac12(9.5 + 7.5) = 8.5. \]Step 5 — the error rate.
\[ P(\text{error}) = \Phi\!\left(-\frac{\sqrt 2}{2}\right) = \Phi(-0.707107) = 0.239750. \]Step 6 — classify \(\mathbf{x}_0 = (9, 12)'\).
\[ y_0 = 0.25(9) + 0.50(12) = 2.25 + 6.00 = 8.25 < 8.5, \]so allocate it to population 2.
Interpretation. Two things deserve attention. First, the error rate is \(24\%\) — the mean vectors differ beyond reasonable doubt (Example 3.2 gave \(p = 0.0059\)) and yet nearly a quarter of individuals will still be misclassified. A significant difference between group means is not the same as separation between group members, and \(D^{2} = 2\) is simply too small for reliable individual classification. Second, \(\mathbf{x}_0\) scored \(8.25\) against a cut-off of \(8.5\): it is assigned to population 2, but barely, and honest practice is to report the score rather than only the verdict. Raising \(D^{2}\) is the only way to improve matters — from the table below, \(D^{2} = 4\) would be needed to bring the error rate below \(16\%\).
| \(D^{2}\) | \(D\) | \(\Phi(-D/2)\) |
|---|---|---|
| 1 | 1.000000 | 0.308538 |
| 2 | 1.414214 | 0.239750 |
| 3 | 1.732051 | 0.193238 |
| 4 | 2.000000 | 0.158655 |