A real quadratic form in \(n\) variables is
\[ Q(\mathbf{x}) = \mathbf{x}'A\mathbf{x} = \sum_{i=1}^{n}\sum_{j=1}^{n} a_{ij}x_i x_j, \]with \(A\) real and, without loss of generality, symmetric: any \(A\) may be replaced by \(\tfrac12(A + A')\) without changing \(Q\), since \(\mathbf{x}'A\mathbf{x}\) is a scalar and therefore equals its own transpose \(\mathbf{x}'A'\mathbf{x}\).
Reading off the matrix. The coefficient of \(x_i^{2}\) is \(a_{ii}\); the coefficient of \(x_ix_j\) for \(i \ne j\) is \(2a_{ij}\), because the term appears twice in the double sum. So a coefficient of \(4x_1x_2\) means \(a_{12} = a_{21} = 2\) — halving is the step most often forgotten.
By the spectral theorem, \(A = P\Lambda P'\) with \(P\) orthogonal. Substituting \(\mathbf{x} = P\mathbf{y}\), so that \(\mathbf{y} = P'\mathbf{x}\),
\[ Q = \mathbf{x}'A\mathbf{x} = (P\mathbf{y})' P\Lambda P' (P\mathbf{y}) = \mathbf{y}'\left(P'P\right)\Lambda\left(P'P\right)\mathbf{y} = \mathbf{y}'\Lambda\mathbf{y} = \sum_{i=1}^{n} \lambda_i y_i^{2}. \]Every cross-product term is gone. This is the canonical (diagonal) form, and the transformation is a rotation of the coordinate axes onto the eigenvectors.
Let \(p\) be the number of positive \(\lambda_i\) and \(q\) the number of negative ones.
\[ \textbf{rank } r = p + q, \qquad \textbf{index} = p, \qquad \textbf{signature } s = p - q. \]Sylvester's law of inertia: \(p\) and \(q\) do not depend on which non-singular transformation was used to diagonalise the form — only the rank and the signature are intrinsic. Then
| classification | condition on the eigenvalues | on \(Q\) |
|---|---|---|
| positive definite | all \(\lambda_i > 0\) | \(Q > 0\) for all \(\mathbf{x} \ne \mathbf{0}\) |
| positive semi-definite | all \(\lambda_i \ge 0\), at least one \(= 0\) | \(Q \ge 0\), and \(=0\) for some \(\mathbf{x} \ne \mathbf{0}\) |
| negative definite | all \(\lambda_i < 0\) | \(Q < 0\) for all \(\mathbf{x} \ne \mathbf{0}\) |
| indefinite | some positive, some negative | \(Q\) takes both signs |
A test that needs no eigenvalues. \(A\) is positive definite if and only if every leading principal minor is positive — \(a_{11} > 0\), \(\left|\begin{smallmatrix}a_{11} & a_{12}\\ a_{21} & a_{22}\end{smallmatrix}\right| > 0\), and so on up to \(\det(A) > 0\). This is quicker by hand and is the usual way to check that a covariance matrix is admissible.
Given. \(Q = 2x_1^{2} + 2x_2^{2} + 2x_3^{2} + 2x_1x_2 + 2x_1x_3 + 2x_2x_3\).
Asked. Write its matrix, reduce it to canonical form, and give its rank, index, signature and classification.
Step 1 — the matrix. The squared terms give the diagonal \(2, 2, 2\). Each cross-product coefficient is \(2\), so each off-diagonal entry is \(2/2 = 1\):
\[ A = \begin{pmatrix} 2 & 1 & 1 \\ 1 & 2 & 1 \\ 1 & 1 & 2 \end{pmatrix}. \]Step 2 — the eigenvalues. From Unit 2, Example 2.1 these are \(4, 1, 1\).
Step 3 — the canonical form.
\[ Q = 4y_1^{2} + y_2^{2} + y_3^{2}, \]with \(y_1\) the coordinate along \((1,1,1)'/\sqrt3\) and \(y_2, y_3\) along an orthonormal pair in the plane \(x_1 + x_2 + x_3 = 0\).
Step 4 — the three counts. All three eigenvalues are positive, so \(p = 3\), \(q = 0\), giving
\[ \text{rank} = 3, \qquad \text{index} = 3, \qquad \text{signature} = 3 - 0 = 3, \]and the form is positive definite.
Step 5 — confirm by leading minors, which uses no eigenvalues.
\[ 2 > 0, \qquad \begin{vmatrix} 2 & 1 \\ 1 & 2 \end{vmatrix} = 4 - 1 = 3 > 0, \qquad \det(A) = 4 > 0. \checkmark \]Step 6 — a spot check on the form itself. At \(\mathbf{x} = (1,-1,0)'\),
\[ Q = 2 + 2 + 0 + 2(1)(-1) + 0 + 0 = 4 - 2 = 2 > 0, \]and this \(\mathbf{x}\) lies in the \(\lambda = 1\) eigenspace with \(\mathbf{x}'\mathbf{x} = 2\), so the canonical form predicts \(1 \times 2 = 2\). \(\checkmark\)
Given. \(Q = x_1^{2} + 4x_1x_2 + x_2^{2}\).
Step 1 — the matrix. The cross-product coefficient is 4, so the off-diagonal entries are \(4/2 = 2\):
\[ A = \begin{pmatrix} 1 & 2 \\ 2 & 1 \end{pmatrix}. \]Step 2 — the eigenvalues. \(|A - \lambda I| = (1-\lambda)^{2} - 4 = \lambda^{2} - 2\lambda - 3 = (\lambda-3)(\lambda+1)\), so \(\lambda = 3\) and \(\lambda = -1\). Checks: \(3 + (-1) = 2 = \operatorname{trace}(A)\) and \(3 \times (-1) = -3 = \det(A)\). \(\checkmark\)
Step 3 — canonical form and counts. \(Q = 3y_1^{2} - y_2^{2}\), so \(p = 1\), \(q = 1\) and
\[ \text{rank} = 2, \qquad \text{index} = 1, \qquad \text{signature} = 1 - 1 = 0, \]and the form is indefinite.
Step 4 — confirm directly by finding both signs. At \(\mathbf{x} = (1,1)'\), \(Q = 1 + 4 + 1 = 6 > 0\); at \(\mathbf{x} = (1,-1)'\), \(Q = 1 - 4 + 1 = -2 < 0\). Both signs occur, as an indefinite form must produce. \(\checkmark\)
Note that the leading-minor test also detects this at once: \(a_{11} = 1 > 0\) but \(\det(A) = -3 < 0\), so \(A\) is not positive definite.
Statement. For symmetric \(A\) with eigenvalues \(\lambda_{\max} \ge \cdots \ge \lambda_{\min}\),
\[ \lambda_{\min} \;\le\; \frac{\mathbf{x}'A\mathbf{x}}{\mathbf{x}'\mathbf{x}} \;\le\; \lambda_{\max} \qquad \text{for every } \mathbf{x} \ne \mathbf{0}, \]with the bounds attained at the corresponding eigenvectors.
Proof. Put \(\mathbf{y} = P'\mathbf{x}\), so \(\mathbf{x}'\mathbf{x} = \mathbf{y}'\mathbf{y}\) because \(P\) is orthogonal, and \(\mathbf{x}'A\mathbf{x} = \sum_i \lambda_i y_i^{2}\) by the reduction above. Then
\[ \frac{\mathbf{x}'A\mathbf{x}}{\mathbf{x}'\mathbf{x}} = \frac{\sum_i \lambda_i y_i^{2}}{\sum_i y_i^{2}}, \]a weighted average of the \(\lambda_i\) with non-negative weights \(y_i^{2}\) summing to the denominator. An average of numbers lies between the smallest and the largest of them, and equals an endpoint exactly when all the weight sits there — that is, when \(\mathbf{x}\) is the corresponding eigenvector. \(\blacksquare\)
Why this matters. Maximising \(\mathbf{a}'S\mathbf{a}\) subject to \(\mathbf{a}'\mathbf{a} = 1\) is precisely the problem that defines the first principal component. The answer — take the eigenvector of the largest eigenvalue, and the maximum is that eigenvalue — is this theorem, and nothing else.
Given two forms \(\mathbf{x}'A\mathbf{x}\) and \(\mathbf{x}'B\mathbf{x}\) with \(B\) positive definite, there is a single non-singular \(T\) such that \(\mathbf{x} = T\mathbf{y}\) makes
\[ \mathbf{x}'B\mathbf{x} = \mathbf{y}'\mathbf{y} = \sum_i y_i^{2}, \qquad \mathbf{x}'A\mathbf{x} = \sum_i \lambda_i y_i^{2}, \]where the \(\lambda_i\) solve the generalized eigenvalue problem
\[ \left| A - \lambda B \right| = 0. \]Why \(B\) must be positive definite. The construction writes \(B = B^{1/2}B^{1/2}\) — possible only when every eigenvalue of \(B\) is positive — and then diagonalises \(B^{-1/2}AB^{-1/2}\), which is symmetric, by the ordinary spectral theorem.
Where it is used. Fisher's linear discriminant maximises the ratio of between-group to within-group variation, which is exactly \(\max_{\mathbf{a}} \dfrac{\mathbf{a}'A\mathbf{a}}{\mathbf{a}'B\mathbf{a}}\) — the largest generalized eigenvalue. Canonical correlation analysis is the same problem again.
Given.
\[ A = \begin{pmatrix} 5 & 2 \\ 2 & 2 \end{pmatrix}, \qquad B = \begin{pmatrix} 2 & 1 \\ 1 & 1 \end{pmatrix}. \]Step 1 — check \(B\) is positive definite. Leading minors: \(2 > 0\) and \(\det(B) = 2 - 1 = 1 > 0\). \(\checkmark\) (And \(A\) too: \(5 > 0\), \(\det(A) = 10 - 4 = 6 > 0\).)
Step 2 — form the determinant.
\[ A - \lambda B = \begin{pmatrix} 5 - 2\lambda & 2 - \lambda \\ 2 - \lambda & 2 - \lambda \end{pmatrix}, \] \[ \left| A - \lambda B \right| = (5 - 2\lambda)(2 - \lambda) - (2 - \lambda)^{2}. \]Step 3 — factor out the common term instead of expanding. Both terms carry \((2 - \lambda)\):
\[ = (2 - \lambda)\left[(5 - 2\lambda) - (2 - \lambda)\right] = (2 - \lambda)(3 - \lambda). \]Step 4 — the roots. \(\lambda_1 = 2\) and \(\lambda_2 = 3\).
Step 5 — verify each directly. At \(\lambda = 2\),
\[ A - 2B = \begin{pmatrix} 1 & 0 \\ 0 & 0 \end{pmatrix}, \qquad \det = 0. \checkmark \]At \(\lambda = 3\),
\[ A - 3B = \begin{pmatrix} -1 & -1 \\ -1 & -1 \end{pmatrix}, \qquad \det = 1 - 1 = 0. \checkmark \]Step 6 — the simultaneous canonical form.
\[ \mathbf{x}'B\mathbf{x} = y_1^{2} + y_2^{2}, \qquad \mathbf{x}'A\mathbf{x} = 2y_1^{2} + 3y_2^{2}. \]Interpretation. The ratio \(\mathbf{x}'A\mathbf{x} / \mathbf{x}'B\mathbf{x}\) is a weighted average of \(2\) and \(3\), so it ranges over \([2, 3]\) and no further. In a discriminant problem those numbers would be the smallest and largest achievable separation, and the direction attaining \(3\) would be the discriminant function.
Statement. For any \(\mathbf{x}, \mathbf{y} \in \mathbb{R}^{n}\),
\[ \left(\mathbf{x}'\mathbf{y}\right)^{2} \le \left(\mathbf{x}'\mathbf{x}\right)\left(\mathbf{y}'\mathbf{y}\right), \]with equality if and only if \(\mathbf{x}\) and \(\mathbf{y}\) are proportional.
Proof. For every real \(t\), \(\|\mathbf{x} - t\mathbf{y}\|^{2} \ge 0\) because it is a squared length. Expanding,
\[ \mathbf{x}'\mathbf{x} - 2t\,\mathbf{x}'\mathbf{y} + t^{2}\,\mathbf{y}'\mathbf{y} \ge 0. \]This is a quadratic in \(t\) that is never negative, so its discriminant cannot be positive:
\[ \left(-2\,\mathbf{x}'\mathbf{y}\right)^{2} - 4\left(\mathbf{y}'\mathbf{y}\right)\left(\mathbf{x}'\mathbf{x}\right) \le 0, \]which rearranges to the statement. Equality needs the quadratic to have a repeated real root, that is \(\mathbf{x} = t\mathbf{y}\) for some \(t\). \(\blacksquare\)
What it gives statistics. Taking centred data vectors, it says exactly \(-1 \le r \le 1\) for the correlation coefficient, with \(|r| = 1\) only for an exact linear relation — the result quoted without proof in Statistical Methods, Unit 2.
Statement. For any \(n \times n\) real matrix \(A\) with columns \(\mathbf{a}_1, \ldots, \mathbf{a}_n\),
\[ \left|\det A\right| \;\le\; \prod_{j=1}^{n} \|\mathbf{a}_j\| = \prod_{j=1}^{n}\left(\sum_{i=1}^{n} a_{ij}^{2}\right)^{1/2}, \]with equality if and only if the columns are mutually orthogonal (or some column is zero).
The geometry. \(|\det A|\) is the volume of the parallelepiped spanned by the columns. That volume is largest, for given edge lengths, when the edges are perpendicular — a box beats any slanted version of itself.
Worked check on \(A\) of Example 3.1. Each column is a permutation of \((2,1,1)\), so each has length \(\sqrt{4 + 1 + 1} = \sqrt6 = 2.449490\). Hence
\[ \prod_{j=1}^{3}\|\mathbf{a}_j\| = \left(\sqrt6\right)^{3} = 6\sqrt6 = 14.696938, \]and \(|\det A| = 4 \le 14.696938\). \(\checkmark\) The gap is large because the columns are far from orthogonal — any two of them have inner product \(2(1) + 1(2) + 1(1) = 5\), not \(0\).
Where it is used. For a covariance matrix \(\Sigma\), Hadamard gives \(\det\Sigma \le \prod_i \sigma_{ii}\), with equality only when the variables are uncorrelated. The ratio of the two sides is therefore a measure of how much the variables overlap, and it is exactly what the generalized variance of Multivariate Analysis (STS-202) reports.