For a square matrix \(A\) of order \(n\), a scalar \(\lambda\) and a non-zero vector \(\mathbf{x}\) satisfying
\[ A\mathbf{x} = \lambda\mathbf{x} \]are a characteristic root (eigenvalue) and a corresponding characteristic vector (eigenvector). Rewriting, \((A - \lambda I)\mathbf{x} = \mathbf{0}\), and a non-zero solution exists exactly when \(A - \lambda I\) is singular. So the roots are the solutions of the characteristic equation
\[ \left| A - \lambda I \right| = 0, \]a polynomial of degree \(n\) in \(\lambda\).
The idempotent line is the one that matters most for statistics: it is exactly the criterion used in Distribution Theory, Unit 4 to decide when a quadratic form is chi-square.
Expanding \(|A - \lambda I|\) by cofactors is laborious for \(n = 3\) and worse beyond. The Leverrier–Faddeev identities give the coefficients from traces alone. Writing the characteristic polynomial as
\[ \lambda^{n} - c_1 \lambda^{n-1} + c_2 \lambda^{n-2} - \cdots \pm c_n = 0, \]and \(t_k = \operatorname{trace}(A^{k})\),
\[ c_1 = t_1, \qquad c_2 = \frac{c_1 t_1 - t_2}{2}, \qquad c_3 = \frac{c_2 t_1 - c_1 t_2 + t_3}{3}, \]and in general \(c_k = \dfrac{1}{k}\left(c_{k-1}t_1 - c_{k-2}t_2 + \cdots \pm t_k\right)\). Only matrix multiplication and addition of diagonal entries are needed, which is why this is the method set as practical 6 of STS-106.
Given.
\[ A = \begin{pmatrix} 2 & 1 & 1 \\ 1 & 2 & 1 \\ 1 & 1 & 2 \end{pmatrix}. \]Asked. Find the characteristic equation by the trace method, the roots, and the eigenvectors.
Step 1 — the three traces. \(\operatorname{trace}(A) = 2+2+2 = 6\). Squaring,
\[ A^{2} = \begin{pmatrix} 6 & 5 & 5 \\ 5 & 6 & 5 \\ 5 & 5 & 6 \end{pmatrix}, \qquad \operatorname{trace}(A^{2}) = 18, \] \[ A^{3} = \begin{pmatrix} 22 & 21 & 21 \\ 21 & 22 & 21 \\ 21 & 21 & 22 \end{pmatrix}, \qquad \operatorname{trace}(A^{3}) = 66. \]Step 2 — the coefficients.
\[ c_1 = 6, \qquad c_2 = \frac{6 \times 6 - 18}{2} = \frac{18}{2} = 9, \qquad c_3 = \frac{9 \times 6 - 6 \times 18 + 66}{3} = \frac{54 - 108 + 66}{3} = \frac{12}{3} = 4. \]Step 3 — the characteristic equation.
\[ \lambda^{3} - 6\lambda^{2} + 9\lambda - 4 = 0. \]A check that costs nothing: \(c_3\) must equal \(\det(A)\), and \(\det(A) = 4\). \(\checkmark\)
Step 4 — factorise. Trying \(\lambda = 1\): \(1 - 6 + 9 - 4 = 0\), so \((\lambda - 1)\) is a factor. Dividing,
\[ \lambda^{3} - 6\lambda^{2} + 9\lambda - 4 = (\lambda - 1)\left(\lambda^{2} - 5\lambda + 4\right) = (\lambda - 1)(\lambda - 1)(\lambda - 4) = (\lambda - 1)^{2}(\lambda - 4). \]So the roots are \(\lambda = 4\) once and \(\lambda = 1\) twice. Two checks: \(4 + 1 + 1 = 6 = \operatorname{trace}(A)\) and \(4 \times 1 \times 1 = 4 = \det(A)\). \(\checkmark\)
Step 5 — the eigenvector for \(\lambda = 4\).
\[ A - 4I = \begin{pmatrix} -2 & 1 & 1 \\ 1 & -2 & 1 \\ 1 & 1 & -2 \end{pmatrix}. \]Its rank is 2, so the null space has dimension \(3 - 2 = 1\). From the first two rows, \(-2x_1 + x_2 + x_3 = 0\) and \(x_1 - 2x_2 + x_3 = 0\); subtracting gives \(-3x_1 + 3x_2 = 0\), so \(x_1 = x_2\), and then \(x_3 = 2x_1 - x_2 = x_1\). Hence
\[ \mathbf{x} = (1, 1, 1)' \quad \text{(up to a scalar)}. \]Step 6 — the eigenvectors for \(\lambda = 1\).
\[ A - I = \begin{pmatrix} 1 & 1 & 1 \\ 1 & 1 & 1 \\ 1 & 1 & 1 \end{pmatrix}, \]whose rank is 1, so the null space has dimension \(3 - 1 = 2\). The single equation is \(x_1 + x_2 + x_3 = 0\), and a basis for its solutions is
\[ (-1, 1, 0)' \quad \text{and} \quad (-1, 0, 1)'. \]Step 7 — check orthogonality across the two roots. \((1,1,1)\cdot(-1,1,0) = -1 + 1 + 0 = 0\) and \((1,1,1)\cdot(-1,0,1) = -1 + 0 + 1 = 0\), as the symmetry of \(A\) requires. (The two vectors for \(\lambda = 1\) are not orthogonal to each other — they need not be — but Gram–Schmidt from Unit 1 makes them so.)
For a root \(\lambda\) of \(|A - \lambda I| = 0\),
Always \(1 \le g(\lambda) \le a(\lambda)\). When \(g < a\) for some root the matrix is called defective and cannot be diagonalised.
For a real symmetric matrix the two always agree, which is why every covariance matrix, every projection matrix and every quadratic-form matrix in statistics is diagonalisable. Defectiveness is a possibility one must know about and then, in this subject, rarely meets.
Case 1 — the symmetric matrix of Example 2.1. For \(\lambda = 1\), the algebraic multiplicity is 2 because \((\lambda-1)^{2}\) divides the polynomial. The geometric multiplicity is \(3 - \operatorname{rank}(A - I) = 3 - 1 = 2\). They agree, so \(A\) is diagonalisable:
\[ a(4) = g(4) = 1, \qquad a(1) = g(1) = 2, \qquad 1 + 2 = 3 = n. \]Case 2 — a defective matrix. Take
\[ B = \begin{pmatrix} 2 & 1 \\ 0 & 2 \end{pmatrix}. \]Its characteristic equation is \(|B - \lambda I| = (2-\lambda)^{2} = 0\), so \(\lambda = 2\) with \(a(2) = 2\). But
\[ B - 2I = \begin{pmatrix} 0 & 1 \\ 0 & 0 \end{pmatrix}, \qquad \operatorname{rank}(B - 2I) = 1, \qquad g(2) = 2 - 1 = 1. \]Solving \((B-2I)\mathbf{x} = \mathbf{0}\) gives \(x_2 = 0\) with \(x_1\) free, so the only eigenvectors are multiples of \((1,0)'\). There is no second independent one, and \(B\) cannot be written as \(P\Lambda P^{-1}\).
Interpretation. \(B\) is not symmetric, and that is exactly what permits the failure. The moment a matrix in this course is a covariance matrix, a correlation matrix or a projection matrix, it is symmetric, and the possibility does not arise.
Every square matrix satisfies its own characteristic equation. If
\[ p(\lambda) = \lambda^{n} - c_1\lambda^{n-1} + \cdots \pm c_n, \]then \(p(A) = O\), the zero matrix.
The practical consequence. Rearranging \(p(A) = O\) gives \(A^{n}\) as a combination of lower powers, so every power of \(A\) reduces to a combination of \(I, A, \ldots, A^{n-1}\). And when \(c_n = \det(A) \ne 0\), multiplying by \(A^{-1}\) expresses the inverse without any division:
\[ A^{-1} = \frac{1}{c_n}\left(A^{n-1} - c_1 A^{n-2} + \cdots \right). \]Given. \(A\) of Example 2.1, with \(p(\lambda) = \lambda^{3} - 6\lambda^{2} + 9\lambda - 4\).
Step 1 — assemble \(p(A)\) entry by entry. Using the powers computed in Example 2.1, the \((1,1)\) entry is
\[ 22 - 6(6) + 9(2) - 4(1) = 22 - 36 + 18 - 4 = 0, \]and the \((1,2)\) entry is
\[ 21 - 6(5) + 9(1) - 4(0) = 21 - 30 + 9 - 0 = 0. \]By the symmetry of every matrix involved, all nine entries follow the same two patterns, so
\[ A^{3} - 6A^{2} + 9A - 4I = O. \checkmark \]Step 2 — rearrange for the inverse. Multiply through by \(A^{-1}\), which exists since \(\det(A) = 4 \ne 0\):
\[ A^{2} - 6A + 9I - 4A^{-1} = O \;\Longrightarrow\; A^{-1} = \tfrac{1}{4}\left(A^{2} - 6A + 9I\right). \]Step 3 — evaluate. The \((1,1)\) entry is \(\tfrac14(6 - 12 + 9) = \tfrac34\); the \((1,2)\) entry is \(\tfrac14(5 - 6 + 0) = -\tfrac14\). So
\[ A^{-1} = \frac{1}{4}\begin{pmatrix} 3 & -1 & -1 \\ -1 & 3 & -1 \\ -1 & -1 & 3 \end{pmatrix}. \]Step 4 — verify against a direct computation. Multiplying \(A A^{-1}\), the \((1,1)\) entry is \(\tfrac14\left[2(3) + 1(-1) + 1(-1)\right] = \tfrac14(6 - 1 - 1) = 1\), and the \((1,2)\) entry is \(\tfrac14\left[2(-1) + 1(3) + 1(-1)\right] = \tfrac14(-2 + 3 - 1) = 0\). The product is \(I\). \(\checkmark\)
Interpretation. No cofactors and no division by a determinant at any intermediate stage — only matrix multiplication and one division at the end. This is also the route used in practical 1 of STS-106, where the same inverse is produced a third way, by partitioning.
Statement. Every real symmetric \(A\) of order \(n\) can be written
\[ A = P \Lambda P', \qquad P'P = PP' = I, \qquad \Lambda = \operatorname{diag}(\lambda_1, \ldots, \lambda_n), \]with \(P\) orthogonal and its columns the orthonormal eigenvectors. Equivalently, grouping the distinct roots \(\lambda_{(1)}, \ldots, \lambda_{(k)}\),
\[ A = \sum_{j=1}^{k} \lambda_{(j)} P_j, \qquad P_j P_j = P_j, \qquad P_j P_l = O \ (j \ne l), \qquad \sum_{j=1}^{k} P_j = I, \]where \(P_j\) is the orthogonal projection onto the eigenspace of \(\lambda_{(j)}\).
Two immediate consequences. Any power follows term by term, \(A^{m} = \sum_j \lambda_{(j)}^{m} P_j\), which gives \(A^{-1}\) at \(m = -1\) and the square root \(A^{1/2} = \sum_j \sqrt{\lambda_{(j)}}\,P_j\) when all roots are non-negative. The square root is what standardises a multivariate normal vector, and it is how the Mahalanobis distance of Multivariate Analysis (STS-202) is constructed.
Given. \(A\) of Example 2.1, with roots \(4\) and \(1\) (twice).
Step 1 — the projection onto the \(\lambda = 4\) eigenspace. That space is spanned by \(\mathbf{v} = (1,1,1)'\), with \(\mathbf{v}'\mathbf{v} = 3\), so
\[ P_1 = \frac{\mathbf{v}\mathbf{v}'}{\mathbf{v}'\mathbf{v}} = \frac{1}{3} \begin{pmatrix} 1 & 1 & 1 \\ 1 & 1 & 1 \\ 1 & 1 & 1 \end{pmatrix} = \frac{J}{3}. \]Step 2 — the other projection. Since the two eigenspaces together fill \(\mathbb{R}^{3}\), \(P_2 = I - P_1 = I - J/3\).
Step 3 — check the projection properties. Using \(J^{2} = 3J\) (each entry of \(J^2\) is a sum of three ones),
\[ P_1^{2} = \frac{J^{2}}{9} = \frac{3J}{9} = \frac{J}{3} = P_1, \qquad \operatorname{trace}(P_1) = \frac{3}{3} = 1 = g(4), \] \[ P_2^{2} = I - \frac{2J}{3} + \frac{J^{2}}{9} = I - \frac{2J}{3} + \frac{J}{3} = I - \frac{J}{3} = P_2, \qquad \operatorname{trace}(P_2) = 3 - 1 = 2 = g(1), \] \[ P_1 P_2 = \frac{J}{3}\left(I - \frac{J}{3}\right) = \frac{J}{3} - \frac{J^{2}}{9} = \frac{J}{3} - \frac{J}{3} = O. \checkmark \]Step 4 — reassemble \(A\).
\[ 4P_1 + 1 \cdot P_2 = \frac{4J}{3} + I - \frac{J}{3} = I + J = \begin{pmatrix} 2 & 1 & 1 \\ 1 & 2 & 1 \\ 1 & 1 & 2 \end{pmatrix} = A. \checkmark \](The last step uses that \(I + J\) has \(1 + 1 = 2\) on the diagonal and \(0 + 1 = 1\) off it.)
Step 5 — get the inverse for free. Applying \(A^{m} = \sum_j \lambda_{(j)}^{m} P_j\) at \(m = -1\),
\[ A^{-1} = \tfrac14 P_1 + 1 \cdot P_2 = \frac{J}{12} + I - \frac{J}{3} = I - \frac{J}{4}, \]whose diagonal is \(1 - \tfrac14 = \tfrac34\) and off-diagonal \(-\tfrac14\) — the same matrix as Example 2.3 produced by Cayley–Hamilton. \(\checkmark\)
Step 6 — and the square root.
\[ A^{1/2} = 2P_1 + 1 \cdot P_2 = \frac{2J}{3} + I - \frac{J}{3} = I + \frac{J}{3}, \]with diagonal \(1 + \tfrac13 = \tfrac43\) and off-diagonal \(\tfrac13\). Squaring it returns \(I + \tfrac{2J}{3} + \tfrac{J^{2}}{9} = I + \tfrac{2J}{3} + \tfrac{J}{3} = I + J = A\). \(\checkmark\)
Interpretation. Three different quantities — the inverse, the square root, and any power — all come from one decomposition, and each was checked against an independent route rather than accepted. \(A = I + J\) is the equicorrelation structure, the covariance matrix of variables that are exchangeable and equally correlated, and it is the standard worked case because its spectrum is so simple.