Skip to the content

Topics Covered

Characteristic Roots Trace Method Algebraic vs Geometric Multiplicity Cayley–Hamilton Spectral Decomposition Matrix Square Root
On this page
  1. 1. Characteristic Roots and Vectors
  2. 2. Algebraic and Geometric Multiplicity
  3. 3. The Cayley–Hamilton Theorem
  4. 4. Spectral Decomposition of a Real Symmetric Matrix
  5. Key Take-aways
Where this unit starts. Eigenvalues, eigenvectors and the Cayley–Hamilton theorem are stated and used at exam level in UGC NET Statistics, Unit 2. This unit proves what that page asserts, and adds the two things it does not reach: the distinction between algebraic and geometric multiplicity, and the spectral decomposition of a real symmetric matrix, which is what every later result about quadratic forms and covariance matrices rests on.

1. Characteristic Roots and Vectors

DEFINITION

For a square matrix \(A\) of order \(n\), a scalar \(\lambda\) and a non-zero vector \(\mathbf{x}\) satisfying

\[ A\mathbf{x} = \lambda\mathbf{x} \]

are a characteristic root (eigenvalue) and a corresponding characteristic vector (eigenvector). Rewriting, \((A - \lambda I)\mathbf{x} = \mathbf{0}\), and a non-zero solution exists exactly when \(A - \lambda I\) is singular. So the roots are the solutions of the characteristic equation

\[ \left| A - \lambda I \right| = 0, \]

a polynomial of degree \(n\) in \(\lambda\).

PROPERTIES \[ \begin{aligned} &\sum_{i=1}^{n} \lambda_i = \operatorname{trace}(A), \qquad \prod_{i=1}^{n} \lambda_i = \det(A) \\ &A \text{ and } A' \text{ have the same characteristic roots} \\ &A \text{ real symmetric} \Rightarrow \text{all } \lambda_i \text{ real, and eigenvectors for distinct } \lambda_i \text{ are orthogonal} \\ &A \text{ idempotent} \Rightarrow \lambda_i \in \{0, 1\}, \text{ and } \operatorname{rank}(A) = \operatorname{trace}(A) \\ &\lambda_i(A^{k}) = \lambda_i^{k}, \qquad \lambda_i(A^{-1}) = 1/\lambda_i \text{ when } A \text{ is non-singular} \end{aligned} \]

The idempotent line is the one that matters most for statistics: it is exactly the criterion used in Distribution Theory, Unit 4 to decide when a quadratic form is chi-square.

BUILDING THE CHARACTERISTIC EQUATION FROM TRACES OF POWERS

Expanding \(|A - \lambda I|\) by cofactors is laborious for \(n = 3\) and worse beyond. The Leverrier–Faddeev identities give the coefficients from traces alone. Writing the characteristic polynomial as

\[ \lambda^{n} - c_1 \lambda^{n-1} + c_2 \lambda^{n-2} - \cdots \pm c_n = 0, \]

and \(t_k = \operatorname{trace}(A^{k})\),

\[ c_1 = t_1, \qquad c_2 = \frac{c_1 t_1 - t_2}{2}, \qquad c_3 = \frac{c_2 t_1 - c_1 t_2 + t_3}{3}, \]

and in general \(c_k = \dfrac{1}{k}\left(c_{k-1}t_1 - c_{k-2}t_2 + \cdots \pm t_k\right)\). Only matrix multiplication and addition of diagonal entries are needed, which is why this is the method set as practical 6 of STS-106.

EXAMPLE 2.1 — ROOTS AND VECTORS OF A 3 × 3 SYMMETRIC MATRIX

Given.

\[ A = \begin{pmatrix} 2 & 1 & 1 \\ 1 & 2 & 1 \\ 1 & 1 & 2 \end{pmatrix}. \]

Asked. Find the characteristic equation by the trace method, the roots, and the eigenvectors.

Step 1 — the three traces. \(\operatorname{trace}(A) = 2+2+2 = 6\). Squaring,

\[ A^{2} = \begin{pmatrix} 6 & 5 & 5 \\ 5 & 6 & 5 \\ 5 & 5 & 6 \end{pmatrix}, \qquad \operatorname{trace}(A^{2}) = 18, \] \[ A^{3} = \begin{pmatrix} 22 & 21 & 21 \\ 21 & 22 & 21 \\ 21 & 21 & 22 \end{pmatrix}, \qquad \operatorname{trace}(A^{3}) = 66. \]

Step 2 — the coefficients.

\[ c_1 = 6, \qquad c_2 = \frac{6 \times 6 - 18}{2} = \frac{18}{2} = 9, \qquad c_3 = \frac{9 \times 6 - 6 \times 18 + 66}{3} = \frac{54 - 108 + 66}{3} = \frac{12}{3} = 4. \]

Step 3 — the characteristic equation.

\[ \lambda^{3} - 6\lambda^{2} + 9\lambda - 4 = 0. \]

A check that costs nothing: \(c_3\) must equal \(\det(A)\), and \(\det(A) = 4\). \(\checkmark\)

Step 4 — factorise. Trying \(\lambda = 1\): \(1 - 6 + 9 - 4 = 0\), so \((\lambda - 1)\) is a factor. Dividing,

\[ \lambda^{3} - 6\lambda^{2} + 9\lambda - 4 = (\lambda - 1)\left(\lambda^{2} - 5\lambda + 4\right) = (\lambda - 1)(\lambda - 1)(\lambda - 4) = (\lambda - 1)^{2}(\lambda - 4). \]

So the roots are \(\lambda = 4\) once and \(\lambda = 1\) twice. Two checks: \(4 + 1 + 1 = 6 = \operatorname{trace}(A)\) and \(4 \times 1 \times 1 = 4 = \det(A)\). \(\checkmark\)

Step 5 — the eigenvector for \(\lambda = 4\).

\[ A - 4I = \begin{pmatrix} -2 & 1 & 1 \\ 1 & -2 & 1 \\ 1 & 1 & -2 \end{pmatrix}. \]

Its rank is 2, so the null space has dimension \(3 - 2 = 1\). From the first two rows, \(-2x_1 + x_2 + x_3 = 0\) and \(x_1 - 2x_2 + x_3 = 0\); subtracting gives \(-3x_1 + 3x_2 = 0\), so \(x_1 = x_2\), and then \(x_3 = 2x_1 - x_2 = x_1\). Hence

\[ \mathbf{x} = (1, 1, 1)' \quad \text{(up to a scalar)}. \]

Step 6 — the eigenvectors for \(\lambda = 1\).

\[ A - I = \begin{pmatrix} 1 & 1 & 1 \\ 1 & 1 & 1 \\ 1 & 1 & 1 \end{pmatrix}, \]

whose rank is 1, so the null space has dimension \(3 - 1 = 2\). The single equation is \(x_1 + x_2 + x_3 = 0\), and a basis for its solutions is

\[ (-1, 1, 0)' \quad \text{and} \quad (-1, 0, 1)'. \]

Step 7 — check orthogonality across the two roots. \((1,1,1)\cdot(-1,1,0) = -1 + 1 + 0 = 0\) and \((1,1,1)\cdot(-1,0,1) = -1 + 0 + 1 = 0\), as the symmetry of \(A\) requires. (The two vectors for \(\lambda = 1\) are not orthogonal to each other — they need not be — but Gram–Schmidt from Unit 1 makes them so.)

2. Algebraic and Geometric Multiplicity

THE TWO COUNTS, AND WHEN THEY DIFFER

For a root \(\lambda\) of \(|A - \lambda I| = 0\),

Always \(1 \le g(\lambda) \le a(\lambda)\). When \(g < a\) for some root the matrix is called defective and cannot be diagonalised.

For a real symmetric matrix the two always agree, which is why every covariance matrix, every projection matrix and every quadratic-form matrix in statistics is diagonalisable. Defectiveness is a possibility one must know about and then, in this subject, rarely meets.

EXAMPLE 2.2 — THE TWO COUNTS, IN BOTH CASES

Case 1 — the symmetric matrix of Example 2.1. For \(\lambda = 1\), the algebraic multiplicity is 2 because \((\lambda-1)^{2}\) divides the polynomial. The geometric multiplicity is \(3 - \operatorname{rank}(A - I) = 3 - 1 = 2\). They agree, so \(A\) is diagonalisable:

\[ a(4) = g(4) = 1, \qquad a(1) = g(1) = 2, \qquad 1 + 2 = 3 = n. \]

Case 2 — a defective matrix. Take

\[ B = \begin{pmatrix} 2 & 1 \\ 0 & 2 \end{pmatrix}. \]

Its characteristic equation is \(|B - \lambda I| = (2-\lambda)^{2} = 0\), so \(\lambda = 2\) with \(a(2) = 2\). But

\[ B - 2I = \begin{pmatrix} 0 & 1 \\ 0 & 0 \end{pmatrix}, \qquad \operatorname{rank}(B - 2I) = 1, \qquad g(2) = 2 - 1 = 1. \]

Solving \((B-2I)\mathbf{x} = \mathbf{0}\) gives \(x_2 = 0\) with \(x_1\) free, so the only eigenvectors are multiples of \((1,0)'\). There is no second independent one, and \(B\) cannot be written as \(P\Lambda P^{-1}\).

Interpretation. \(B\) is not symmetric, and that is exactly what permits the failure. The moment a matrix in this course is a covariance matrix, a correlation matrix or a projection matrix, it is symmetric, and the possibility does not arise.

3. The Cayley–Hamilton Theorem

STATEMENT AND ITS USE

Every square matrix satisfies its own characteristic equation. If

\[ p(\lambda) = \lambda^{n} - c_1\lambda^{n-1} + \cdots \pm c_n, \]

then \(p(A) = O\), the zero matrix.

The practical consequence. Rearranging \(p(A) = O\) gives \(A^{n}\) as a combination of lower powers, so every power of \(A\) reduces to a combination of \(I, A, \ldots, A^{n-1}\). And when \(c_n = \det(A) \ne 0\), multiplying by \(A^{-1}\) expresses the inverse without any division:

\[ A^{-1} = \frac{1}{c_n}\left(A^{n-1} - c_1 A^{n-2} + \cdots \right). \]
EXAMPLE 2.3 — VERIFYING CAYLEY–HAMILTON AND INVERTING WITH IT

Given. \(A\) of Example 2.1, with \(p(\lambda) = \lambda^{3} - 6\lambda^{2} + 9\lambda - 4\).

Step 1 — assemble \(p(A)\) entry by entry. Using the powers computed in Example 2.1, the \((1,1)\) entry is

\[ 22 - 6(6) + 9(2) - 4(1) = 22 - 36 + 18 - 4 = 0, \]

and the \((1,2)\) entry is

\[ 21 - 6(5) + 9(1) - 4(0) = 21 - 30 + 9 - 0 = 0. \]

By the symmetry of every matrix involved, all nine entries follow the same two patterns, so

\[ A^{3} - 6A^{2} + 9A - 4I = O. \checkmark \]

Step 2 — rearrange for the inverse. Multiply through by \(A^{-1}\), which exists since \(\det(A) = 4 \ne 0\):

\[ A^{2} - 6A + 9I - 4A^{-1} = O \;\Longrightarrow\; A^{-1} = \tfrac{1}{4}\left(A^{2} - 6A + 9I\right). \]

Step 3 — evaluate. The \((1,1)\) entry is \(\tfrac14(6 - 12 + 9) = \tfrac34\); the \((1,2)\) entry is \(\tfrac14(5 - 6 + 0) = -\tfrac14\). So

\[ A^{-1} = \frac{1}{4}\begin{pmatrix} 3 & -1 & -1 \\ -1 & 3 & -1 \\ -1 & -1 & 3 \end{pmatrix}. \]

Step 4 — verify against a direct computation. Multiplying \(A A^{-1}\), the \((1,1)\) entry is \(\tfrac14\left[2(3) + 1(-1) + 1(-1)\right] = \tfrac14(6 - 1 - 1) = 1\), and the \((1,2)\) entry is \(\tfrac14\left[2(-1) + 1(3) + 1(-1)\right] = \tfrac14(-2 + 3 - 1) = 0\). The product is \(I\). \(\checkmark\)

Interpretation. No cofactors and no division by a determinant at any intermediate stage — only matrix multiplication and one division at the end. This is also the route used in practical 1 of STS-106, where the same inverse is produced a third way, by partitioning.

4. Spectral Decomposition of a Real Symmetric Matrix

THE THEOREM

Statement. Every real symmetric \(A\) of order \(n\) can be written

\[ A = P \Lambda P', \qquad P'P = PP' = I, \qquad \Lambda = \operatorname{diag}(\lambda_1, \ldots, \lambda_n), \]

with \(P\) orthogonal and its columns the orthonormal eigenvectors. Equivalently, grouping the distinct roots \(\lambda_{(1)}, \ldots, \lambda_{(k)}\),

\[ A = \sum_{j=1}^{k} \lambda_{(j)} P_j, \qquad P_j P_j = P_j, \qquad P_j P_l = O \ (j \ne l), \qquad \sum_{j=1}^{k} P_j = I, \]

where \(P_j\) is the orthogonal projection onto the eigenspace of \(\lambda_{(j)}\).

Two immediate consequences. Any power follows term by term, \(A^{m} = \sum_j \lambda_{(j)}^{m} P_j\), which gives \(A^{-1}\) at \(m = -1\) and the square root \(A^{1/2} = \sum_j \sqrt{\lambda_{(j)}}\,P_j\) when all roots are non-negative. The square root is what standardises a multivariate normal vector, and it is how the Mahalanobis distance of Multivariate Analysis (STS-202) is constructed.

EXAMPLE 2.4 — THE SPECTRAL DECOMPOSITION OF THE SAME MATRIX

Given. \(A\) of Example 2.1, with roots \(4\) and \(1\) (twice).

Step 1 — the projection onto the \(\lambda = 4\) eigenspace. That space is spanned by \(\mathbf{v} = (1,1,1)'\), with \(\mathbf{v}'\mathbf{v} = 3\), so

\[ P_1 = \frac{\mathbf{v}\mathbf{v}'}{\mathbf{v}'\mathbf{v}} = \frac{1}{3} \begin{pmatrix} 1 & 1 & 1 \\ 1 & 1 & 1 \\ 1 & 1 & 1 \end{pmatrix} = \frac{J}{3}. \]

Step 2 — the other projection. Since the two eigenspaces together fill \(\mathbb{R}^{3}\), \(P_2 = I - P_1 = I - J/3\).

Step 3 — check the projection properties. Using \(J^{2} = 3J\) (each entry of \(J^2\) is a sum of three ones),

\[ P_1^{2} = \frac{J^{2}}{9} = \frac{3J}{9} = \frac{J}{3} = P_1, \qquad \operatorname{trace}(P_1) = \frac{3}{3} = 1 = g(4), \] \[ P_2^{2} = I - \frac{2J}{3} + \frac{J^{2}}{9} = I - \frac{2J}{3} + \frac{J}{3} = I - \frac{J}{3} = P_2, \qquad \operatorname{trace}(P_2) = 3 - 1 = 2 = g(1), \] \[ P_1 P_2 = \frac{J}{3}\left(I - \frac{J}{3}\right) = \frac{J}{3} - \frac{J^{2}}{9} = \frac{J}{3} - \frac{J}{3} = O. \checkmark \]

Step 4 — reassemble \(A\).

\[ 4P_1 + 1 \cdot P_2 = \frac{4J}{3} + I - \frac{J}{3} = I + J = \begin{pmatrix} 2 & 1 & 1 \\ 1 & 2 & 1 \\ 1 & 1 & 2 \end{pmatrix} = A. \checkmark \]

(The last step uses that \(I + J\) has \(1 + 1 = 2\) on the diagonal and \(0 + 1 = 1\) off it.)

Step 5 — get the inverse for free. Applying \(A^{m} = \sum_j \lambda_{(j)}^{m} P_j\) at \(m = -1\),

\[ A^{-1} = \tfrac14 P_1 + 1 \cdot P_2 = \frac{J}{12} + I - \frac{J}{3} = I - \frac{J}{4}, \]

whose diagonal is \(1 - \tfrac14 = \tfrac34\) and off-diagonal \(-\tfrac14\) — the same matrix as Example 2.3 produced by Cayley–Hamilton. \(\checkmark\)

Step 6 — and the square root.

\[ A^{1/2} = 2P_1 + 1 \cdot P_2 = \frac{2J}{3} + I - \frac{J}{3} = I + \frac{J}{3}, \]

with diagonal \(1 + \tfrac13 = \tfrac43\) and off-diagonal \(\tfrac13\). Squaring it returns \(I + \tfrac{2J}{3} + \tfrac{J^{2}}{9} = I + \tfrac{2J}{3} + \tfrac{J}{3} = I + J = A\). \(\checkmark\)

Interpretation. Three different quantities — the inverse, the square root, and any power — all come from one decomposition, and each was checked against an independent route rather than accepted. \(A = I + J\) is the equicorrelation structure, the covariance matrix of variables that are exchangeable and equally correlated, and it is the standard worked case because its spectrum is so simple.

Key Take-aways