A general matrix changes both length and direction. Eigenvectors identify exceptional directions that remain on the same line and are only scaled or reversed.

An eigenvector turns a matrix action into scaling

A nonzero vector v\mathbf v is an eigenvector of AA when

Av=λv.A\mathbf v=\lambda\mathbf v.

The scalar λ\lambda is its eigenvalue. The zero vector is excluded because it satisfies the equation for every scalar and identifies no direction. Rearranging gives

(A−λI)v=0.(A-\lambda I)\mathbf v=\mathbf0.

A nonzero solution exists exactly when A−λIA-\lambda I is singular, so

det⁡(A−λI)=0.\det(A-\lambda I)=0.

This is the characteristic equation. Finding a root is only the first step; solving the corresponding homogeneous system gives the eigenvectors.

We use the monic convention pA(t)=det⁡(tI−A)p_A(t)=\det(tI-A). Since det⁡(A−tI)=(−1)npA(t)\det(A-tI)=(-1)^np_A(t), either characteristic equation has the same roots and multiplicities.

An eigenspace is a kernel

For fixed λ\lambda, include zero with all corresponding eigenvectors:

Eλ=ker⁡(A−λI).E_\lambda=\ker(A-\lambda I).

This eigenspace is a subspace. Its dimension is the geometric multiplicity. The root multiplicity of λ\lambda in the characteristic polynomial is its algebraic multiplicity, and

1≤dim⁡Eλ≤mult⁡λ(pA).1\le \dim E_\lambda \le \operatorname{mult}_{\lambda}(p_A).

Eigenvectors belonging to distinct eigenvalues are independent. This permits several one-dimensional invariant directions to form a basis.

Why distinct eigenvalues give independent directions

Proceed by induction. Suppose the first k−1k-1 eigenvectors are independent and a combination of all kk vanishes. Apply A−λkIA-\lambda_k I to obtain

∑i=1k−1ci(λi−λk)vi=0.\sum_{i=1}^{k-1}c_i(\lambda_i-\lambda_k)\mathbf v_i=0.

Independence and distinct eigenvalues force the first k−1k-1 coefficients to vanish; substitution gives ck=0c_k=0. The proof uses eigenvalue information, not merely pairwise nonparallelism.

Why geometric multiplicity cannot exceed algebraic multiplicity

Let g=dim⁡Eλg=\dim E_\lambda. Extend a basis of the eigenspace to a basis of the entire space. In these coordinates,

A~=(λIgC0B).\widetilde A=\begin{pmatrix}\lambda I_g&C\\0&B\end{pmatrix}.

The lower-left block vanishes because the first gg basis vectors map to their own λ\lambda multiples. The block-triangular determinant gives

det⁡(tI−A~)=(t−λ)gdet⁡(tI−B).\det(tI-\widetilde A)=(t-\lambda)^g\det(tI-B).

Hence the root multiplicity is at least gg. Distinct eigenspaces have a direct sum, by the same elimination argument used for independent eigenvectors, applied to nonzero combinations in each eigenspace. When the characteristic polynomial splits, their dimensions sum to nn exactly when every geometric multiplicity equals its algebraic multiplicity, equivalently when an eigenvector basis exists.

Invariant subspaces are more general

A subspace WW is invariant under AA when A(W)⊆WA(W)\subseteq W. Every eigenvector spans an invariant line, but an entire rotation plane can be invariant without containing a real eigenvector. A basis adapted to invariant subspaces gives a block matrix and can split a large problem into smaller ones.

Diagonalization means finding an eigenvector basis

If AA has nn independent eigenvectors, put them in the columns of SS and their eigenvalues on the diagonal of DD. Then

AS=SD,A=SDS−1.AS=SD, \qquad A=SDS^{-1}.

Conversely, the columns of SS in such a factorization are eigenvectors. A square matrix is therefore diagonalizable exactly when the space has an eigenvector basis.

Distinct eigenvalues guarantee enough independent eigenvectors. Repeated eigenvalues require checking eigenspace dimensions. For example,

(1101)\begin{pmatrix}1&1\\0&1\end{pmatrix}

has only a one-dimensional eigenspace and cannot supply a basis of two eigenvectors.

Construct a complete diagonalization

For A=(2103)A=\begin{pmatrix}2&1\\0&3\end{pmatrix}, choose eigenvectors (1,0)T(1,0)^{\mathsf T} and (1,1)T(1,1)^{\mathsf T}. Then

S=(1101),D=(2003),S−1=(1−101).S=\begin{pmatrix}1&1\\0&1\end{pmatrix}, \quad D=\begin{pmatrix}2&0\\0&3\end{pmatrix}, \quad S^{-1}=\begin{pmatrix}1&-1\\0&1\end{pmatrix}.

An input (x,y)(x,y) has new coordinates (x−y,y)(x-y,y). Scale by 2,32,3 and convert back to obtain (2x+y,3y)(2x+y,3y), exactly the original action. This explains the order of S−1S^{-1}, DD, and SS.

Diagonalizability depends on the scalar field. A ninety-degree rotation is diagonalizable over the complex numbers but not the reals. In general the characteristic polynomial must split over the chosen field and every geometric multiplicity must equal its algebraic multiplicity.

Real matrices may need complex scalars

The ninety-degree rotation

R=(0−110)R=\begin{pmatrix}0&-1\\1&0\end{pmatrix}

has characteristic equation λ2+1=0\lambda^2+1=0. No real direction remains on its original line. Over the complex numbers, the eigenvalues are ii and −i-i. The existing complex-numbers chapter reviews ii, conjugation, and polar form.

For complex vectors, use the conjugate inner product ⟨x,y⟩=x∗y\langle\mathbf x,\mathbf y\rangle=\mathbf x^*\mathbf y, where ∗^* is conjugate transpose. Ordinary transpose can fail to produce a nonnegative squared norm.

Similarity preserves spectral information

A change of basis replaces AA by S−1ASS^{-1}AS. Similar matrices have the same characteristic polynomial, eigenvalues, determinant, and trace. Their entries change while the scaling behavior of the underlying map does not.

Why similarity preserves characteristic polynomial and trace

Use tI−S−1AS=S−1(tI−A)StI-S^{-1}AS=S^{-1}(tI-A)S and determinant multiplicativity to obtain equal characteristic polynomials. Trace is the sum of diagonal entries. Interchanging finite sums gives

tr⁡(XY)=∑i∑jxijyji=tr⁡(YX).\operatorname{tr}(XY)=\sum_i\sum_j x_{ij}y_{ji} =\operatorname{tr}(YX).

Consequently tr⁡(S−1AS)=tr⁡(ASS−1)=tr⁡(A)\operatorname{tr}(S^{-1}AS)=\operatorname{tr}(ASS^{-1})=\operatorname{tr}(A).

Exercises

ExerciseFrom eigenvalue to eigenspace

Find the eigenvalues of A=(2103)A=\begin{pmatrix}2&1\\0&3\end{pmatrix} and an eigenvector basis.

Solution

The eigenvalues are 22 and 33. Corresponding vectors (1,0)T(1,0)^{\mathsf T} and (1,1)T(1,1)^{\mathsf T} are independent, so AA is diagonalizable.

ExerciseA stationary state is an eigenvector

Explain why Pπ=πP\boldsymbol\pi=\boldsymbol\pi makes a stationary state an eigenvector for eigenvalue 11.

Solution

The equation is exactly Pπ=1πP\boldsymbol\pi=1\boldsymbol\pi. Probability normalization selects from this eigenspace a vector whose entries sum to one.

ExerciseOne distinct root can still suffice

Explain why the identity matrix is diagonalizable although it has only one distinct eigenvalue.

Solution

Its eigenspace for eigenvalue 11 is all of Rn\mathbb R^n, so it contains a basis of nn eigenvectors.