Eigenvalue decomposition mainly concerns square matrices. Singular value decomposition applies to every matrix and uses separate orthogonal bases in the domain and codomain to describe one linear map.
Every matrix is orthogonal change, scaling, and orthogonal change
For every real matrix , there are orthogonal matrices and a nonnegative diagonal-shaped matrix such that
expresses inputs in the right-singular-vector basis, scales perpendicular directions, and places them in the output space. The positive entries
are the singular values. Their count is .
Check dimensions in a rectangular factorization
In the full SVD, is , is , and is . Keeping only positive modes gives the compact form , with columns on each side. Orthogonal transformations can include reflections, so they are not always rotations.
Singular values come from symmetric matrices
The matrix is symmetric positive semidefinite. The spectral theorem gives
For , define . These vectors are orthonormal and satisfy . Completing bases on both sides produces the full SVD.
Eigenvalues may be negative or complex; singular values are always nonnegative real numbers. SVD exists even when a square matrix is not diagonalizable.
Symmetry follows by transposition, and positive semidefiniteness follows from . For the completed bases, the equations on every give , so right multiplication by proves the factorization. The rank is because multiplication by invertible basis matrices preserves the image dimension and has exactly independent nonzero columns. When , and any orthogonal bases work.
Why the constructed left vectors are orthonormal
For positive singular values,
A zero eigenvalue gives , hence . These two cases complete the construction from the spectral theorem.
SVD aligns the four fundamental subspaces
Right singular vectors with positive singular values span the row space; those with zero singular value span . Positive left singular vectors span the column space; the remaining ones span . Thus
This places the row, column, and null spaces computed in vector spaces and bases into one orthogonal structure.
Writing gives . Orthogonality makes this zero exactly when , and it shows that the image is exactly the span of . Apply the same argument to to identify its kernel and image. The row space is the image of , proving both orthogonal-complement identities.
The pseudoinverse unifies exact and least-squares solutions
Invert each positive singular value to form . The Moore–Penrose pseudoinverse is
The vector is the minimum-norm least-squares solution. For a consistent system it is the shortest exact solution, and for an invertible matrix it equals .
A small singular value amplifies data error by roughly . Truncating very small singular values sacrifices some fit in exchange for stability.
General proof of the minimum-norm property
For the full SVD , put and . Orthogonal transformations preserve norms, giving
The final sum is independent of the input. Each term in the first sum vanishes precisely when . The remaining input coordinates are free and describe the kernel. Since , setting every free coordinate to zero gives the unique minimum norm. This is exactly .
All solutions of a rectangular example
For the single-row matrix , a compact SVD is , , and . Thus
For , it returns . Every exact solution is , with squared length . The pseudoinverse selects the unique shortest one.
Truncated SVD gives the best low-rank approximation
Write SVD as an outer-product sum:
Keeping the first terms gives . The Eckart–Young theorem states that is a closest matrix of rank at most in the operator or Frobenius norm. The discarded singular values determine the error exactly.
Precise low-rank errors and PCA conventions
The operator two-norm is the largest output norm from a unit input. The Frobenius norm is the square root of the sum of squared entries. For ,
For operator-norm optimality, any rank-at-most- matrix has a unit null vector in the span of the first right singular vectors. Therefore . Truncation attains this lower bound.
To verify the error formulas, write a unit input in the right singular basis. The squared output under is , attained at . For the Frobenius norm, orthogonal left multiplication preserves each column length, and orthogonal right multiplication preserves each row length. Thus it preserves the sum of squared entries, reducing the computation to the discarded diagonal entries of .
Proof of Frobenius-norm low-rank optimality
Let be the column space of , of dimension at most , and let be its orthogonal projector. Apply Pythagoras column by column:
For a full left singular basis, set . Then and : expand using an orthonormal basis of and sum over the complete basis. Orthogonality in the outer-product expansion gives
The inequality assigns at most units of weight to the largest coefficients. Moving weight from a smaller coefficient to an unfilled larger one never decreases the sum. Thus every has squared error at least , attained by truncated SVD. For , taking gives zero error.
PCA also requires a centering decision
Place mean-centered samples in a data matrix. Right singular vectors give principal directions, and squared singular values are proportional to variance along them. The existing expectation and variance chapter supplies the probabilistic language. If feature units differ greatly, one must also decide whether to standardize, since raw scale can dominate the components.
For PCA, explicitly put samples in rows of centered . The sample covariance is
Principal directions are right singular vectors and their variances are . Putting samples in columns exchanges the roles of left and right singular vectors.
For a unit feature direction , the projected sample column is , with mean zero and sample variance . The Rayleigh bound therefore proves maximal variance at the top right singular vector. Restricting to its orthogonal complement and repeating proves the successive principal directions; tied eigenvalues allow any orthonormal basis within the tied eigenspace.
Exercises
Explain why has exactly one positive singular value.
Solution
The second column is twice the first, so the rank is one. The number of positive singular values equals the rank.
Use SVD to show that when is invertible.
Solution
All singular values are positive, so .
If the singular values are , what is the operator-norm error of the best rank-one approximation?
Solution
The error is the first discarded singular value, which is .
Comments