A vector space specifies operations within a collection. A linear map preserves them between spaces. After studying bases, we can now state exactly when a matrix represents the whole map.

Preserving combinations

A map T:V→WT:V\to W is linear if

T(au+bv)=aT(u)+bT(v).T(a\mathbf u+b\mathbf v) =aT(\mathbf u)+bT(\mathbf v).

It makes combining before processing equivalent to processing before combining. The first chapter’s response model assumes this property.

The two concepts answer different questions. A vector space specifies addition and scaling within a collection of objects. A linear map preserves those operations while mapping one space into another. Merely accepting vector inputs does not make a process linear.

A direct consequence is T(0)=0T(\mathbf0)=\mathbf0: apply additivity to 0+0\mathbf0+\mathbf0 and cancel one term. A fixed nonzero offset therefore prevents linearity. The map T(v)=v+cT(\mathbf v)=\mathbf v+\mathbf c is instead affine. Subtracting a baseline can make a meaningful difference in modelling.

Linearity must hold for arbitrary inputs

Preserving addition and scaling means

T(u+v)=T(u)+T(v),T(av)=aT(v).T(\mathbf u+\mathbf v)=T(\mathbf u)+T(\mathbf v),\qquad T(a\mathbf v)=aT(\mathbf v).

Together they imply preservation of two-term linear combinations. Conversely, choosing unit coefficients or a zero vector in the combination formula recovers these laws. Either form is useful, but its quantifiers cover every input and scalar.

For T(x,y)=(x+2y,3x−y)T(x,y)=(x+2y,3x-y), each output is a fixed linear combination of inputs, so expansion verifies linearity componentwise.

For T(x,y)=(x,y2)T(x,y)=(x,y^2), preserving zero is insufficient:

T(0,2)=(0,4)≠2T(0,1)=(0,2).T(0,2)=(0,4)\ne2T(0,1)=(0,2).

One counterexample disproves linearity; a few successful inputs do not prove it.

An affine rule T(x)=Ax+bT(\mathbf x)=A\mathbf x+\mathbf b with nonzero fixed offset fails to preserve zero. Measuring output relative to the baseline b\mathbf b leaves the linear rule x↦Ax\mathbf x\mapsto A\mathbf x. Subtracting a reference state can therefore change the appropriate mathematical description of a model.

Kernel, image, and solution sets

Define

ker⁡T={v:T(v)=0},im⁡T={T(v):v∈V}.\ker T=\{\mathbf v:T(\mathbf v)=\mathbf0\},\qquad \operatorname{im}T=\{T(\mathbf v):\mathbf v\in V\}.

The kernel contains changes invisible at the output; the image contains reachable outputs. Linearity makes both subspaces of their respective spaces.

Why kernel and image are subspaces

For u,v∈ker⁡T\mathbf u,\mathbf v\in\ker T, linearity gives

T(au+bv)=a0+b0=0.T(a\mathbf u+b\mathbf v)=a\mathbf0+b\mathbf0=\mathbf0.

The kernel is closed under combinations and contains zero.

For image vectors y1=T(u)\mathbf y_1=T(\mathbf u) and y2=T(v)\mathbf y_2=T(\mathbf v),

ay1+by2=T(au+bv),a\mathbf y_1+b\mathbf y_2=T(a\mathbf u+b\mathbf v),

which is another output. The kernel belongs to the domain VV; the image belongs to the codomain WW. Their ambient spaces may have different dimensions.

Rank-nullity for a finite-dimensional map follows directly from bases. Extend a basis of the kernel to a basis of the domain. Images of the added vectors span the image because kernel terms vanish. Those images are independent: a nontrivial dependence would place a combination of added basis vectors in the kernel, contradicting independence of the extended basis. Hence

dim⁡V=dim⁡ker⁡T+dim⁡im⁡T.\dim V=\dim\ker T+\dim\operatorname{im}T.

Injectivity, surjectivity, and invertibility

The map is injective exactly when ker⁡T={0}\ker T=\{\mathbf0\}. Injectivity permits only zero to map to zero. Conversely, equal outputs imply u−v∈ker⁡T\mathbf u-\mathbf v\in\ker T, so a trivial kernel forces equal inputs.

Surjectivity means im⁡T=W\operatorname{im}T=W. Invertibility requires both: every output must be reachable, with exactly one preimage.

When both spaces have the same finite dimension nn, rank-nullity makes injectivity and surjectivity equivalent. A zero-dimensional kernel gives an nn-dimensional image, and an image equal to the codomain leaves a zero-dimensional kernel. Different dimensions break this equivalence. For example,

J:R2→R3,J(x,y)=(x,y,0)J:\mathbb R^2\to\mathbb R^3,\quad J(x,y)=(x,y,0)

is injective but not surjective, while

R:R3→R2,R(x,y,z)=(x,y)R:\mathbb R^3\to\mathbb R^2,\quad R(x,y,z)=(x,y)

is surjective but not injective: it discards the entire zz direction.

Computing kernel, image, and preimages

Consider

T:R3→R2,T(x,y,z)=(x+y,y+z).\begin{gathered} T:\mathbb R^3\to\mathbb R^2,\\ T(x,y,z)=(x+y,y+z). \end{gathered}

Setting the output to zero gives x=−y,z=−yx=-y,z=-y, hence

ker⁡T=span⁡((−1,1,−1)).\ker T=\operatorname{span}((-1,1,-1)).

Every target (p,q)(p,q) is reached by (p,0,q)(p,0,q), so the image is all of R2\mathbb R^2. The map is surjective, but its one-dimensional kernel makes preimages nonunique:

T−1({(p,q)})={(p−t,t,q−t):t∈R}.T^{-1}(\{(p,q)\}) =\{(p-t,t,q-t):t\in\mathbb R\}.

Here T−1T^{-1} denotes a set preimage, not an inverse function. Kernel, image, and preimage identify discarded changes, reachable outputs, and the inputs producing a specified output.

All preimages of one output

If T(v)=bT(\mathbf v)=\mathbf b has a particular solution vp\mathbf v_p, all solutions are exactly

vp+ker⁡T.\mathbf v_p+\ker T.

The difference of any two solutions maps to zero, and adding any kernel element to a particular solution preserves its output. For nonzero b\mathbf b, this solution set excludes zero and is not a linear subspace.

This is the structure behind the traffic parameterization. Later, homogeneous-plus-particular solutions of circuit differential equations will use the same argument. Nonlinear equations do not automatically admit superposition.

Why a finite-dimensional linear map is determined by a matrix

Let VV have ordered basis B=(b1,…,bn)\mathcal B=(\mathbf b_1,\ldots,\mathbf b_n) and WW have ordered basis C=(c1,…,cm)\mathcal C=(\mathbf c_1,\ldots,\mathbf c_m). Knowing the images of basis vectors determines every output:

v=∑jxjbj⟹T(v)=∑jxjT(bj).\mathbf v=\sum_j x_j\mathbf b_j \quad\Longrightarrow\quad T(\mathbf v)=\sum_j x_jT(\mathbf b_j).

Put the C\mathcal C-coordinates of T(bj)T(\mathbf b_j) in column jj of AA. Then

[T(v)]C=A[v]B.[T(\mathbf v)]_{\mathcal C}=A[\mathbf v]_{\mathcal B}.

The left side is the output’s coordinate vector. The matrix itself is not an element of VV.

The representation is unique: equal matrices give equal images of basis vectors, hence equal images of every input. Conversely, any m×nm\times n matrix defines a linear map through this formula. Thus matrices and linear maps correspond bijectively when both spaces are finite-dimensional and both bases are fixed.

Differentiation can have a matrix representation

On the space P2\mathcal P_2 of polynomials of degree at most two, choose basis (1,t,t2)(1,t,t^2). Differentiation satisfies

D(1)=0,D(t)=1,D(t2)=2t.D(1)=0,\qquad D(t)=1,\qquad D(t^2)=2t.

Collect output coordinates as columns:

[D]B←B=(010002000).[D]_{\mathcal B\leftarrow\mathcal B} =\begin{pmatrix}0&1&0\\0&0&2\\0&0&0\end{pmatrix}.

The coefficient column (a,b,c)T(a,b,c)^{\mathsf T} becomes (b,2c,0)T(b,2c,0)^{\mathsf T}, representing (a+bt+ct2)′=b+2ct(a+bt+ct^2)'=b+2ct. Cubing the matrix gives zero because three derivatives annihilate every such polynomial.

This uses a finite-dimensional space closed under differentiation. The space of all polynomials is infinite-dimensional, so this fixed-size matrix cannot represent differentiation on the whole space. General function spaces require the same care.

A change of basis changes the matrix, not the map

For simplicity, consider T:V→VT:V\to V. Let AA represent it in the old basis, and let columns of SS contain the new basis vectors in old coordinates. Then

[v]old=S[v]new.[\mathbf v]_{\mathrm{old}}=S[\mathbf v]_{\mathrm{new}}.

Convert the input to old coordinates, apply AA, then convert the output back:

Anew=S−1AS.A_{\mathrm{new}}=S^{-1}AS.

For A=diag⁡(2,1)A=\operatorname{diag}(2,1) and new basis (1,1),(1,−1)(1,1),(1,-1),

S=(111−1),Anew=(3/21/21/23/2).S=\begin{pmatrix}1&1\\1&-1\end{pmatrix},\qquad A_{\mathrm{new}}=\begin{pmatrix}3/2&1/2\\1/2&3/2\end{pmatrix}.

The first column says that (1,1)(1,1) maps to (2,1)(2,1), whose new coordinates are (3/2,1/2)(3/2,1/2). The operation still doubles the first old coordinate; only its description changes.

Different bases at the two ends

A general map T:V→WT:V\to W need not have the same domain and codomain. Suppose its old matrix is AA. Let SS contain new domain basis vectors in old coordinates, and let RR contain new codomain basis vectors in old coordinates. Then

Anew=R−1AS.A_{\mathrm{new}}=R^{-1}AS.

The right factor converts input coordinates; the left factor converts outputs back. The formula becomes S−1ASS^{-1}AS only when the same basis change applies at both ends.

For the identity map on R2\mathbb R^2, use domain basis (1,1),(1,−1)(1,1),(1,-1) and the standard codomain basis. Its matrix is

[I]standard←new=(111−1).[I]_{\mathrm{standard}\leftarrow\mathrm{new}} =\begin{pmatrix}1&1\\1&-1\end{pmatrix}.

It is not the identity matrix because the coordinate systems differ. Input coordinates (1,0)(1,0) represent the vector (1,1)(1,1), whose standard output coordinates are (1,1)(1,1). The identity map has identity matrix only when both ends use the same basis.

Compositions require matching intermediate coordinates

If T:U→VT:U\to V has matrix AA and S:V→WS:V\to W has matrix BB, using the same basis for intermediate space VV, then S∘TS\circ T has matrix BABA.

If one output uses old coordinates and the next input expects new coordinates, a change-of-basis matrix must be inserted. Matching array dimensions alone does not establish matching coordinate meanings.

Exercises

ExerciseDoes preserving zero guarantee linearity?

Let T(x,y)=(x,y2)T(x,y)=(x,y^2). It satisfies T(0,0)=(0,0)T(0,0)=(0,0). Is it linear?

Solution

No. For v=(0,1)\mathbf v=(0,1), we have T(2v)=(0,4)T(2\mathbf v)=(0,4) but 2T(v)=(0,2)2T(\mathbf v)=(0,2). Scaling is not preserved. Mapping zero to zero is necessary for linearity, but not sufficient.

ExerciseMapping a function to a number

For real-valued functions on the reals, define E(f)=f(0)E(f)=f(0). Prove that EE is a linear map to R\mathbb R, and explain why it cannot uniquely recover the original function.

Solution

Pointwise operations give

E(af+bg)=af(0)+bg(0)=aE(f)+bE(g).E(af+bg)=a f(0)+b g(0)=aE(f)+bE(g).

Thus EE is linear. But the zero function and f(t)=tf(t)=t both map to zero, so the map is not injective. Preserving combinations and preserving all information are different requirements.

ExerciseWhat sampling discards

On P2\mathcal P_2, define E(p)=(p(0),p(1))E(p)=(p(0),p(1)). Find a kernel basis and determine whether it is surjective.

Solution

For p(t)=a+bt+ct2p(t)=a+bt+ct^2, the kernel conditions are a=0,b+c=0a=0,b+c=0, so t(t−1)t(t-1) is a basis. Any pair (u,v)(u,v) is reached by p(t)=u+(v−u)tp(t)=u+(v-u)t. The map is surjective but not injective.

ExerciseThe order of a basis matters

Swap the first two domain basis vectors while keeping the codomain basis fixed. What happens to the matrix?

Solution

Columns are the images of domain basis vectors, so swap the first two columns only. Output coordinates have not changed. This differs from simultaneously relabelling both ends of a graph adjacency matrix.