The first chapters represented responses by number columns, and linear systems exposed free directions. We now separate the operations from that particular representation. Why can whole functions be vectors, and what makes a coordinate description complete and unique?

Abstracting the operations from the objects

So far, pairs of numbers have recorded responses, and arrows have displayed those pairs. What makes a vector abstract? The next step is to set aside its appearance and identify the rules that made our reasoning possible.

Objects, representations, and coefficients are different

In the first chapter’s response model, the object is an entire response. After choosing measurements and units, we record it as a pair such as (2,1)(2,1). We then use coefficients a,ba,b to combine responses. These are three different roles.

Functions provide another example. Take two whole polynomials:

f(t)=2+t,g(t)=−1+2t.f(t)=2+t,\qquad g(t)=-1+2t.

Define addition and scaling point by point:

(af+bg)(t)=af(t)+bg(t).(af+bg)(t)=a f(t)+b g(t).

Then

2f(t)+g(t)=3+4t.2f(t)+g(t)=3+4t.

The result is an entire function, not its value at one particular input or a single point on its graph. Representing x+ytx+yt by its coefficient pair (x,y)(x,y) produces exactly the same combination rules as the earlier number pairs.

A representation must fit its space of objects. Two coefficients completely describe a polynomial of degree at most one, but two samples cannot determine an arbitrary function. The zero function and t(t−1)t(t-1) agree at t=0t=0 and t=1t=1 without being the same function.

This interactive figure needs JavaScript.

Adjust the coefficients, then switch between the function and coordinate-plane views. The numerical operations stay the same; their interpretation and display change. In the dependent setting, the second function is twice the first, so its constant and linear terms can no longer be controlled independently.

A vector space specifies the rules to preserve

DefinitionReal vector space

A real vector space consists of a set VV with addition and real scalar multiplication, both defined and closed within VV. For any u,v,w∈V\mathbf u,\mathbf v,\mathbf w\in V and a,b∈Ra,b\in\mathbb R, the following laws hold:

RuleExpression
Commutative additionu+v=v+u\mathbf u+\mathbf v=\mathbf v+\mathbf u
Associative addition(u+v)+w=u+(v+w)(\mathbf u+\mathbf v)+\mathbf w=\mathbf u+(\mathbf v+\mathbf w)
Zero vectorThere is 0\mathbf0 with v+0=v\mathbf v+\mathbf0=\mathbf v
Additive inversesEach v\mathbf v has −v-\mathbf v with v+(−v)=0\mathbf v+(-\mathbf v)=\mathbf0
Associative scalinga(bv)=(ab)va(b\mathbf v)=(ab)\mathbf v
Identity scalar1v=v1\mathbf v=\mathbf v
Distribution over vector additiona(u+v)=au+ava(\mathbf u+\mathbf v)=a\mathbf u+a\mathbf v
Distribution over scalar addition(a+b)v=av+bv(a+b)\mathbf v=a\mathbf v+b\mathbf v

Elements of VV are called vectors.

This is the abstraction: when the laws hold, the same reasoning applies without being rebuilt separately for arrows, number columns, and polynomials.

For real-valued functions on a fixed domain with pointwise operations, each law reduces to real arithmetic at every input. The zero vector is the zero function, and the additive inverse is the negative function. Polynomials of degree at most one are closed under these operations, so they too form a vector space.

In contrast, RGB colours restricted to components in [0,1][0,1] do not form a real vector space. Additive inverses can leave the range, and sums can exceed its upper bound. Embedding colours in R3\mathbb R^3 makes them convenient to compute with without making every result a valid colour. Clipping the result introduces a nonlinear operation.

The collection and its operations both matter

A vector space is not determined by the appearance of its elements. Real polynomials of degree at most two form P2\mathcal P_2. Polynomials of degree exactly two do not: adding t2t^2 and −t2-t^2 leaves that collection.

Real matrices of a fixed size also form a vector space under entrywise operations. An m×nm\times n matrix has mnmn freely chosen entries, but its size must be fixed: arbitrary matrices of different shapes cannot all be added together.

The axioms imply familiar consequences rather than assuming them separately. For example,

0v=(0+0)v=0v+0v.0\mathbf v=(0+0)\mathbf v=0\mathbf v+0\mathbf v.

Cancellation gives 0v=00\mathbf v=\mathbf0. Similarly, (−1)v+v=0v=0(-1)\mathbf v+\mathbf v=0\mathbf v=\mathbf0, so multiplication by −1-1 produces the additive inverse.

No coordinates appeared in these proofs. That is why they apply equally to columns, functions, and matrices.

For the coordinate examples below, retain the first chapter’s vectors:

u=(2,1),v=(−1,2),w=2u.\mathbf u=(2,1),\qquad\mathbf v=(-1,2),\qquad\mathbf w=2\mathbf u.

Why is it a subspace?

A subset WW of a real vector space VV is a linear subspace if it contains the zero vector and is closed under addition and real scalar multiplication. Closure means that these operations on elements of WW stay in WW. The remaining vector-space laws are inherited from VV.

PropositionA span is a subspace

The span of finitely many vectors satisfies these conditions.

Proof

Choosing all coefficients zero gives the zero vector. If x\mathbf x and y\mathbf y have coefficients aia_i and bib_i, then

x+y=∑i(ai+bi)vi,\mathbf x+\mathbf y=\sum_i(a_i+b_i)\mathbf v_i,

which is another combination of the original vectors. For any real scalar cc,

cx=∑i(cai)vic\mathbf x=\sum_i(ca_i)\mathbf v_i

also remains in the same set.

The span is the smallest subspace containing the given vectors: any subspace containing them must contain their scalar multiples and finite sums, hence all their linear combinations.

The line y=1y=1 is not a linear subspace because it misses the origin. Looking like a straight line is not sufficient. The singleton {0}\{\mathbf0\}, however, is a subspace.

Applying the subspace test

A nonempty subset WW of VV is a subspace precisely when

αu+βv∈Wfor all u,v∈W, α,β∈R.\alpha\mathbf u+\beta\mathbf v\in W \quad\text{for all }\mathbf u,\mathbf v\in W,\ \alpha,\beta\in\mathbb R.

Zero coefficients supply the zero vector, coefficients (1,1)(1,1) give addition, and (α,0)(\alpha,0) give scalar multiplication. Conversely, closure under the two operations gives closure under these combinations.

For example, W={(x,y,z):x−2y+z=0}W=\{(x,y,z):x-2y+z=0\} is a subspace. A linear combination of vectors satisfying the homogeneous equation still satisfies it. Solving the condition also gives

(x,y,z)=(2y−z,y,z)=y(2,1,0)+z(−1,0,1),(x,y,z)=(2y-z,y,z) =y(2,1,0)+z(-1,0,1),

so W=span⁡((2,1,0),(−1,0,1))W=\operatorname{span}((2,1,0),(-1,0,1)). The same equation provides both a test and a parameterization. Changing its right-hand side to one excludes zero, so the resulting set is not a subspace.

Intersection, sum, and union

The intersection U∩WU\cap W of two subspaces is a subspace: combinations of vectors belonging to both remain in both.

Their union usually is not. The two coordinate axes in a plane are subspaces, but their union excludes (1,0)+(0,1)(1,0)+(0,1). To combine subspaces while retaining linear combinations, use their sum:

U+W={u+w:u∈U,w∈W}.U+W=\{\mathbf u+\mathbf w:\mathbf u\in U,\mathbf w\in W\}.

It contains zero. Addition and scaling can be regrouped into a part in UU and a part in WW, proving closure. It is the smallest subspace containing U∪WU\cup W.

In three dimensions, let UU be the xyxy plane and WW the xzxz plane. Their intersection is the xx axis and their sum is all of space. Simply adding their dimensions would count the common direction twice.

If U∩W={0}U\cap W=\{\mathbf0\}, every vector in their sum has a unique decomposition. Two decompositions would give

u1−u2=w2−w1∈U∩W,\mathbf u_1-\mathbf u_2=\mathbf w_2-\mathbf w_1\in U\cap W,

forcing both sides to vanish. This is a direct sum, written U⊕WU\oplus W. The xyxy plane and zz axis provide an example; the two planes sharing the xx axis do not.

Reachability does not guarantee uniqueness

When w=2u\mathbf w=2\mathbf u, the target 2u2\mathbf u has multiple representations:

2u=2u+0w=0u+w.2\mathbf u=2\mathbf u+0\mathbf w=0\mathbf u+\mathbf w.

Subtracting gives 2u−w=02\mathbf u-\mathbf w=\mathbf0. Coefficients that are not all zero have cancelled to zero, revealing redundancy between the directions.

DefinitionLinear independence and dependence

Vectors are linearly independent if

a1v1+⋯+amvm=0a_1\mathbf v_1+\cdots+a_m\mathbf v_m=\mathbf0

forces every coefficient to be zero. They are linearly dependent if some coefficients that are not all zero satisfy the equation.

“Not all zero” does not mean “every coefficient is nonzero.” In a relation involving three or more vectors, some coefficients may be zero. Any family containing the zero vector is dependent: give that vector coefficient 11 and all others coefficient zero.

For a finite family, dependence is equivalent to at least one vector being expressible using the others. In a nontrivial zero combination, choose a nonzero coefficient, rearrange, and divide by it. Conversely, move such an expression to one side to obtain a nontrivial zero combination.

TheoremIndependence is exactly uniqueness of coefficients

A finite family is linearly independent if and only if every vector in its span has exactly one coefficient representation.

Proof

If coefficients aia_i and bib_i represent the same target, subtraction gives

∑i(ai−bi)vi=0.\sum_i(a_i-b_i)\mathbf v_i=\mathbf0.

Independence forces ai=bia_i=b_i for every ii.

Conversely, a nontrivial zero combination gives two representations of the zero vector: the all-zero coefficients and the nontrivial coefficients. Representations therefore cannot all be unique.

In fact, if the family is dependent, every reachable target has infinitely many real coefficient representations. Add any real multiple of a nontrivial zero combination to an existing coefficient list; the target does not change. Unreachable targets still have no representation.

A basis: coverage without redundancy

DefinitionBasis

A family is a basis of a space WW if it spans WW and is linearly independent.

The conditions have separate jobs: spanning gives existence of a representation, while independence gives uniqueness. After ordering the basis vectors, each vector has a unique coefficient column, its coordinates in that basis.

The standard vectors e1=(1,0)\mathbf e_1=(1,0) and e2=(0,1)\mathbf e_2=(0,1) form a basis of the plane. So do our u=(2,1)\mathbf u=(2,1) and v=(−1,2)\mathbf v=(-1,2): the coefficients found for an arbitrary target both exist and are unique. Basis vectors need not follow the coordinate axes, have unit length, or be perpendicular. For example, (1,0)(1,0) and (1,1)(1,1) also form a basis: an arbitrary (x,y)(x,y) has unique coefficients (x−y,y)(x-y,y).

The same vector p=(3,4)\mathbf p=(3,4) has standard coordinates (3,4)(3,4) and coordinates (2,1)(2,1) in the ordered basis (u,v)(\mathbf u,\mathbf v):

p=3e1+4e2=2u+v.\mathbf p=3\mathbf e_1+4\mathbf e_2 =2\mathbf u+\mathbf v.

The vector is unchanged; the reference used to describe it has changed. A later treatment of change of basis will formalize this relationship.

The number of vectors in a basis of a finite-dimensional space is its dimension. The next section proves that all bases have the same length. Here the plane has dimension two, while the line spanned by a nonzero vector has dimension one. The zero subspace has the empty family as a basis and dimension zero, with the empty linear combination defined to be zero.

Why dimension does not depend on the basis

The key fact is that an independent family cannot be longer than a finite spanning family of the same space.

Let u1,…,ur\mathbf u_1,\ldots,\mathbf u_r be independent and let w1,…,ws\mathbf w_1,\ldots,\mathbf w_s span the space. Express u1\mathbf u_1 using the wj\mathbf w_j. At least one coefficient is nonzero. Solve for that wj\mathbf w_j and replace it by u1\mathbf u_1 without losing the spanning property.

Next introduce u2\mathbf u_2. Independence prevents it from being a combination of u1\mathbf u_1 alone, so at least one remaining wj\mathbf w_j has a nonzero coefficient and can be replaced. Each independent vector consumes one position from the original spanning family. At most ss replacements are possible, giving r≤sr\le s.

Apply this inequality to two bases in both directions. Their lengths agree, making dimension a property of the space rather than the chosen basis.

How elimination measures dimension

Let AA have nn columns and rr pivots after elimination. Row operations left-multiply by an invertible matrix EE, so for every coefficient column c\mathbf c,

Ac=0⟺EAc=0.A\mathbf c=\mathbf0\quad\Longleftrightarrow\quad EA\mathbf c=\mathbf0.

Column relations are therefore preserved. Echelon-form pivot columns are independent and span its other columns, so the corresponding original columns form a basis of the original column space. Row operations may change the column space itself; the transformed columns are not generally a basis of the original space.

The column-space dimension, or rank, is therefore rr. The homogeneous system has n−rn-r free variables. Setting these to successive standard basis vectors produces n−rn-r independent solution directions spanning the nullspace. Hence

dim⁡ker⁡A+rank⁡(A)=n.\dim\ker A+\operatorname{rank}(A)=n.

The traffic system has five input coordinates and rank three, leaving a two-dimensional nullspace. Its nonhomogeneous solution set is a translate of that space; writing it as a particular solution plus directions does not make it a vector subspace.

Extracting a basis from generators

A spanning family may contain redundant vectors. Removing a vector expressible through the others preserves its span. Repeated removal from a finite spanning family eventually gives a basis. Conversely, adjoining a vector outside the span of an independent family preserves independence. In finite dimensions this extends the family to a basis.

For computation, place the generators in columns. For example,

A=(123011134)⟶R=(101011000).A=\begin{pmatrix}1&2&3\\0&1&1\\1&3&4\end{pmatrix} \quad\longrightarrow\quad R=\begin{pmatrix}1&0&1\\0&1&1\\0&0&0\end{pmatrix}.

The first two columns are pivot columns. Take the original columns

a1=(1,0,1)T,a2=(2,1,3)T\mathbf a_1=(1,0,1)^{\mathsf T},\qquad \mathbf a_2=(2,1,3)^{\mathsf T}

as a basis of the column space. Since a3=a1+a2\mathbf a_3=\mathbf a_1+\mathbf a_2, removing column three loses no reachable output.

For the row space, the nonzero echelon rows do form a basis: reversible row operations preserve the span of the rows themselves. Here these are (1,0,1)(1,0,1) and (0,1,1)(0,1,1). Distinguish this rule from selecting original pivot columns for the column space.

Computing a nullspace basis

The homogeneous system Ax=0A\mathbf x=\mathbf0 becomes

x1+x3=0,x2+x3=0.x_1+x_3=0,\qquad x_2+x_3=0.

Set x3=tx_3=t to obtain

x=t(−1,−1,1)T.\mathbf x=t(-1,-1,1)^{\mathsf T}.

Thus (−1,−1,1)T(-1,-1,1)^{\mathsf T} is a basis of the nullspace. A column-space basis describes reachable outputs; a nullspace basis describes input changes invisible in the output. They answer different questions.

With several free variables, set one to one and the rest to zero in turn, recovering the complete solution each time. These directions are independent because their free coordinates are standard basis vectors, and they span because they realize every possible free-coordinate choice.

Counting common directions once

Finite-dimensional subspaces satisfy

dim⁡(U+W)=dim⁡U+dim⁡W−dim⁡(U∩W).\dim(U+W)=\dim U+\dim W-\dim(U\cap W).

Choose a basis c1,…,ck\mathbf c_1,\ldots,\mathbf c_k of the intersection. Extend it to bases (ci,uj)(\mathbf c_i,\mathbf u_j) of UU and (ci,wℓ)(\mathbf c_i,\mathbf w_\ell) of WW. Combining these with only one copy of the intersection basis spans U+WU+W.

To check independence, suppose a combination vanishes. Move the wℓ\mathbf w_\ell terms to the other side; their sum belongs to both UU and WW, hence to the intersection. Independence of (ci,wℓ)(\mathbf c_i,\mathbf w_\ell) forces every wℓ\mathbf w_\ell coefficient to vanish. Independence of the basis of UU then forces the remaining coefficients to vanish. Counting the combined basis proves the formula.

Exercises

ExerciseIs pairwise nonparallel enough?

Are (1,0)(1,0), (0,1)(0,1), and (1,1)(1,1) independent? Every pair is nonparallel; does that contradict your answer?

Solution

They are dependent because (1,0)+(0,1)−(1,1)=0(1,0)+(0,1)-(1,1)=\mathbf0. Pairwise nonparallel vectors rule out scalar-multiple relations between pairs, but a larger family may still contain redundancy.

ExerciseIs containing the origin enough?

Which subsets of R2\mathbb R^2 are linear subspaces: the line y=2xy=2x, the line y=2x+1y=2x+1, and the set xy=0xy=0?

Solution

The first line is span⁡((1,2))\operatorname{span}((1,2)), hence a subspace. The second misses the origin. The last set is the union of the coordinate axes; it contains the origin but is not closed under addition, since (1,0)+(0,1)(1,0)+(0,1) lies outside it.

ExercisePolynomials can be vectors

In the real vector space of polynomials of degree at most one, prove that 1+t1+t and 1−t1-t form a basis.

Solution

For any a+bta+bt, compare coefficients in

α(1+t)+β(1−t)=a+bt.\alpha(1+t)+\beta(1-t)=a+bt.

The unique solution is

α=a+b2,β=a−b2.\alpha=\frac{a+b}{2},\qquad\beta=\frac{a-b}{2}.

Every polynomial has a unique representation, so these two polynomials form a basis. The vectors are polynomials; the scalars remain real numbers.

ExerciseColumn relations and the nullspace

For the chapter’s matrix AA, column three is the sum of the first two. Why does this not make the first two dependent? Give a nullspace basis.

Solution

The first two are not scalar multiples: their second entries are zero and one, with the first column nonzero. The three-column relation gives A(−1,−1,1)T=0A(-1,-1,1)^{\mathsf T}=0. Elimination leaves one free variable, so this vector is a nullspace basis.

ExerciseUnique decompositions

Let U=span⁡((1,0,0),(0,1,0))U=\operatorname{span}((1,0,0),(0,1,0)) and W=span⁡((1,0,1))W=\operatorname{span}((1,0,1)). Is R3=U⊕W\mathbb R^3=U\oplus W?

Solution

If t(1,0,1)t(1,0,1) belongs to UU, its third component forces t=0t=0. Also (x,y,z)=(x−z,y,0)+z(1,0,1)(x,y,z)=(x-z,y,0)+z(1,0,1), so every vector has a decomposition and the trivial intersection makes it unique.