From response vectors to a matrix

Vectors and linear combinations described each input by its complete response. With many inputs, we collect these response vectors as columns:

A=(2−112).A=\begin{pmatrix}2&-1\\1&2\end{pmatrix}.

Column jj records the response to one unit of input jj. Row ii records how output ii depends on all inputs. These are two views of the same model.

A real matrix A∈Rm×nA\in\mathbb R^{m\times n} has mm rows and nn columns; its entry aija_{ij} occupies row ii, column jj. When the matrix represents a response rule, rows count outputs and columns count inputs.

Matrix-vector multiplication combines the columns

DefinitionMatrix-vector product

If the columns of AA are a1,…,an\mathbf a_1,\ldots,\mathbf a_n and x∈Rn\mathbf x\in\mathbb R^n, define

Ax=x1a1+⋯+xnan.A\mathbf x=x_1\mathbf a_1+\cdots+x_n\mathbf a_n.

The result belongs to Rm\mathbb R^m. Componentwise,

(Ax)i=∑j=1naijxj.(A\mathbf x)_i=\sum_{j=1}^n a_{ij}x_j.

For our example,

A(21)=2(21)+(−12)=(34).A\begin{pmatrix}2\\1\end{pmatrix} =2\begin{pmatrix}2\\1\end{pmatrix} +\begin{pmatrix}-1\\2\end{pmatrix} =\begin{pmatrix}3\\4\end{pmatrix}.

Dimensions must match because each column needs a coefficient. The matrix need not be square: two parameters may produce three observations.

Check dimensions before performing an operation

A matrix’s shape specifies which inputs and operations make sense. For A∈Rm×nA\in\mathbb R^{m\times n}, the product AxA\mathbf x requires nn input components and produces mm output components. Addition requires two matrices of exactly the same shape; multiplication requires only the shared intermediate dimension to match.

If AA is 2×32\times3 and BB is 3×43\times4, then ABAB is 2×42\times4, while BABA is undefined. Even when both matrices are square and both orders exist, equality is a separate question.

Equality of matrices requires equal shapes and equal corresponding entries. Matrix factors cannot be cancelled as casually as nonzero real numbers.

Zero, diagonal, and triangular matrices

The zero matrix has all entries zero and maps every input to zero. Its dimensions are usually inferred from context.

A diagonal matrix can have nonzero entries only on its main diagonal:

D=(d1000d2000d3),Dx=(d1x1,d2x2,d3x3)T.D=\begin{pmatrix}d_1&0&0\\0&d_2&0\\0&0&d_3\end{pmatrix}, \qquad D\mathbf x=(d_1x_1,d_2x_2,d_3x_3)^{\mathsf T}.

It scales coordinates separately, without mixing them. If di=0d_i=0, information about input coordinate ii disappears entirely.

An upper triangular matrix has zeros below the main diagonal; a lower triangular matrix has zeros above it. Their equations can be solved sequentially. In an upper triangular system, the last equation involves only the last unknown; after solving it, work upward. LU factorization exploits this structure.

Interpolation is linear in the coefficients

Suppose we want a polynomial of degree at most two whose values at the inputs 1,2,31,2,3 are 6,7,56,7,5, respectively [1][1] X. Yang, “ENG1005 Week 3: Interpolation and Fitting, Personal Workshop Solutions,” 2024. Personal solutions to Monash ENG1005 workshop problems; source snapshot 77ebe58de2fea53d62533d6dd23caa16108ed109. Repository access may be restricted.. https://github.com/Eryc123Y/ENG1005-2024S2/blob/77ebe58de2fea53d62533d6dd23caa16108ed109/Source%20Code/W3.tex. Write

p(t)=αt2+βt+γ.p(t)=\alpha t^2+\beta t+\gamma.

Evaluating at the three inputs gives

(111421931)⏟V(αβγ)⏟c=(675)⏟y.\underbrace{\begin{pmatrix}1&1&1\\4&2&1\\9&3&1\end{pmatrix}}_{V} \underbrace{\begin{pmatrix}\alpha\\\beta\\\gamma\end{pmatrix}}_{\mathbf c} =\underbrace{\begin{pmatrix}6\\7\\5\end{pmatrix}}_{\mathbf y}.

The columns sample t2t^2, tt, and 11, respectively. Their coefficients combine the sample vectors into the target data.

Linearity concerns the unknown coefficients. A squared input variable does not prevent the coefficient-to-observation rule from being linear. Fixed sample locations determine VV; new observations change only the right-hand side.

The coefficients are

c=(−3/211/22).\mathbf c=\begin{pmatrix}-3/2\\11/2\\2\end{pmatrix}.

Substitution yields 6,7,56,7,5. Linear systems and LU derives these coefficients by elimination.

Matrix multiplication composes two steps

Suppose B∈Rn×pB\in\mathbb R^{n\times p} first produces nn intermediate quantities, then A∈Rm×nA\in\mathbb R^{m\times n} produces the outputs. We want

(AB)x=A(Bx).(AB)\mathbf x=A(B\mathbf x).

Apply this requirement to each standard basis vector. Column jj of the product must be AA applied to column jj of BB, so

(AB)ij=∑k=1naikbkj.(AB)_{ij}=\sum_{k=1}^n a_{ik}b_{kj}.

The formula adds the contributions through every intermediate coordinate. The product is m×pm\times p; the shared dimension is summed over.

ExampleShearing and scaling in different orders

Take

S=(1101),D=(2001).S=\begin{pmatrix}1&1\\0&1\end{pmatrix},\qquad D=\begin{pmatrix}2&0\\0&1\end{pmatrix}.

Then S(x,y)=(x+y,y)S(x,y)=(x+y,y) and D(x,y)=(2x,y)D(x,y)=(2x,y), while

DS=(2201),SD=(2101).DS=\begin{pmatrix}2&2\\0&1\end{pmatrix},\qquad SD=\begin{pmatrix}2&1\\0&1\end{pmatrix}.

The input (0,1)(0,1) gives (2,1)(2,1) in the first case and (1,1)(1,1) in the second. Multiplication generally does not commute because order changes the result. In DSDS, the rightmost matrix acts first.

Grouping three steps differently leaves their order intact. Matrix multiplication is associative; entrywise, either grouping gives the finite sum

∑j,kaijbjkckℓ.\sum_{j,k}a_{ij}b_{jk}c_{k\ell}.

Three ways to read a product

For A∈Rm×nA\in\mathbb R^{m\times n} and B∈Rn×pB\in\mathbb R^{n\times p}, the entrywise formula is only one useful organization of the computation.

By columns: column jj of ABAB is AbjA\mathbf b_j. This reads the product as one transformation applied to several inputs.

By rows: row ii of ABAB is row ii of AA multiplied by BB. It combines intermediate output rules into the rule for final output ii.

By intermediate coordinates: if ak\mathbf a_k is column kk of AA and rk\mathbf r_k is row kk of BB, then

AB=∑k=1nakrk.AB=\sum_{k=1}^n\mathbf a_k\mathbf r_k.

Each m×1m\times1 column times a 1×p1\times p row is an m×pm\times p outer product. Its entry aikbkja_{ik}b_{kj} records the contribution through intermediate coordinate kk from input jj to output ii.

For example, take

A=(1234),B=(5678).A=\begin{pmatrix}1&2\\3&4\end{pmatrix},\qquad B=\begin{pmatrix}5&6\\7&8\end{pmatrix}.

Pair the first column with the first row, then the second column with the second row:

AB=(13)(56)+(24)(78).AB=\begin{pmatrix}1\\3\end{pmatrix}\begin{pmatrix}5&6\end{pmatrix} +\begin{pmatrix}2\\4\end{pmatrix}\begin{pmatrix}7&8\end{pmatrix}.

Compute the two outer products and add their entries:

AB=(561518)+(14162832)=(19224350).AB=\begin{pmatrix}5&6\\15&18\end{pmatrix} +\begin{pmatrix}14&16\\28&32\end{pmatrix} =\begin{pmatrix}19&22\\43&50\end{pmatrix}.

All three views describe the same multiplication, organizing its sum around different structures.

Block multiplication preserves dimension matching

Split an input into two groups and partition the matrix’s columns accordingly:

(A1A2)(x1x2)=A1x1+A2x2.\begin{pmatrix}A_1&A_2\end{pmatrix} \begin{pmatrix}\mathbf x_1\\\mathbf x_2\end{pmatrix} =A_1\mathbf x_1+A_2\mathbf x_2.

More generally, compatible blocks satisfy

(ABCD)(xy)=(Ax+ByCx+Dy).\begin{pmatrix}A&B\\C&D\end{pmatrix} \begin{pmatrix}\mathbf x\\\mathbf y\end{pmatrix} = \begin{pmatrix}A\mathbf x+B\mathbf y\\C\mathbf x+D\mathbf y\end{pmatrix}.

Blocks need not have equal sizes or be square. Their shared boundaries must match. This is ordinary multiplication organized around groups of variables.

Addition, identity, and inverses

Matrices of equal size add and scale entrywise. The column definition gives

(A+B)x=Ax+Bx,A(x+z)=Ax+Az.(A+B)\mathbf x=A\mathbf x+B\mathbf x,\qquad A(\mathbf x+\mathbf z)=A\mathbf x+A\mathbf z.

The identity InI_n has ones on its diagonal and zeros elsewhere. Its columns are the standard basis, so Inx=xI_n\mathbf x=\mathbf x.

A square matrix AA is invertible if a matrix CC satisfies both CA=AC=ICA=AC=I; write C=A−1C=A^{-1}. It recovers inputs from outputs. For example,

(2−112)−1=15(21−12).\begin{pmatrix}2&-1\\1&2\end{pmatrix}^{-1} =\frac15\begin{pmatrix}2&1\\-1&2\end{pmatrix}.

Multiplication in either order verifies this identity. When AA is invertible, every equation Ax=bA\mathbf x=\mathbf b has the unique solution A−1bA^{-1}\mathbf b.

Conversely, suppose a square matrix gives a unique solution for every right-hand side. Solve Acj=ejA\mathbf c_j=\mathbf e_j and collect the solutions in CC, yielding AC=IAC=I. Since the homogeneous equation has only the zero solution, A(CA−I)=0A(CA-I)=0 also implies CA=ICA=I, column by column. Thus invertibility means that every target is reachable with unique coefficients. In computation, elimination usually solves the system without constructing the entire inverse.

Cancellation requires an invertible factor

Take

A=(1000),B=(0001).A=\begin{pmatrix}1&0\\0&0\end{pmatrix},\qquad B=\begin{pmatrix}0&0\\0&1\end{pmatrix}.

Both matrices are nonzero, yet AB=0AB=0. A zero product therefore need not have a zero factor. Likewise, AX=AYAX=AY need not imply X=YX=Y: AA may discard a nonzero difference X−YX-Y.

If AA is invertible, multiplication on the left by A−1A^{-1} does justify cancellation. To cancel a right factor, multiply its inverse on the right instead. The side matters because multiplication does not commute.

The inverse is unique. If C,DC,D are both inverses of AA, then

C=C(AD)=(CA)D=D.C=C(AD)=(CA)D=D.

For a two-by-two matrix, direct multiplication gives a useful formula when ad−bc≠0ad-bc\ne0:

(abcd)−1=1ad−bc(d−b−ca).\begin{pmatrix}a&b\\c&d\end{pmatrix}^{-1} =\frac1{ad-bc}\begin{pmatrix}d&-b\\-c&a\end{pmatrix}.

The off-diagonal entries cancel and both diagonal entries become ad−bcad-bc. This can be checked as a multiplication identity before studying determinants.

If ad−bc=0ad-bc=0, no inverse exists. When the first row is nonzero, the nonzero vector (b,−a)T(b,-a)^{\mathsf T} maps to zero. If only the second row is nonzero, use (d,−c)T(d,-c)^{\mathsf T} instead. If both rows vanish, every vector maps to zero. An invertible matrix cannot annihilate a nonzero input: applying its inverse would force that input to equal zero.

Transposition exchanges indices

The transpose swaps rows and columns:

(AT)ij=aji.(A^{\mathsf T})_{ij}=a_{ji}.

It changes an m×nm\times n matrix into an n×mn\times m matrix and is generally different from the inverse. The product formula gives

(AB)T=BTAT,(AB)^{\mathsf T}=B^{\mathsf T}A^{\mathsf T},

because entry (i,j)(i,j) on the right is ∑kbkiajk\sum_k b_{ki}a_{jk}, exactly entry (j,i)(j,i) of ABAB. Matrix as Graph will use this distinction when interpreting directed edges.

Exercises

ExerciseReading a rectangular matrix

For

B=(101101),B=\begin{pmatrix}1&0\\1&1\\0&1\end{pmatrix},

compute B(2,3)TB(2,3)^{\mathsf T}. Which input determines the third output?

Solution

The output is (2,5,3)T(2,5,3)^{\mathsf T}. The third row is (0,1)(0,1), so only the second input contributes. Three outputs do not require three inputs.

ExerciseUndoing a composition

If square matrices A,BA,B are invertible, prove (AB)−1=B−1A−1(AB)^{-1}=B^{-1}A^{-1}.

Solution

Associativity gives (AB)(B−1A−1)=AIA−1=I(AB)(B^{-1}A^{-1})=AIA^{-1}=I; the reverse product is also II. Undo the last action, AA, before undoing BB.

ExerciseTranspose versus inverse

For D=diag⁡(2,3)D=\operatorname{diag}(2,3), find its transpose and inverse and explain their different roles.

Solution

The transpose is DD; the inverse is diag⁡(1/2,1/3)\operatorname{diag}(1/2,1/3). Transposition exchanges indices; inversion undoes the map. Symmetry alone does not make them equal.

ExerciseRecovering a matrix from unit inputs

A matrix maps (1,0)(1,0) to (1,2,3)(1,2,3) and (0,1)(0,1) to (−1,0,2)(-1,0,2). Find its output for (2,−1)(2,-1).

Solution

The known outputs are its columns, giving 2(1,2,3)−(−1,0,2)=(3,4,4)2(1,2,3)-(-1,0,2)=(3,4,4). The matrix has size 3×23\times2.

References

  1. [1] X. Yang, “ENG1005 Week 3: Interpolation and Fitting, Personal Workshop Solutions,” 2024. Personal solutions to Monash ENG1005 workshop problems; source snapshot 77ebe58de2fea53d62533d6dd23caa16108ed109. Repository access may be restricted.. https://github.com/Eryc123Y/ENG1005-2024S2/blob/77ebe58de2fea53d62533d6dd23caa16108ed109/Source%20Code/W3.tex ↩