A vector space specifies which objects can be added and scaled. It does not yet say how to compare lengths and directions. Questions about nearest vectors, error magnitude, or perpendicular directions require additional geometric structure.

An inner product turns two vectors into a scalar

The standard inner product on Rn\mathbb R^n is

⟨x,y⟩=xTy=∑i=1nxiyi.\langle \mathbf x,\mathbf y\rangle =\mathbf x^{\mathsf T}\mathbf y =\sum_{i=1}^n x_i y_i.

It is symmetric, linear in each input, and positive definite: ⟨x,x⟩≥0\langle\mathbf x,\mathbf x\rangle\ge0, with equality only for x=0\mathbf x=\mathbf0. Any real vector-space operation with these properties is an inner product. Choosing a different inner product changes how the space measures length and angle.

For positive weights wiw_i, for example,

⟨x,y⟩W=xTWy,W=diag⁡(w1,…,wn)\langle\mathbf x,\mathbf y\rangle_W =\mathbf x^{\mathsf T}W\mathbf y, \qquad W=\operatorname{diag}(w_1,\ldots,w_n)

makes errors in some coordinates more important. A negative weight can make ⟨x,x⟩W\langle\mathbf x,\mathbf x\rangle_W negative, so the resulting rule is not an inner product.

Norm, distance, and angle

An inner product induces a norm and a distance:

∥x∥=⟨x,x⟩,d(x,y)=∥x−y∥.\|\mathbf x\|=\sqrt{\langle\mathbf x,\mathbf x\rangle}, \qquad d(\mathbf x,\mathbf y)=\|\mathbf x-\mathbf y\|.

For nonzero vectors, define θ∈[0,π]\theta\in[0,\pi] by

cos⁡θ=⟨x,y⟩∥x∥ ∥y∥.\cos\theta =\frac{\langle\mathbf x,\mathbf y\rangle} {\|\mathbf x\|\,\|\mathbf y\|}.

Here cos⁡\cos is the familiar trigonometric function. The existing Cauchy–Schwarz theorem guarantees that the ratio lies in [−1,1][-1,1]. The zero vector has no direction, so its angle with another vector is undefined.

A positive inner product gives an acute angle, zero gives a right angle, and a negative value gives an obtuse angle. Cosine similarity retains directional agreement while discarding overall scale. It is therefore useful for comparing representations, but it is different from Euclidean distance.

Why Cauchy–Schwarz holds

For y≠0\mathbf y\ne\mathbf0, a squared length is nonnegative:

0≤∥x−⟨x,y⟩∥y∥2y∥2.0\le \left\|\mathbf x- \frac{\langle\mathbf x,\mathbf y\rangle}{\|\mathbf y\|^2} \mathbf y\right\|^2.

Expanding gives

∣⟨x,y⟩∣2≤∥x∥2∥y∥2.|\langle\mathbf x,\mathbf y\rangle|^2 \le \|\mathbf x\|^2\|\mathbf y\|^2.

Equality holds exactly when the vectors are linearly dependent. Applying this estimate to the expansion of ∥x+y∥2\|\mathbf x+\mathbf y\|^2 yields the triangle inequality

∥x+y∥≤∥x∥+∥y∥.\|\mathbf x+\mathbf y\| \le \|\mathbf x\|+\|\mathbf y\|.

The inequalities chapter proves the same result from sums of products. The present derivation shows how it creates geometry on a vector space.

If y=0\mathbf y=0, both sides of Cauchy–Schwarz vanish and the pair is dependent. For nonzero y\mathbf y, equality in the displayed squared norm means exactly that x\mathbf x is the specified scalar multiple of y\mathbf y, proving both directions of the equality condition.

Norm axioms

Triangle inequality and equality

Cauchy–Schwarz gives

∥x+y∥2=∥x∥2+2⟨x,y⟩+∥y∥2≤(∥x∥+∥y∥)2.\begin{aligned} \|\mathbf x+\mathbf y\|^2 &=\|\mathbf x\|^2+2\langle\mathbf x,\mathbf y\rangle+\|\mathbf y\|^2\\ &\le(\|\mathbf x\|+\|\mathbf y\|)^2. \end{aligned}

Both sides are nonnegative, so taking square roots preserves the inequality. For nonzero vectors, equality requires dependence and a nonnegative inner product, hence the same direction. A zero vector also gives equality. Positive definiteness supplies positivity, and bilinearity gives ∥cx∥=∣c∣∥x∥\|c\mathbf x\|=|c|\|\mathbf x\|, establishing all norm axioms.

Orthogonality and Pythagorean decomposition

Vectors are orthogonal when ⟨x,y⟩=0\langle\mathbf x,\mathbf y\rangle=0, written x⊥y\mathbf x\perp\mathbf y. Orthogonal vectors satisfy

∥x+y∥2=∥x∥2+∥y∥2.\|\mathbf x+\mathbf y\|^2 =\|\mathbf x\|^2+\|\mathbf y\|^2.

For a subspace WW, define

W⊥={v:⟨v,w⟩=0 for every w∈W}.W^\perp =\{\mathbf v:\langle\mathbf v,\mathbf w\rangle=0 \text{ for every }\mathbf w\in W\}.

The set W⊥W^\perp is a subspace. In a finite-dimensional inner-product space, every vector has a unique decomposition into one component in WW and another in W⊥W^\perp. The next chapter turns this existence statement into a projection algorithm.

Orthonormal bases expose coordinates directly

A basis is orthonormal when its vectors have unit length and are pairwise orthogonal. Then

v=∑i=1n⟨v,qi⟩qi.\mathbf v=\sum_{i=1}^n \langle\mathbf v,\mathbf q_i\rangle\mathbf q_i.

Coordinates in a general basis require solving a system. In an orthonormal basis, each coordinate is one inner product. If the columns of QQ form such a basis, then QTQ=IQ^{\mathsf T}Q=I and Q−1=QTQ^{-1}=Q^{\mathsf T}.

The matrix identity here uses standard Euclidean coordinates and the standard inner product. If the coordinate inner product is xTGy\mathbf x^{\mathsf T}G\mathbf y, the corresponding identity is QTGQ=IQ^{\mathsf T}GQ=I. For a square basis matrix it gives Q−1=QTGQ^{-1}=Q^{\mathsf T}G. For a rectangular orthonormal family, QTQ=IQ^{\mathsf T}Q=I gives a left inverse, not a two-sided inverse.

Closure of the complement and the coordinate formula

For u,v∈W⊥\mathbf u,\mathbf v\in W^\perp and every w∈W\mathbf w\in W, bilinearity gives ⟨au+bv,w⟩=0\langle a\mathbf u+b\mathbf v,\mathbf w\rangle=0. Zero also satisfies the condition, so the complement is a subspace.

If v=∑iciqi\mathbf v=\sum_i c_i\mathbf q_i, take the inner product with qj\mathbf q_j. Every other term vanishes, yielding cj=⟨v,qj⟩c_j=\langle\mathbf v,\mathbf q_j\rangle. Squared length is consequently ∑ici2\sum_i c_i^2, so orthogonal coordinate changes preserve length.

Gram–Schmidt constructs an orthonormal basis

Starting from linearly independent vectors v1,…,vk\mathbf v_1,\ldots,\mathbf v_k, successively remove components in directions already constructed:

uj=vj−∑i=1j−1⟨vj,ui⟩⟨ui,ui⟩ui,qj=uj∥uj∥.\mathbf u_j =\mathbf v_j- \sum_{i=1}^{j-1} \frac{\langle\mathbf v_j,\mathbf u_i\rangle} {\langle\mathbf u_i,\mathbf u_i\rangle}\mathbf u_i, \qquad \mathbf q_j=\frac{\mathbf u_j}{\|\mathbf u_j\|}.

Each step subtracts only a combination of preceding directions, so the span is unchanged. The remainder is orthogonal to every preceding direction. A zero remainder reveals that the original inputs were dependent and cannot be normalized into a basis.

For v1=(1,1)\mathbf v_1=(1,1) and v2=(1,0)\mathbf v_2=(1,0), first take q1=(1,1)/2\mathbf q_1=(1,1)/\sqrt2. Removing the q1\mathbf q_1 component from v2\mathbf v_2 leaves (1/2,−1/2)(1/2,-1/2), whose normalized version is q2=(1,−1)/2\mathbf q_2=(1,-1)/\sqrt2.

Induction for Gram–Schmidt

Assume the preceding nonzero ui\mathbf u_i are orthogonal. Take the inner product of the defining formula for uj\mathbf u_j with uℓ\mathbf u_\ell. Only the i=ℓi=\ell summand survives and exactly cancels ⟨vj,uℓ⟩\langle\mathbf v_j,\mathbf u_\ell\rangle. A zero remainder would place vj\mathbf v_j in the preceding span, contradicting independence. Old and new vectors express each other using preceding vectors, so every prefix span is preserved. This proves orthogonality, nonzero remainders, and completeness together.

Why the orthogonal complement fills the remaining space

Apply Gram–Schmidt to a basis of WW, obtaining q1,…,qk\mathbf q_1,\ldots,\mathbf q_k. For any v\mathbf v, set

w=∑i=1k⟨v,qi⟩qi,r=v−w.\mathbf w=\sum_{i=1}^k\langle\mathbf v,\mathbf q_i\rangle\mathbf q_i, \qquad \mathbf r=\mathbf v-\mathbf w.

Then ⟨r,qj⟩=0\langle\mathbf r,\mathbf q_j\rangle=0 for every jj, so r∈W⊥\mathbf r\in W^\perp. This proves existence. Two decompositions would have a difference in both WW and W⊥W^\perp. Such a vector is orthogonal to itself and must vanish, proving uniqueness. Hence

V=W⊕W⊥,dim⁡W+dim⁡W⊥=dim⁡V.V=W\oplus W^\perp, \qquad \dim W+\dim W^\perp=\dim V.

Finite dimension matters here. Infinite-dimensional spaces require an additional discussion of closed subspaces.

Geometry is not built into a coordinate column

The same column can have different lengths under different inner products. For x=(1,1)\mathbf x=(1,1), the standard norm is 2\sqrt2, while the norm induced by W=diag⁡(1,4)W=\operatorname{diag}(1,4) is 5\sqrt5. Coordinates therefore need a specified measurement rule.

The polar form of complex numbers writes a unit-circle point as (cos⁡θ,sin⁡θ)(\cos\theta,\sin\theta). Its standard inner product with (1,0)(1,0) is cos⁡θ\cos\theta. An angle defined by a different inner product need not equal the Euclidean angle drawn on paper. The angle formula must use norms induced by that same inner product.

Exercises

ExerciseAngle and the zero vector

Find the angle between (1,1,0)(1,1,0) and (1,0,1)(1,0,1). Explain why the same formula cannot be used with the zero vector.

Solution

The inner product is 11, and both norms are 2\sqrt2. Thus cos⁡θ=1/2\cos\theta=1/2 and θ=π/3\theta=\pi/3. A zero vector makes the denominator zero and has no direction.

ExerciseCheck a weighted inner product

Determine whether ⟨x,y⟩=2x1y1+3x2y2\langle\mathbf x,\mathbf y\rangle=2x_1y_1+3x_2y_2 is an inner product on R2\mathbb R^2.

Solution

It is symmetric and bilinear, while 2x12+3x22≥02x_1^2+3x_2^2\ge0 with equality only when both coordinates vanish. It is therefore an inner product.

ExercisePerform one orthogonalization

Apply Gram–Schmidt to (1,0,1)(1,0,1) and (1,1,0)(1,1,0).

Solution

Take q1=(1,0,1)/2\mathbf q_1=(1,0,1)/\sqrt2. The second remainder is (1,1,0)−12(1,0,1)=(1/2,1,−1/2)(1,1,0)-\tfrac12(1,0,1)=(1/2,1,-1/2), which normalizes to (1,2,−1)/6(1,2,-1)/\sqrt6.