Expectation and Variance

Expectation and variance are two fundamental concepts in probability theory that describe the central tendency and spread of a random variable's distribution.

Throughout, identities between finite expectations assume integrability: E∣X∣<∞E|X|<\infty, and variance/covariance formulas assume finite second moments. A nonnegative expectation may be +∞+\infty, but subtraction of infinities is not allowed. The continuous formulas below use measurable functions and the usual integral rules.

Expected Value (Mean)

The expected value, also known as the mean or expectation, represents the average value of a random variable over many trials.

For Discrete Random Variables

E[X]=μX=∑xx⋅pX(x)\mathbb{E}[X] = \mu_X = \sum_{x} x \cdot p_X(x)

where:

  • pX(x)p_X(x) is the probability mass function (PMF)
  • The sum is taken over all possible values of XX

Properties:

  • Linearity and Homogeneity: E[aX+b]=aE[X]+b\mathbb{E}[aX + b] = a\mathbb{E}[X] + b
  • For two random variables: E[X+Y]=E[X]+E[Y]\mathbb{E}[X + Y] = \mathbb{E}[X] + \mathbb{E}[Y]
  • For independent random variables: E[XY]=E[X]E[Y]\mathbb{E}[XY] = \mathbb{E}[X]\mathbb{E}[Y]

Proofs of Key Properties

ProofLinearity

For a,b∈Ra, b \in \mathbb{R}

E[aX+b]=∑x(ax+b)⋅pX(x)=a∑xx⋅pX(x)+b∑xpX(x)=aE[X]+b\begin{aligned} \mathbb{E}[aX + b] &= \sum_{x} (ax + b) \cdot p_X(x) \\ &= a\sum_{x} x \cdot p_X(x) + b\sum_{x} p_X(x) \\ &= a\mathbb{E}[X] + b \end{aligned}
ProofAdditivity
E[X+Y]=∑x∑y(x+y)⋅pX,Y(x,y)=∑x∑yx⋅pX,Y(x,y)+∑x∑yy⋅pX,Y(x,y)=E[X]+E[Y]\begin{aligned} \mathbb{E}[X + Y] &= \sum_{x}\sum_{y} (x + y) \cdot p_{X,Y}(x,y) \\ &= \sum_{x}\sum_{y} x \cdot p_{X,Y}(x,y) + \sum_{x}\sum_{y} y \cdot p_{X,Y}(x,y) \\ &= \mathbb{E}[X] + \mathbb{E}[Y] \end{aligned}
ProofProduct for Independent Variables

If XX and YY are independent, then pX,Y(x,y)=pX(x)pY(y)p_{X,Y}(x,y) = p_X(x)p_Y(y), so:

E[XY]=∑x∑yxy⋅pX,Y(x,y)=∑x∑yxy⋅pX(x)pY(y)=(∑xxpX(x))(∑yypY(y))=E[X]E[Y]\begin{aligned} \mathbb{E}[XY] &= \sum_{x}\sum_{y} xy \cdot p_{X,Y}(x,y) \\ &= \sum_{x}\sum_{y} xy \cdot p_X(x)p_Y(y) \\ &= \left(\sum_{x} x p_X(x)\right)\left(\sum_{y} y p_Y(y)\right) \\ &= \mathbb{E}[X]\mathbb{E}[Y] \end{aligned}

For Continuous Random Variables

E[X]=μX=∫−∞∞x⋅fX(x)dx\mathbb{E}[X] = \mu_X = \int_{-\infty}^{\infty} x \cdot f_X(x)dx

where:

  • fX(x)f_X(x) is the probability density function (PDF)

Variance

Variance measures how much the values of a random variable deviate from its mean.

DefinitionVariance
V(X)=σX2=E[(X−μX)2]=E[X2]−(E[X])2\mathbb{V}(X) = \sigma_X^2 = \mathbb{E}[(X - \mu_X)^2] = \mathbb{E}[X^2] - (\mathbb{E}[X])^2

Same mean, different variance. These illustrative distributions have the same mean but different spreads, hence different variances.

Same mean, different variance. These illustrative distributions have the same mean but different spreads, hence different variances.

These illustrative distributions have the same mean but different spreads, hence different variances.

For Discrete Random Variables

V(X)=∑x(x−μX)2⋅pX(x)\mathbb{V}(X) = \sum_{x} (x - \mu_X)^2 \cdot p_X(x)

For Continuous Random Variables

V(X)=∫−∞∞(x−μX)2⋅fX(x)dx\mathbb{V}(X) = \int_{-\infty}^{\infty} (x - \mu_X)^2 \cdot f_X(x)dx

Standard Deviation

The standard deviation is the square root of the variance:

σX=V(X)\sigma_X = \sqrt{\mathbb{V}(X)}

Properties of Variance

  • V(X)≥0\mathbb{V}(X) \geq 0
  • V(a)=0\mathbb{V}(a) = 0 for any constant aa
  • V(aX)=a2V(X)\mathbb{V}(aX) = a^2 \mathbb{V}(X)
  • V(X+a)=V(X)\mathbb{V}(X + a) = \mathbb{V}(X)
  • For independent random variables: V(X+Y)=V(X)+V(Y)\mathbb{V}(X + Y) = \mathbb{V}(X) + \mathbb{V}(Y)

Proofs of Variance Properties

ProofScaling

For a∈Ra \in \mathbb{R}

V(aX)=E[(aX−E[aX])2]=E[(aX−aE[X])2]=E[a2(X−E[X])2]=a2E[(X−E[X])2]=a2V(X)\begin{aligned} \mathbb{V}(aX) &= \mathbb{E}[(aX - \mathbb{E}[aX])^2] \\ &= \mathbb{E}[(aX - a\mathbb{E}[X])^2] \\ &= \mathbb{E}[a^2(X - \mathbb{E}[X])^2] \\ &= a^2\mathbb{E}[(X - \mathbb{E}[X])^2] \\ &= a^2\mathbb{V}(X) \end{aligned}
ProofShift Invariance
V(X+a)=E[(X+a−E[X+a])2]=E[(X+a−E[X]−a)2]=E[(X−E[X])2]=V(X)\begin{aligned} \mathbb{V}(X + a) &= \mathbb{E}[(X + a - \mathbb{E}[X + a])^2] \\ &= \mathbb{E}[(X + a - \mathbb{E}[X] - a)^2] \\ &= \mathbb{E}[(X - \mathbb{E}[X])^2] \\ &= \mathbb{V}(X) \end{aligned}
ProofAdditivity for Independent Variables

If XX and YY are independent:

V(X+Y)=E[(X+Y)2]−(E[X+Y])2=E[X2+2XY+Y2]−(E[X]+E[Y])2=E[X2]+2E[X]E[Y]+E[Y2]−E[X]2−2E[X]E[Y]−E[Y]2=(E[X2]−E[X]2)+(E[Y2]−E[Y]2)=V(X)+V(Y)\begin{aligned} \mathbb{V}(X + Y) &= \mathbb{E}[(X + Y)^2] - (\mathbb{E}[X + Y])^2 \\ &= \mathbb{E}[X^2 + 2XY + Y^2] - (\mathbb{E}[X] + \mathbb{E}[Y])^2 \\ &= \mathbb{E}[X^2] + 2\mathbb{E}[X]\mathbb{E}[Y] + \mathbb{E}[Y^2] - \mathbb{E}[X]^2 - 2\mathbb{E}[X]\mathbb{E}[Y] - \mathbb{E}[Y]^2 \\ &= (\mathbb{E}[X^2] - \mathbb{E}[X]^2) + (\mathbb{E}[Y^2] - \mathbb{E}[Y]^2) \\ &= \mathbb{V}(X) + \mathbb{V}(Y) \end{aligned}
ProofSecond-moment identity and zero variance

Put μ=E[X]\mu=E[X]. Expansion and linearity give E[(X−μ)2]=E[X2]−2μE[X]+μ2=E[X2]−μ2E[(X-\mu)^2]=E[X^2]-2\mu E[X]+\mu^2=E[X^2]-\mu^2. Nonnegativity follows from the squared integrand. A constant has zero centered square. Conversely, if the expectation of that square is zero, each event {∣X−μ∣≥1/k}\{|X-\mu|\ge1/k\} has probability zero by the pointwise lower bound (X−μ)2≥k−2(X-\mu)^2\ge k^{-2} on the event. Their countable union is {X≠μ}\{X\ne\mu\}, so zero variance means X=μX=\mu almost surely.

Examples

ExampleDiscrete Case (Die Roll)

For a fair six-sided die:

  • PMF: pX(x)=16p_X(x) = \frac{1}{6} for x∈{1,2,3,4,5,6}x \in \{1, 2, 3, 4, 5, 6\}

Expected Value:

E[X]=∑x=16x⋅16=1+2+3+4+5+66=216=3.5\mathbb{E}[X] = \sum_{x=1}^{6} x \cdot \frac{1}{6} = \frac{1+2+3+4+5+6}{6} = \frac{21}{6} = 3.5

Variance:

E[X2]=∑x=16x2⋅16=1+4+9+16+25+366=916\mathbb{E}[X^2] = \sum_{x=1}^{6} x^2 \cdot \frac{1}{6} = \frac{1+4+9+16+25+36}{6} = \frac{91}{6}V(X)=E[X2]−(E[X])2=916−(3.5)2=916−494=182−14712=3512≈2.92\mathbb{V}(X) = \mathbb{E}[X^2] - (\mathbb{E}[X])^2 = \frac{91}{6} - (3.5)^2 = \frac{91}{6} - \frac{49}{4} = \frac{182 - 147}{12} = \frac{35}{12} \approx 2.92
ExampleContinuous Case (Normal Distribution)

For X∼N(μ,σ2)X \sim N(\mu, \sigma^2):

  • PDF: fX(x)=1σ2πe−(x−μ)22σ2f_X(x) = \frac{1}{\sigma\sqrt{2\pi}} e^{-\frac{(x-\mu)^2}{2\sigma^2}}

Expected Value: E[X]=μ\mathbb{E}[X] = \mu

Variance: V(X)=σ2\mathbb{V}(X) = \sigma^2

ExampleContinuous Case (Uniform Distribution)

For X∼U(a,b)X \sim U(a, b):

  • PDF: fX(x)=1b−af_X(x) = \frac{1}{b-a} for a≤x≤ba \leq x \leq b

Expected Value:

E[X]=∫abx⋅1b−adx=a+b2\mathbb{E}[X] = \int_a^b x \cdot \frac{1}{b-a} dx = \frac{a+b}{2}

Variance:

V(X)=∫ab(x−a+b2)2⋅1b−adx=(b−a)212\mathbb{V}(X) = \int_a^b \left(x - \frac{a+b}{2}\right)^2 \cdot \frac{1}{b-a} dx = \frac{(b-a)^2}{12}

Expectation of Functions of Random Variables

When we apply a function to a random variable, we obtain a new random variable. Computing the expectation of this new random variable is a fundamental problem in probability theory.

Law of the Unconscious Statistician (LOTUS)

The core principle for computing expectations of functions of random variables is the Law of the Unconscious Statistician (LOTUS). This law states that to compute E[g(X)]\mathbb{E}[g(X)], we don't need to first find the distribution of g(X)g(X). Instead, we can work directly with the original distribution of XX.

Computation Formula

For a function g:R→Rg: \mathbb{R} \to \mathbb{R} and random variable XX, the expectation of g(X)g(X) is:

E[g(X)]={∑xg(x)⋅pX(x)(discrete)∫−∞∞g(x)⋅fX(x)dx(continuous)\mathbb{E}[g(X)] = \begin{cases} \sum_{x} g(x) \cdot p_X(x) & \text{(discrete)} \\ \int_{-\infty}^{\infty} g(x) \cdot f_X(x) dx & \text{(continuous)} \end{cases}

Important Properties

  1. Linearity: E[a⋅g(X)+b⋅h(X)]=aE[g(X)]+bE[h(X)]\mathbb{E}[a \cdot g(X) + b \cdot h(X)] = a\mathbb{E}[g(X)] + b\mathbb{E}[h(X)]
  2. Monotonicity: If g(x)≤h(x)g(x) \leq h(x) for all xx, then E[g(X)]≤E[h(X)]\mathbb{E}[g(X)] \leq \mathbb{E}[h(X)]
ProofLOTUS and monotonicity

In the discrete case, group values by y=g(x)y=g(x). Since P(g(X)=y)=∑x:g(x)=ypX(x)P(g(X)=y)=\sum_{x:g(x)=y}p_X(x), substitution in E[g(X)]E[g(X)] and regrouping gives ∑xg(x)pX(x)\sum_x g(x)p_X(x). Regrouping is valid for nonnegative terms or an absolutely convergent sum.

For a density, first take an indicator g=1Bg=\mathbf1_B: both sides equal P(X∈B)=∫BfXP(X\in B)=\int_B f_X. Linearity gives the claim for simple functions, finite linear combinations of indicators. Approximate a nonnegative measurable gg increasingly by simple functions and apply monotone convergence; for integrable signed gg, subtract positive and negative parts. This proof uses the monotone convergence theorem as an integration prerequisite, whose general proof is not yet in the calculus notes. The same argument works for joint distributions. Finally h−g≥0h-g\ge0 implies E[h(X)]−E[g(X)]≥0E[h(X)]-E[g(X)]\ge0, proving monotonicity whenever the difference is defined.

Application Examples

ExampleExpectation of Square Function

For any random variable XX, computing E[X2]\mathbb{E}[X^2]:

  • Discrete case: E[X2]=∑xx2⋅pX(x)\mathbb{E}[X^2] = \sum_{x} x^2 \cdot p_X(x)
  • Continuous case: E[X2]=∫−∞∞x2⋅fX(x)dx\mathbb{E}[X^2] = \int_{-\infty}^{\infty} x^2 \cdot f_X(x) dx

This result gives the standard variance formula: V(X)=E[X2]−(E[X])2\mathbb{V}(X) = \mathbb{E}[X^2] - (\mathbb{E}[X])^2

ExampleExpectation of Exponential Function

For any random variable XX, computing E[etX]\mathbb{E}[e^{tX}]:

  • Discrete case: E[etX]=∑xetx⋅pX(x)\mathbb{E}[e^{tX}] = \sum_{x} e^{tx} \cdot p_X(x)
  • Continuous case: E[etX]=∫−∞∞etx⋅fX(x)dx\mathbb{E}[e^{tX}] = \int_{-\infty}^{\infty} e^{tx} \cdot f_X(x) dx

This is the definition of the moment generating function, which has wide applications in probability theory.

Numerical Estimation Methods

When functions are complex or distributions are non-standard, analytical solutions may be difficult to obtain. In such cases, we can use Taylor series approximation for numerical estimation.

Taylor Series Approximation Method

For a random variable XX with mean μ\mu and variance σ2\sigma^2, the expectation and variance of f(X)f(X) can be approximated using Taylor expansion.

ProofApproximation Derivation for Expectation
  1. Perform second-order Taylor expansion of f(X)f(X) around μ\mu:
f(X)=f(μ)+f′(μ)(X−μ)+f′′(μ)2(X−μ)2+R2f(X) = f(\mu) + f'(\mu)(X-\mu) + \frac{f''(\mu)}{2}(X-\mu)^2 + R_2

where R2R_2 is the remainder term.

  1. Take expectation of both sides:
E[f(X)]=E[f(μ)]+E[f′(μ)(X−μ)]+E[f′′(μ)2(X−μ)2]+E[R2]\mathbb{E}[f(X)] = \mathbb{E}[f(\mu)] + \mathbb{E}[f'(\mu)(X-\mu)] + \mathbb{E}\left[\frac{f''(\mu)}{2}(X-\mu)^2\right] + \mathbb{E}[R_2]
  1. Since f(μ)f(\mu), f′(μ)f'(\mu), and f′′(μ)f''(\mu) are constants:
E[f(X)]=f(μ)+f′(μ)E[X−μ]+f′′(μ)2E[(X−μ)2]+E[R2]\mathbb{E}[f(X)] = f(\mu) + f'(\mu)\mathbb{E}[X-\mu] + \frac{f''(\mu)}{2}\mathbb{E}[(X-\mu)^2] + \mathbb{E}[R_2]
  1. Using E[X−μ]=0\mathbb{E}[X-\mu] = 0 and E[(X−μ)2]=σ2\mathbb{E}[(X-\mu)^2] = \sigma^2, and ignoring higher-order remainder terms:
E[f(X)]≈f(μ)+f′′(μ)2σ2\mathbb{E}[f(X)] \approx f(\mu) + \frac{f''(\mu)}{2}\sigma^2
ProofApproximation Derivation for Variance
  1. Use first-order Taylor expansion (usually sufficient for variance calculation):
f(X)≈f(μ)+f′(μ)(X−μ)f(X) \approx f(\mu) + f'(\mu)(X-\mu)
  1. Since f(μ)f(\mu) is constant, it doesn't affect variance:
V[f(X)]≈V[f′(μ)(X−μ)]\mathbb{V}[f(X)] \approx \mathbb{V}[f'(\mu)(X-\mu)]
  1. Constant factors can be factored out:
V[f(X)]≈[f′(μ)]2V[X−μ]\mathbb{V}[f(X)] \approx [f'(\mu)]^2 \mathbb{V}[X-\mu]
  1. Since V[X−μ]=V[X]=σ2\mathbb{V}[X-\mu] = \mathbb{V}[X] = \sigma^2:
V[f(X)]≈[f′(μ)]2σ2\mathbb{V}[f(X)] \approx [f'(\mu)]^2 \sigma^2

Summary Formulas:

E[f(X)]≈f(μ)+f′′(μ)σ22V[f(X)]≈(f′(μ))2σ2\begin{aligned} \mathbb{E}\left[f(X)\right] &\approx f(\mu) + f''(\mu)\frac{\sigma^2}{2} \\ \mathbb{V}\left[f(X)\right] &\approx \left(f'(\mu)\right)^2\sigma^2 \end{aligned}

Approximation Accuracy Notes

  • The expectation approximation uses second-order expansion, providing higher accuracy
  • The variance approximation uses first-order expansion; for strongly nonlinear functions, higher-order terms may be needed
  • When f(X)f(X) is a linear function, the approximation is exact
  • Approximation quality depends on the remainder bounds below.

A small variance alone does not control Taylor remainders. If ∣f(3)∣≤M|f^{(3)}|\le M on every segment between μ\mu and a possible value of XX, Taylor's remainder gives

∣E[f(X)]−f(μ)−12f′′(μ)σ2∣≤M6E∣X−μ∣3.\left|E[f(X)]-f(\mu)-\tfrac12f''(\mu)\sigma^2\right| \le\frac{M}{6}E|X-\mu|^3.

For the linear approximation write f(X)=f(μ)+L+Rf(X)=f(\mu)+L+R, with L=f′(μ)(X−μ)L=f'(\mu)(X-\mu). When the terms have finite second moments, covariance expansion and Cauchy–Schwarz give

∣Var⁡(f(X))−Var⁡(L)∣≤2Var⁡(L)Var⁡(R)+Var⁡(R).|\operatorname{Var}(f(X))-\operatorname{Var}(L)| \le2\sqrt{\operatorname{Var}(L)\operatorname{Var}(R)}+\operatorname{Var}(R).

These bounds state the actual quantities that must be small; concentration without derivative and tail control does not guarantee accuracy.

Covariance and Correlation

When working with multiple random variables, we often want to measure their relationship.

Covariance

Cov(X,Y)=E[(X−μX)(Y−μY)]=E[XY]−E[X]E[Y]\text{Cov}(X,Y) = \mathbb{E}[(X - \mu_X)(Y - \mu_Y)] = \mathbb{E}[XY] - \mathbb{E}[X]\mathbb{E}[Y]

Correlation Coefficient

ρX,Y=Cov(X,Y)σXσY\rho_{X,Y} = \frac{\text{Cov}(X,Y)}{\sigma_X \sigma_Y}

Properties:

  • −1≤ρX,Y≤1-1 \leq \rho_{X,Y} \leq 1
  • ρ=1\rho = 1: Perfect positive linear relationship
  • ρ=−1\rho = -1: Perfect negative linear relationship
  • ρ=0\rho = 0: No linear relationship (but may have non-linear relationship)
ProofCovariance and the correlation bound

Expanding centered products gives Cov⁡(X,Y)=E[XY]−E[X]E[Y]\operatorname{Cov}(X,Y)=E[XY]-E[X]E[Y] and

Var⁡(X+Y)=Var⁡(X)+Var⁡(Y)+2Cov⁡(X,Y).\operatorname{Var}(X+Y)=\operatorname{Var}(X)+\operatorname{Var}(Y)+2\operatorname{Cov}(X,Y).

For U=X−E[X]U=X-E[X] and V=Y−E[Y]V=Y-E[Y], nonnegativity of E[(U−tV)2]E[(U-tV)^2] at t=E[UV]/E[V2]t=E[UV]/E[V^2] gives E[UV]2≤E[U2]E[V2]E[UV]^2\le E[U^2]E[V^2]. Thus ∣ρ∣≤1|\rho|\le1 when both variances are positive. Equality is equivalent to U=tVU=tV almost surely by the zero-variance argument. Positive tt gives ρ=1\rho=1, negative tt gives ρ=−1\rho=-1. With a zero variance, correlation is undefined.

Independence gives zero covariance by the product formula. The converse fails: if XX is uniform on {−1,0,1}\{-1,0,1\} and Y=X2Y=X^2, then E[X]=E[X3]=0E[X]=E[X^3]=0, so covariance is zero, yet YY is determined by XX and is not constant.

Common Distributions and Their Moments

DistributionExpected ValueVariance
Bernoulli(p)ppp(1−p)p(1-p)
Binomial(n,p)npnpnp(1−p)np(1-p)
Poisson(λ)λ\lambdaλ\lambda
Uniform(a,b)a+b2\frac{a+b}{2}(b−a)212\frac{(b-a)^2}{12}
Normal(μ,σ²)μ\muσ2\sigma^2
Exponential(λ)1λ\frac{1}{\lambda}1λ2\frac{1}{\lambda^2}

Important Theorems

TheoremLaw of Large Numbers

For i.i.d. random variables X1,X2,...,XnX_1, X_2, ..., X_n with mean μ\mu:

1n∑i=1nXi→Pμ as n→∞\frac{1}{n}\sum_{i=1}^{n} X_i \xrightarrow{P} \mu \text{ as } n \to \infty

This chapter uses the finite-variance version: an infinite i.i.d. sequence with E[Xi]=μE[X_i]=\mu and Var⁡(Xi)<∞\operatorname{Var}(X_i)<\infty. Its complete Chebyshev proof is given in the weak-law note. A finite first moment alone also suffices for the general i.i.d. weak law, but that requires a truncation argument beyond the displayed variance proof.

TheoremCentral Limit Theorem

For i.i.d. random variables with mean μ\mu and variance σ2\sigma^2:

∑i=1nXi−nμσn→DN(0,1) as n→∞\frac{\sum_{i=1}^{n} X_i - n\mu}{\sigma\sqrt{n}} \xrightarrow{D} N(0,1) \text{ as } n \to \infty

The assumptions here are an infinite i.i.d. sequence and 0<σ2<∞0<\sigma^2<\infty. For zero variance the displayed normalization is undefined. A proof via characteristic functions expands the standardized one-variable transform as 1−t2/2+o(t2)1-t^2/2+o(t^2), takes its nn-th power at t/nt/\sqrt n to obtain e−t2/2e^{-t^2/2}, and then applies Lévy's continuity theorem. That last theorem and the expansion under a finite second moment are not established in this collection yet; this statement records a dependency for the planned probability-limit chapter, not a completed proof.

Expectation with Multiple Random Variables

When working with functions of multiple random variables, we need to understand how to compute their expectations.

Expectation of Functions of Multiple Variables

For a function g(X,Y)g(X,Y) of two random variables, the expectation is computed using the joint distribution:

E[g(X,Y)]={∑x∑yg(x,y)⋅pX,Y(x,y)(discrete)∬R2g(x,y)⋅fX,Y(x,y)dxdy(continuous)\mathbb{E}[g(X,Y)] = \begin{cases} \sum_{x}\sum_{y} g(x,y) \cdot p_{X,Y}(x,y) & \text{(discrete)} \\ \iint_{\mathbb{R}^2} g(x,y) \cdot f_{X,Y}(x,y) dx dy & \text{(continuous)} \end{cases}

Key Properties

From this definition, we derive important properties:

  1. Linearity: E[X+Y]=E[X]+E[Y]\mathbb{E}[X + Y] = \mathbb{E}[X] + \mathbb{E}[Y] (always holds)
  2. Products: E[XY]=E[X]E[Y]\mathbb{E}[XY] = \mathbb{E}[X]\mathbb{E}[Y] (holds only when X and Y are independent)

Computing Expectations from Joint Distributions

Geometric Interpretation for Continuous Case

For a joint probability density function f(x,y)f(x,y), computing E[X]\mathbb{E}[X] involves integrating over the entire plane:

E[X]=∬R2x⋅f(x,y)dxdy\mathbb{E}[X] = \iint_{\mathbb{R}^2} x \cdot f(x,y) dx dy

This can be understood geometrically as finding the "center of mass" in the x-direction of the 3D surface formed by the joint density.

The computation can be done in two equivalent ways:

  1. Direct integration: Integrate x⋅f(x,y)x \cdot f(x,y) over the entire plane
  2. Using marginal density: First find fX(x)=∫−∞∞f(x,y)dyf_X(x) = \int_{-\infty}^{\infty} f(x,y) dy, then compute E[X]=∫−∞∞x⋅fX(x)dx\mathbb{E}[X] = \int_{-\infty}^{\infty} x \cdot f_X(x) dx

The second approach works because:

E[X]=∫−∞∞∫−∞∞x⋅f(x,y)dydx=∫−∞∞x(∫−∞∞f(x,y)dy)dx=∫−∞∞x⋅fX(x)dx\mathbb{E}[X] = \int_{-\infty}^{\infty} \int_{-\infty}^{\infty} x \cdot f(x,y) dy dx = \int_{-\infty}^{\infty} x \left(\int_{-\infty}^{\infty} f(x,y) dy\right) dx = \int_{-\infty}^{\infty} x \cdot f_X(x) dx

Connection to Discrete Case

Similarly, for discrete random variables:

E[X]=∑x∑yx⋅pX,Y(x,y)=∑xx(∑ypX,Y(x,y))=∑xx⋅pX(x)\mathbb{E}[X] = \sum_{x}\sum_{y} x \cdot p_{X,Y}(x,y) = \sum_{x} x \left(\sum_{y} p_{X,Y}(x,y)\right) = \sum_{x} x \cdot p_X(x)

This shows that whether we work with joint distributions directly or first compute marginal distributions, we arrive at the same expectation.

Conditional Expectation

The conditional expectation of YY given X=xX = x is:

E[Y∣X=x]={∑yy⋅pY∣X(y∣x)(discrete)∫−∞∞y⋅fY∣X(y∣x)dy(continuous)\mathbb{E}[Y|X = x] = \begin{cases} \sum_{y} y \cdot p_{Y|X}(y|x) & \text{(discrete)} \\ \int_{-\infty}^{\infty} y \cdot f_{Y|X}(y|x) dy & \text{(continuous)} \end{cases}

This leads to the law of total expectation:

E[Y]=E[E[Y∣X]]\mathbb{E}[Y] = \mathbb{E}[\mathbb{E}[Y|X]]
ProofLaw of total expectation

In the discrete case define the conditional mean only at xx with pX(x)>0p_X(x)>0. Then

∑xE[Y∣X=x]pX(x)=∑x:pX(x)>0∑yy pX,Y(x,y)=E[Y].\sum_x E[Y\mid X=x]p_X(x) =\sum_{x:p_X(x)>0}\sum_y y\,p_{X,Y}(x,y) =E[Y].

Zero-marginal points carry no mass and their conditional values may be chosen arbitrarily. Absolute integrability of YY justifies rearrangement. For a joint density replace sums by integrals and use fY∣X(y∣x)fX(x)=fX,Y(x,y)f_{Y\mid X}(y\mid x)f_X(x)=f_{X,Y}(x,y) wherever fX(x)>0f_X(x)>0, followed by Fubini's theorem. This proves the displayed cases; the general measure-theoretic conditional expectation needs additional construction.

Exercises

ExerciseUniform Discrete Integer Sampling

An integer NN is selected uniformly at random from {1,2,…,103}\{1, 2, \ldots, 10^3\}.

  1. What is the probability that NN is divisible by 3? By 5? By 7?
  2. Compute the expected value E[N]\mathbb{E}[N].
  3. Compute the variance V(N)\mathbb{V}(N).
Solution
  1. Divisibility probabilities:

    • Divisible by 3: ⌊1000/3⌋=333\lfloor 1000/3 \rfloor = 333 integers, so P(3∣N)=3331000=0.333P(3 \mid N) = \frac{333}{1000} = 0.333.
    • Divisible by 5: ⌊1000/5⌋=200\lfloor 1000/5 \rfloor = 200 integers, so P(5∣N)=2001000=0.200P(5 \mid N) = \frac{200}{1000} = 0.200.
    • Divisible by 7: ⌊1000/7⌋=142\lfloor 1000/7 \rfloor = 142 integers, so P(7∣N)=1421000=0.142P(7 \mid N) = \frac{142}{1000} = 0.142.
  2. Expected value: For a uniform discrete random variable on {1,2,…,n}\{1, 2, \ldots, n\} with n=1000n = 1000:

E[N]=1n∑k=1nk=n(n+1)2n=n+12=10012=500.5.\mathbb{E}[N] = \frac{1}{n} \sum_{k=1}^{n} k = \frac{n(n+1)}{2n} = \frac{n+1}{2} = \frac{1001}{2} = 500.5.
  1. Variance: Using the sum of squares formula ∑k=1nk2=n(n+1)(2n+1)6\sum_{k=1}^n k^2 = \frac{n(n+1)(2n+1)}{6}:
E[N2]=(n+1)(2n+1)6=1001⋅20016=333833.5.\mathbb{E}[N^2] = \frac{(n+1)(2n+1)}{6} = \frac{1001 \cdot 2001}{6} = 333833.5.V(N)=E[N2]−(E[N])2=n2−112=10002−112=99999912=83333.25.\mathbb{V}(N) = \mathbb{E}[N^2] - (\mathbb{E}[N])^2 = \frac{n^2 - 1}{12} = \frac{1000^2 - 1}{12} = \frac{999999}{12} = 83333.25.
ExerciseTotal Heads in Two Biased Coin Flips

Two coins are flipped independently. The first coin lands on heads with probability 0.60.6, and the second lands on heads with probability 0.70.7. Let XX denote the total number of heads. Find E[X]\mathbb{E}[X] and V(X)\mathbb{V}(X).

Solution

Let X1∼Bernoulli(0.6)X_1 \sim \text{Bernoulli}(0.6) and X2∼Bernoulli(0.7)X_2 \sim \text{Bernoulli}(0.7) be indicator variables for the first and second coin landing heads. Then X=X1+X2X = X_1 + X_2.

  • By linearity of expectation:
E[X]=E[X1]+E[X2]=0.6+0.7=1.3.\mathbb{E}[X] = \mathbb{E}[X_1] + \mathbb{E}[X_2] = 0.6 + 0.7 = 1.3.
  • Since the coin flips are independent:
V(X)=V(X1)+V(X2)=0.6(1−0.6)+0.7(1−0.7)=0.24+0.21=0.45.\mathbb{V}(X) = \mathbb{V}(X_1) + \mathbb{V}(X_2) = 0.6(1 - 0.6) + 0.7(1 - 0.7) = 0.24 + 0.21 = 0.45.
ExerciseBinary Search Questioning

One of the numbers {1,2,…,10}\{1, 2, \dots, 10\} is chosen uniformly at random. You guess the number by asking yes-no questions of the form "Is the number strictly greater than kk?" following an optimal binary search tree strategy. Compute the expected number of questions asked.

Solution

An optimal binary search strategy on 10 elements queries:

  • Query 1: "Is N>5N > 5?" (splits into {1..5}\{1..5\} and {6..10}\{6..10\}).
  • For {1..5}\{1..5\}: query "Is N>2N > 2?" (splits into {1,2}\{1,2\} and {3,4,5}\{3,4,5\}).
  • For {6..10}\{6..10\}: query "Is N>7N > 7?" (splits into {6,7}\{6,7\} and {8,9,10}\{8,9,10\}).

The depths in the decision tree:

  • Numbers resolved in 3 questions: 6 numbers.
  • Numbers resolved in 4 questions: 4 numbers.

Since all 10 outcomes are equally likely:

E[Questions]=6×3+4×410=18+1610=3.4.\mathbb{E}[\text{Questions}] = \frac{6 \times 3 + 4 \times 4}{10} = \frac{18 + 16}{10} = 3.4.
ExerciseSt. Petersburg Game

A player tosses a fair coin repeatedly until a tail appears for the first time. If the first tail appears on the nn-th flip, the player wins 2n2^n dollars. Let XX denote the player's winnings. Prove that E[X]=+∞\mathbb{E}[X] = +\infty.

Solution

The number of flips until the first tail follows a Geometric(1/21/2) distribution:

P(N=n)=(12)n−1⋅12=12n,n∈{1,2,3,…}P(N = n) = \left(\frac{1}{2}\right)^{n-1} \cdot \frac{1}{2} = \frac{1}{2^n}, \quad n \in \{1, 2, 3, \ldots\}

When N=nN = n, the winnings are X=2nX = 2^n. The expected payoff is:

E[X]=∑n=1∞2n⋅P(N=n)=∑n=1∞2n⋅12n=∑n=1∞1=+∞.\mathbb{E}[X] = \sum_{n=1}^{\infty} 2^n \cdot P(N = n) = \sum_{n=1}^{\infty} 2^n \cdot \frac{1}{2^n} = \sum_{n=1}^{\infty} 1 = +\infty.
ExerciseInspection Paradox on School Buses

Four buses carrying 148 students arrive at a school. The buses carry 40, 33, 25, and 50 students, respectively.

  1. One student is selected uniformly at random from the 148 students. Let XX denote the number of students on that selected student's bus.
  2. One bus is selected uniformly at random from the 4 buses. Let YY denote the number of students on that selected bus. Compute E[X]\mathbb{E}[X] and E[Y]\mathbb{E}[Y], and explain the discrepancy.
Solution
  1. Student perspective (XX): The probability that the chosen student is on a bus with ii passengers is proportional to the bus size:
P(X=i)=i148,i∈{40,33,25,50}P(X = i) = \frac{i}{148}, \quad i \in \{40, 33, 25, 50\}E[X]=∑i⋅i148=402+332+252+502148=1600+1089+625+2500148=5814148≈39.28.\mathbb{E}[X] = \sum i \cdot \frac{i}{148} = \frac{40^2 + 33^2 + 25^2 + 50^2}{148} = \frac{1600 + 1089 + 625 + 2500}{148} = \frac{5814}{148} \approx 39.28.
  1. Bus perspective (YY): Each bus is chosen with probability 14\frac{1}{4}:
E[Y]=40+33+25+504=1484=37.0.\mathbb{E}[Y] = \frac{40 + 33 + 25 + 50}{4} = \frac{148}{4} = 37.0.

The difference E[X]>E[Y]\mathbb{E}[X] > \mathbb{E}[Y] is an instance of the inspection paradox: larger buses contain more students and are therefore more likely to be sampled when choosing a student at random. In fact, E[X]=E[Y]+V(Y)E[Y]\mathbb{E}[X] = \mathbb{E}[Y] + \frac{\mathbb{V}(Y)}{\mathbb{E}[Y]}.

ExerciseVariance in the Bus Inspection Paradox

Compute the variances V(X)\mathbb{V}(X) and V(Y)\mathbb{V}(Y) for the random variables XX and YY defined in the preceding bus problem.

Solution
  1. Variance of XX:
E[X2]=403+333+253+503148=64000+35937+15625+125000148=240562148≈1625.42.\mathbb{E}[X^2] = \frac{40^3 + 33^3 + 25^3 + 50^3}{148} = \frac{64000 + 35937 + 15625 + 125000}{148} = \frac{240562}{148} \approx 1625.42.V(X)=E[X2]−(E[X])2≈1625.42−(39.2838)2≈1625.42−1543.22=82.20.\mathbb{V}(X) = \mathbb{E}[X^2] - (\mathbb{E}[X])^2 \approx 1625.42 - (39.2838)^2 \approx 1625.42 - 1543.22 = 82.20.
  1. Variance of YY:
E[Y2]=402+332+252+5024=58144=1453.5.\mathbb{E}[Y^2] = \frac{40^2 + 33^2 + 25^2 + 50^2}{4} = \frac{5814}{4} = 1453.5.V(Y)=E[Y2]−(E[Y])2=1453.5−372=1453.5−1369=84.50.\mathbb{V}(Y) = \mathbb{E}[Y^2] - (\mathbb{E}[Y])^2 = 1453.5 - 37^2 = 1453.5 - 1369 = 84.50.
ExerciseProper Scoring Rule for Weather Forecasting

A weather forecaster announces a probability p∈[0,1]p \in [0, 1] that it will rain tomorrow. The forecaster receives a score of 1−(1−p)21 - (1 - p)^2 if it rains, and 1−p21 - p^2 if it does not rain. Suppose the forecaster's true internal belief is that it will rain with probability p∗p^*. What probability pp should the forecaster report to maximize their expected score?

Solution

The expected score as a function of the announced probability pp is:

E[Score]=p∗[1−(1−p)2]+(1−p∗)[1−p2]=p∗[2p−p2]+(1−p∗)(1−p2)=2pp∗−p∗p2+1−p2−p∗+p∗p2=2pp∗−p2+1−p∗.\begin{aligned} \mathbb{E}[\text{Score}] &= p^* \left[1 - (1 - p)^2\right] + (1 - p^*)\left[1 - p^2\right] \\ &= p^* [2p - p^2] + (1 - p^*)(1 - p^2) \\ &= 2p p^* - p^* p^2 + 1 - p^2 - p^* + p^* p^2 \\ &= 2p p^* - p^2 + 1 - p^*. \end{aligned}

To find the maximum, take the derivative with respect to pp:

ddpE[Score]=2p∗−2p.\frac{d}{dp}\mathbb{E}[\text{Score}] = 2p^* - 2p.

Setting the derivative to zero yields p=p∗p = p^*. The second derivative d2dp2E[Score]=−2<0\frac{d^2}{dp^2}\mathbb{E}[\text{Score}] = -2 < 0 confirms this is a strict global maximum. Thus, the Brier scoring rule is strictly proper: the forecaster maximizes expected score if and only if reporting their honest probability p=p∗p = p^*.

ExerciseQuotient Form and Independence of CDFs

Explain why the independence of random variables X1,X2,…,XnX_1, X_2, \ldots, X_n cannot be formulated using a quotient ratio of CDFs:

FX1,…,Xn(x1,…,xn)FX1(x1)=FX2(x2)⋯FXn(xn).\frac{F_{X_1, \ldots, X_n}(x_1, \ldots, x_n)}{F_{X_1}(x_1)} = F_{X_2}(x_2) \cdots F_{X_n}(x_n).
Solution

Independence is defined by the product rule for joint probabilities: FX1,…,Xn(x1,…,xn)=∏i=1nFXi(xi)F_{X_1, \ldots, X_n}(x_1, \ldots, x_n) = \prod_{i=1}^n F_{X_i}(x_i).

A quotient of CDFs FX,Y(x,y)FX(x)\frac{F_{X, Y}(x, y)}{F_X(x)} does not correspond to the conditional CDF P(Y≤y∣X≤x)P(Y \leq y \mid X \leq x) in general because conditioning on X≤xX \leq x requires the joint probability P(X≤x,Y≤y)P(X \leq x, Y \leq y) divided by P(X≤x)P(X \leq x), which is not FY(y)F_Y(y) unless the variables are independent.

Moreover, writing the condition as a quotient requires assuming FX1(x1)>0F_{X_1}(x_1) > 0, failing for all points outside the support. Independence is an intrinsic symmetry between unconditional distributions, correctly stated via products rather than ratios.

ExerciseMoments and Linear Combinations of Independent RVs

Let XX and YY be independent discrete random variables with distributions:

x−101P(X=x)161312y024P(Y=y)141214\begin{array}{c|ccc} x & -1 & 0 & 1 \\ \hline P(X = x) & \frac{1}{6} & \frac{1}{3} & \frac{1}{2} \end{array} \qquad \begin{array}{c|ccc} y & 0 & 2 & 4 \\ \hline P(Y = y) & \frac{1}{4} & \frac{1}{2} & \frac{1}{4} \end{array}

Let W=3Y−6XW = 3Y - 6X and Z=2X+YZ = 2X + Y.

  1. Compute E[X]\mathbb{E}[X] and V(X)\mathbb{V}(X).
  2. Compute E[W]\mathbb{E}[W].
  3. Tabulate the probability distribution of ZZ.
Solution
  1. Moments of XX:
E[X]=(−1)(16)+0(13)+1(12)=−16+36=13.\mathbb{E}[X] = (-1)\left(\frac{1}{6}\right) + 0\left(\frac{1}{3}\right) + 1\left(\frac{1}{2}\right) = -\frac{1}{6} + \frac{3}{6} = \frac{1}{3}.E[X2]=(−1)2(16)+02(13)+12(12)=16+36=23.\mathbb{E}[X^2] = (-1)^2\left(\frac{1}{6}\right) + 0^2\left(\frac{1}{3}\right) + 1^2\left(\frac{1}{2}\right) = \frac{1}{6} + \frac{3}{6} = \frac{2}{3}.V(X)=E[X2]−(E[X])2=23−19=59.\mathbb{V}(X) = \mathbb{E}[X^2] - (\mathbb{E}[X])^2 = \frac{2}{3} - \frac{1}{9} = \frac{5}{9}.
  1. Expectation E[W]\mathbb{E}[W]: First, compute E[Y]\mathbb{E}[Y]:
E[Y]=0(14)+2(12)+4(14)=0+1+1=2.\mathbb{E}[Y] = 0\left(\frac{1}{4}\right) + 2\left(\frac{1}{2}\right) + 4\left(\frac{1}{4}\right) = 0 + 1 + 1 = 2.

By linearity:

E[W]=3E[Y]−6E[X]=3(2)−6(13)=6−2=4.\mathbb{E}[W] = 3\mathbb{E}[Y] - 6\mathbb{E}[X] = 3(2) - 6\left(\frac{1}{3}\right) = 6 - 2 = 4.
  1. Distribution of Z=2X+YZ = 2X + Y: Possible values of 2X2X are {−2,0,2}\{-2, 0, 2\}. Values of YY are {0,2,4}\{0, 2, 4\}. The possible values of Z=2X+YZ = 2X + Y are {−2,0,2,4,6}\{-2, 0, 2, 4, 6\}. Since XX and YY are independent, P(X=x,Y=y)=P(X=x)P(Y=y)P(X=x, Y=y) = P(X=x)P(Y=y):
    • P(Z=−2)=P(X=−1,Y=0)=16⋅14=124P(Z = -2) = P(X=-1, Y=0) = \frac{1}{6} \cdot \frac{1}{4} = \frac{1}{24}.
P(Z=0)=P(X=−1,Y=2)+P(X=0,Y=0)=16⋅12+13⋅14=112+112=16P(Z = 0) = P(X=-1, Y=2) + P(X=0, Y=0) = \frac{1}{6} \cdot \frac{1}{2} + \frac{1}{3} \cdot \frac{1}{4} = \frac{1}{12} + \frac{1}{12} = \frac{1}{6}

.

P(Z=2)=P(X=−1,Y=4)+P(X=0,Y=2)+P(X=1,Y=0)=16⋅14+13⋅12+12⋅14=124+16+18=13P(Z = 2) = P(X=-1, Y=4) + P(X=0, Y=2) + P(X=1, Y=0) = \frac{1}{6}\cdot\frac{1}{4} + \frac{1}{3}\cdot\frac{1}{2} + \frac{1}{2}\cdot\frac{1}{4} = \frac{1}{24} + \frac{1}{6} + \frac{1}{8} = \frac{1}{3}

.

P(Z=4)=P(X=0,Y=4)+P(X=1,Y=2)=13⋅14+12⋅12=112+14=13P(Z = 4) = P(X=0, Y=4) + P(X=1, Y=2) = \frac{1}{3}\cdot\frac{1}{4} + \frac{1}{2}\cdot\frac{1}{2} = \frac{1}{12} + \frac{1}{4} = \frac{1}{3}

.

  • P(Z=6)=P(X=1,Y=4)=12⋅14=18P(Z = 6) = P(X=1, Y=4) = \frac{1}{2}\cdot\frac{1}{4} = \frac{1}{8}.
z−20246P(Z=z)12416131318\begin{array}{c|ccccc} z & -2 & 0 & 2 & 4 & 6 \\ \hline P(Z = z) & \frac{1}{24} & \frac{1}{6} & \frac{1}{3} & \frac{1}{3} & \frac{1}{8} \end{array}
ExerciseAdditivity of Variance by Mathematical Induction

Prove that for any sequence of pairwise independent random variables {X1,X2,…,Xn}\{X_1, X_2, \ldots, X_n\} with finite variances:

V(∑i=1nXi)=∑i=1nV(Xi).\mathbb{V}\left(\sum_{i=1}^n X_i\right) = \sum_{i=1}^n \mathbb{V}(X_i).
Proof

We prove the statement by mathematical induction on nn.

Base Case (n=1n = 1): V(X1)=V(X1)\mathbb{V}(X_1) = \mathbb{V}(X_1), which holds trivially. For n=2n = 2, by definition:

V(X1+X2)=E[(X1+X2)2]−(E[X1+X2])2=E[X12]+2E[X1X2]+E[X22]−(E[X1]+E[X2])2.\begin{aligned} \mathbb{V}(X_1 + X_2) &= \mathbb{E}[(X_1 + X_2)^2] - (\mathbb{E}[X_1 + X_2])^2 \\ &= \mathbb{E}[X_1^2] + 2\mathbb{E}[X_1 X_2] + \mathbb{E}[X_2^2] - (\mathbb{E}[X_1] + \mathbb{E}[X_2])^2. \end{aligned}

By independence, E[X1X2]=E[X1]E[X2]\mathbb{E}[X_1 X_2] = \mathbb{E}[X_1]\mathbb{E}[X_2], so the cross-terms cancel:

V(X1+X2)=(E[X12]−E[X1]2)+(E[X22]−E[X2]2)=V(X1)+V(X2).\mathbb{V}(X_1 + X_2) = (\mathbb{E}[X_1^2] - \mathbb{E}[X_1]^2) + (\mathbb{E}[X_2^2] - \mathbb{E}[X_2]^2) = \mathbb{V}(X_1) + \mathbb{V}(X_2).

Inductive Step: Assume the property holds for n=kn = k, i.e. V(∑i=1kXi)=∑i=1kV(Xi)\mathbb{V}\left(\sum_{i=1}^k X_i\right) = \sum_{i=1}^k \mathbb{V}(X_i). For n=k+1n = k + 1, let Sk=∑i=1kXiS_k = \sum_{i=1}^k X_i. Since Xk+1X_{k+1} is independent of each X1,…,XkX_1, \dots, X_k, it is independent of their sum SkS_k. Applying the two-variable additivity rule:

V(Sk+Xk+1)=V(Sk)+V(Xk+1)=∑i=1kV(Xi)+V(Xk+1)=∑i=1k+1V(Xi).\mathbb{V}(S_k + X_{k+1}) = \mathbb{V}(S_k) + \mathbb{V}(X_{k+1}) = \sum_{i=1}^k \mathbb{V}(X_i) + \mathbb{V}(X_{k+1}) = \sum_{i=1}^{k+1} \mathbb{V}(X_i).

By induction, the property holds for all n≥1n \geq 1.