Weak Law of Large Numbers

Overview

The Weak Law of Large Numbers (WLLN) is a fundamental theorem in probability theory that describes the behavior of the average of a large number of independent and identically distributed (i.i.d.) random variables.

TheoremWeak Law of Large Numbers

Let X1,X2,...,XnX_1, X_2, ..., X_n be a sequence of i.i.d. random variables with finite expected value E[Xi]=μ\mathbb{E}[X_i] = \mu and finite variance V(Xi)=σ2\mathbb{V}(X_i) = \sigma^2.

Define the sample mean as:

Xˉn=1n∑i=1nXi\bar{X}_n = \frac{1}{n}\sum_{i=1}^{n} X_i

The Weak Law of Large Numbers states that for any ϵ>0\epsilon > 0:

lim⁡n→∞P(∣Xˉn−μ∣≥ϵ)=0\lim_{n \to \infty} P(|\bar{X}_n - \mu| \geq \epsilon) = 0

Or equivalently:

Xˉn→Pμ as n→∞\bar{X}_n \xrightarrow{P} \mu \text{ as } n \to \infty

This is called convergence in probability.

LemmaMarkov and Chebyshev inequalities

For a nonnegative random variable YY and a>0a>0, P(Y≥a)≤E[Y]/aP(Y\ge a)\le E[Y]/a. Indeed Y≥a1{Y≥a}Y\ge a\mathbf1_{\{Y\ge a\}} pointwise; taking expectations proves the bound. Apply it to Y=(X−E[X])2Y=(X-E[X])^2 and a=ϵ2a=\epsilon^2 to obtain

P(∣X−E[X]∣≥ϵ)≤Var⁡(X)ϵ2.P(|X-E[X]|\ge\epsilon)\le\frac{\operatorname{Var}(X)}{\epsilon^2}.

The theorem concerns an infinite sequence (Xi)i≥1(X_i)_{i\ge1} and its first nn terms. Finite variance is the assumption of this proof; pairwise independence with a common mean and variance is already enough for the variance calculation.

ProofUsing Chebyshev's Inequality

Step 1: Compute the Expected Value of the Sample Mean

E[Xˉn]=E[1n∑i=1nXi]=1n∑i=1nE[Xi]=1n⋅nμ=μ\mathbb{E}[\bar{X}_n] = \mathbb{E}\left[\frac{1}{n}\sum_{i=1}^{n} X_i\right] = \frac{1}{n}\sum_{i=1}^{n} \mathbb{E}[X_i] = \frac{1}{n} \cdot n\mu = \mu

Step 2: Compute the Variance of the Sample Mean

Since the XiX_i are independent:

V(Xˉn)=V(1n∑i=1nXi)=1n2∑i=1nV(Xi)=1n2⋅nσ2=σ2n\mathbb{V}(\bar{X}_n) = \mathbb{V}\left(\frac{1}{n}\sum_{i=1}^{n} X_i\right) = \frac{1}{n^2}\sum_{i=1}^{n} \mathbb{V}(X_i) = \frac{1}{n^2} \cdot n\sigma^2 = \frac{\sigma^2}{n}

Step 3: Apply Chebyshev's Inequality

For any ϵ>0\epsilon > 0:

P(∣Xˉn−μ∣≥ϵ)≤V(Xˉn)ϵ2=σ2nϵ2P(|\bar{X}_n - \mu| \geq \epsilon) \leq \frac{\mathbb{V}(\bar{X}_n)}{\epsilon^2} = \frac{\sigma^2}{n\epsilon^2}

Step 4: Take the limit

lim⁡n→∞P(∣Xˉn−μ∣≥ϵ)≤lim⁡n→∞σ2nϵ2=0\lim_{n \to \infty} P(|\bar{X}_n - \mu| \geq \epsilon) \leq \lim_{n \to \infty} \frac{\sigma^2}{n\epsilon^2} = 0

Since probabilities are non-negative, the limit must be exactly 0.

Weak law: analytical concentration. The curves show the analytical variance and Chebyshev bound, not simulated sample paths.

Weak law: analytical concentration. The curves show the analytical variance and Chebyshev bound, not simulated sample paths.

The curves show the analytical variance and Chebyshev bound, not simulated sample paths.

Interpretation

The WLLN tells us that as the sample size increases, the sample mean Xˉn\bar{X}_n converges in probability to the true mean μ\mu. This means that for large nn, the sample mean will be close to the population mean with high probability.

Applications

  1. Statistics: Justifies using sample averages to estimate population parameters
  2. Gambling: Explains why casinos have consistent profits
  3. Insurance: Forms the basis for risk pooling and premium calculation
  4. Quality Control: Validates using sample means to monitor processes

Relationship to Strong Law

The Strong Law of Large Numbers (SLLN) states almost sure convergence:

P(lim⁡n→∞Xˉn=μ)=1P\left(\lim_{n \to \infty} \bar{X}_n = \mu\right) = 1

Almost-sure convergence implies convergence in probability. For fixed ϵ>0\epsilon>0, let BN=⋃n≥N{∣Yn−Y∣≥ϵ}B_N=\bigcup_{n\ge N}\{|Y_n-Y|\ge\epsilon\}. These events decrease with NN. Almost-sure convergence makes their intersection a null event, so continuity of probability from above gives P(BN)→0P(B_N)\to0. For n≥Nn\ge N, P(∣Yn−Y∣≥ϵ)≤P(BN)P(|Y_n-Y|\ge\epsilon)\le P(B_N).

The converse fails for general sequences. On [0,1)[0,1) with uniform probability, list the indicators of the 2k2^k dyadic half-open intervals of length 2−k2^{-k}, level by level for k=1,2,…k=1,2,\ldots. Each indicator has probability 2−k2^{-k} of being one, so the sequence converges in probability to zero. Every point lies in one interval at every level and outside another, so its indicators take both values infinitely often and do not converge pointwise. This distinguishes convergence notions; it does not deny that the i.i.d. finite-variance sample means in this chapter also satisfy a strong law. A proof of that stronger theorem belongs to the planned probability-limit chapter.

ExampleCoin Flipping

For a fair coin with P(Heads)=0.5P(\text{Heads}) = 0.5, let Xi=1X_i = 1 if the ii-th flip is heads, 00 otherwise.

  • μ=E[Xi]=0.5\mu = \mathbb{E}[X_i] = 0.5
  • σ2=V(Xi)=0.25\sigma^2 = \mathbb{V}(X_i) = 0.25

The proportion of heads in nn flips is Xˉn\bar{X}_n. By WLLN:

lim⁡n→∞P(∣Xˉn−0.5∣≥ϵ)=0\lim_{n \to \infty} P(|\bar{X}_n - 0.5| \geq \epsilon) = 0

This means that as we flip the coin more times, the proportion of heads will approach 0.5.

For more details on expectation and variance, see Expectation and Variance.