Moment Generating Functions

A moment generating function packages expectations into one function, but its useful differentiation and uniqueness properties require a domain condition.

DefinitionMoment generating function

For a real random variable XX, define MX(t)=E[etX]M_X(t)=E[e^{tX}] wherever this expectation is finite. In the discrete and density cases,

MX(t)=∑xetxpX(x),MX(t)=∫RetxfX(x) dx,M_X(t)=\sum_x e^{tx}p_X(x),\qquad M_X(t)=\int_{\mathbb R}e^{tx}f_X(x)\,dx,

respectively. Always MX(0)=1M_X(0)=1. We say the MGF exists near zero if it is finite for every tt in some interval (−a,a)(-a,a) with a>0a>0.

Having all polynomial moments does not by itself justify an MGF near zero. For example, if ZZ is standard normal and X=eZX=e^Z, completing the square gives E[Xn]=en2/2E[X^n]=e^{n^2/2} for each nonnegative integer nn. Yet E[etX]=∞E[e^{tX}]=\infty for every t>0t>0: in the defining normal integral, tez−z2/2→∞te^z-z^2/2\to\infty, so the integrand eventually exceeds a positive constant.

Moment generation

TheoremDifferentiate to obtain moments

If MXM_X is finite on (−a,a)(-a,a), then every absolute moment is finite, and

MX(n)(t)=E[XnetX],∣t∣<a.M_X^{(n)}(t)=E[X^ne^{tX}],\qquad |t|<a.

In particular MX(n)(0)=E[Xn]M_X^{(n)}(0)=E[X^n].

Proof

Choose ∣t∣<b<c<a|t|<b<c<a. Since ec∣X∣≤ecX+e−cXe^{c|X|}\le e^{cX}+e^{-cX}, its expectation is finite. The function une−(c−b)uu^n e^{-(c-b)u} is bounded for u≥0u\ge0, so for some finite CnC_n,

∣X∣neb∣X∣≤Cnec∣X∣.|X|^n e^{b|X|}\le C_n e^{c|X|}.

The same bound for n+1n+1 dominates difference quotients of the nn-th derivative in a small interval around tt, by the mean value theorem. Dominated convergence therefore passes each derivative through the expectation. Repeating establishes all orders. At t=0t=0, the bound also gives finiteness of every absolute moment. This uses dominated convergence as an integration prerequisite; its general proof remains outside this note.

Affine transformations and independent sums

For Y=aX+bY=aX+b, the definition gives

MY(t)=E[et(aX+b)]=ebtMX(at),M_Y(t)=E[e^{t(aX+b)}]=e^{bt}M_X(at),

whenever MX(at)M_X(at) is finite. If X,YX,Y are independent, their joint law factors, giving

MX+Y(t)=E[etXetY]=MX(t)MY(t)M_{X+Y}(t)=E[e^{tX}e^{tY}]=M_X(t)M_Y(t)

on the common finite domain. In a discrete model this follows by factoring the double sum; for densities it follows by factoring the double integral. Nonnegative integrands justify the interchange, and finiteness makes the product finite. Iteration proves the rule for a finite mutually independent family. Pairwise independence alone does not justify the many-variable factorization.

What uniqueness requires

TheoremMGF uniqueness with a neighborhood condition

If MX(t)=MY(t)<∞M_X(t)=M_Y(t)<\infty throughout an open interval containing zero, then XX and YY have the same distribution.

For a common finite support {x1,…,xm}\{x_1,\ldots,x_m\} of distinct values, here is an elementary proof. Equal derivatives at zero give equal expectations of all polynomials. The Lagrange polynomial

ℓj(x)=∏i≠jx−xixj−xi\ell_j(x)=\prod_{i\ne j}\frac{x-x_i}{x_j-x_i}

is one at xjx_j and zero at every other support point. Therefore P(X=xj)=E[ℓj(X)]=E[ℓj(Y)]=P(Y=xj)P(X=x_j)=E[\ell_j(X)]=E[\ell_j(Y)]=P(Y=x_j) for every jj.

For general real distributions, the standard proof extends E[ezX]E[e^{zX}] analytically to a vertical strip, uses the identity theorem to obtain equality on the imaginary axis, and applies uniqueness of characteristic functions. Those two analytic uniqueness theorems have not yet been proved in this collection, so the general theorem is an explicit prerequisite rather than a claimed elementary proof here. MIT’s probability notes state the neighborhood condition and discuss transform uniqueness [1][1] M. I. of Technology, “Fundamentals of Probability: Lecture 13, Moment Generating Functions,” 2018. https://ocw.mit.edu/courses/6-436j-fundamentals-of-probability-fall-2018/1a592ed184fb4c444547f67c9bcdd8ec_MIT6_436JF18_lec13.pdf. Equality only at t=0t=0 gives no information: every probability distribution has that value one.

Common MGFs and their derivations

Assume 0≤p≤10\le p\le1, integer n≥0n\ge0, λ>0\lambda>0, σ>0\sigma>0, and a<ba<b.

DistributionMGFFinite domain
Bernoulli1−p+pet1-p+pe^tAll real tt
Binomial(1−p+pet)n(1-p+pe^t)^nAll real tt
Poissonexp⁡(λ(et−1))\exp(\lambda(e^t-1))All real tt
Normalexp⁡(μt+σ2t2/2)\exp(\mu t+\sigma^2t^2/2)All real tt
Exponentialλ/(λ−t)\lambda/(\lambda-t)t<λt<\lambda
Uniform(etb−eta)/(t(b−a))(e^{tb}-e^{ta})/(t(b-a)) if t≠0t\ne0; 11 at t=0t=0All real tt

For Bernoulli, sum the two outcomes. A binomial variable is the sum of nn independent Bernoulli variables, so the product rule gives its MGF. For Poisson, the exponential series gives

MX(t)=e−λ∑k=0∞(λet)kk!=e−λeλet.M_X(t)=e^{-\lambda}\sum_{k=0}^{\infty}\frac{(\lambda e^t)^k}{k!} =e^{-\lambda}e^{\lambda e^t}.

For a normal density, complete the square:

tx−(x−μ)22σ2=−(x−μ−σ2t)22σ2+μt+σ2t22.tx-\frac{(x-\mu)^2}{2\sigma^2} =-\frac{(x-\mu-\sigma^2t)^2}{2\sigma^2} +\mu t+\frac{\sigma^2t^2}{2}.

The shifted normal density integrates to one, leaving the stated factor. For exponential variables integrate λe−(λ−t)x\lambda e^{-(\lambda-t)x} on [0,∞)[0,\infty), which is finite exactly when t<λt<\lambda. For uniform variables integrate etx/(b−a)e^{tx}/(b-a) on [a,b][a,b]. At zero the integrand is constant, so the apparent singularity in the quotient is removable.

A complete moment calculation

For a Poisson variable,

MX′(t)=λetMX(t),MX′′(t)=(λet+λ2e2t)MX(t).M'_X(t)=\lambda e^tM_X(t),\qquad M''_X(t)=(\lambda e^t+\lambda^2e^{2t})M_X(t).

Thus E[X]=λE[X]=\lambda, E[X2]=λ+λ2E[X^2]=\lambda+\lambda^2, and Var⁡(X)=λ\operatorname{Var}(X)=\lambda. The expectation and variance note proves the second-moment identity used in the final step.

References

  1. [1] M. I. of Technology, “Fundamentals of Probability: Lecture 13, Moment Generating Functions,” 2018. https://ocw.mit.edu/courses/6-436j-fundamentals-of-probability-fall-2018/1a592ed184fb4c444547f67c9bcdd8ec_MIT6_436JF18_lec13.pdf ↩