Probability distributions are mathematical functions that describe the likelihood of different outcomes for a random variable. They provide a complete description of the probability structure of random phenomena and are fundamental to statistical analysis and machine learning.

Overview

Probability distributions can be classified based on the nature of the random variable: discrete (countable outcomes), continuous (uncountable outcomes within intervals), or mixed (combinations). Each distribution is characterized by its support (possible values), probability function (PMF for discrete, PDF for continuous), cumulative distribution function, parameters, and moments.

DefinitionProbability Distribution

A probability distribution is a function or rule that assigns probabilities to the outcomes of a random experiment or, more generally, to the events in a sample space. Let XX be a random variable, then the probability distribution of XX is defined by its probability mass function (PMF) for discrete variables or probability density function (PDF) for continuous variables.

Without loss of generality, we can define the distribution of a random variable XX as follows:

P(X=x)=f(x)P(X = x) = f(x)

for discrete variables, where f(x)f(x) is the PMF, and

P(X≤x)=F(x)P(X \leq x) = F(x)

for continuous variables, where F(x)F(x) is the cumulative distribution function (CDF). The PMF and PDF must satisfy the properties of non-negativity and normalization:

  • For discrete variables: ∑xP(X=x)=1\sum_{x} P(X = x) = 1
  • For continuous variables: ∫−∞∞f(x)dx=1\int_{-\infty}^{\infty} f(x) dx = 1

Discrete Probability Distributions

Four discrete PMF scenarios. The quantity being counted separates Bernoulli, binomial, geometric, and Poisson models; the plotted parameters are illustrative.

Four discrete PMF scenarios. The quantity being counted separates Bernoulli, binomial, geometric, and Poisson models; the plotted parameters are illustrative.

The quantity being counted separates Bernoulli, binomial, geometric, and Poisson models; the plotted parameters are illustrative.

Bernoulli Distribution

Models a single trial with two possible outcomes (success/failure)

Parameters: pp (probability of success), where 0≤p≤10 \leq p \leq 1

Support: x∈{0,1}x \in \{0, 1\}

PMF: P(X=x)=px(1−p)1−xP(X = x) = p^x(1-p)^{1-x}

Moment Calculations:

For the expected value:

E[X]=∑x=01x⋅P(X=x)=0⋅(1−p)+1⋅p=p\begin{aligned} \mathbb{E}[X] &= \sum_{x=0}^{1} x \cdot P(X = x) \\ &= 0 \cdot (1-p) + 1 \cdot p \\ &= p \end{aligned}

For the second moment:

E[X2]=∑x=01x2⋅P(X=x)=02⋅(1−p)+12⋅p=p\begin{aligned} \mathbb{E}[X^2] &= \sum_{x=0}^{1} x^2 \cdot P(X = x) \\ &= 0^2 \cdot (1-p) + 1^2 \cdot p \\ &= p \end{aligned}

Therefore, the variance is:

V(X)=E[X2]−(E[X])2=p−p2=p(1−p)\begin{aligned} \mathbb{V}(X) &= \mathbb{E}[X^2] - (\mathbb{E}[X])^2 \\ &= p - p^2 \\ &= p(1-p) \end{aligned}

Applications: Coin flips, binary outcomes, indicator variables


Binomial Distribution

Models the number of successes in nn independent Bernoulli trials

Parameters: nn (number of trials), pp (success probability)

Support: x∈{0,1,2,...,n}x \in \{0, 1, 2, ..., n\}

PMF: P(X=x)=(nx)px(1−p)n−xP(X = x) = \binom{n}{x} p^x(1-p)^{n-x}

Moment Calculations:

The expected value can be derived using the linearity of expectation. Since X=∑i=1nXiX = \sum_{i=1}^{n} X_i where Xi∼Bernoulli(p)X_i \sim \text{Bernoulli}(p):

E[X]=E[∑i=1nXi]=∑i=1nE[Xi]=∑i=1np=np\begin{aligned} \mathbb{E}[X] &= \mathbb{E}\left[\sum_{i=1}^{n} X_i\right] \\ &= \sum_{i=1}^{n} \mathbb{E}[X_i] \\ &= \sum_{i=1}^{n} p \\ &= np \end{aligned}

For the variance, since the XiX_i are independent:

V(X)=V(∑i=1nXi)=∑i=1nV(Xi)=∑i=1np(1−p)=np(1−p)\begin{aligned} \mathbb{V}(X) &= \mathbb{V}\left(\sum_{i=1}^{n} X_i\right) \\ &= \sum_{i=1}^{n} \mathbb{V}(X_i) \\ &= \sum_{i=1}^{n} p(1-p) \\ &= np(1-p) \end{aligned}

Alternatively, we can compute directly:

E[X]=∑x=0nx(nx)px(1−p)n−x=np∑x=1n(n−1x−1)px−1(1−p)n−x=np\begin{aligned} \mathbb{E}[X] &= \sum_{x=0}^{n} x \binom{n}{x} p^x(1-p)^{n-x} \\ &= np\sum_{x=1}^{n} \binom{n-1}{x-1} p^{x-1}(1-p)^{n-x} \\ &= np \end{aligned}

Applications: Quality control, survey sampling, clinical trials

The source chapter plots a binomial example with n=5n=5 fair coin tosses. The probabilities of 00 through 55 heads are 132,532,1032,1032,532,132\frac{1}{32},\frac{5}{32},\frac{10}{32},\frac{10}{32},\frac{5}{32},\frac{1}{32}.

Binomial(n=5, p=1/2) probabilities for the number of heads.


Hypergeometric Distribution

The name looks related to the geometric distribution, but it is not. It is named after the hypergeometric function that appears in the PMF. Unlike Bernoulli-type families, a hypergeometric random variable models draws without replacement, so successive draws are dependent.

DefinitionHypergeometric distribution

The hypergeometric distribution is the law of the number of successes kk in nn draws from a finite population of size NN that contains KK successes. Its PMF is

P(X=k)=(Kk)(N−Kn−k)(Nn),max⁡(0,n−(N−K))≤k≤min⁡(K,n).P(X = k) = \frac{\binom{K}{k} \binom{N-K}{n-k}}{\binom{N}{n}}, \quad \max(0, n-(N-K)) \leq k \leq \min(K,n).

Combinatorially: choose kk successes from KK, choose n−kn-k failures from N−KN-K, and divide by the number of ways to choose nn items from NN.

ExampleDrawing red balls without replacement

A population of 2020 balls has 88 red and 1212 blue. Draw 55 balls without replacement. Let XX be the number of red balls drawn. Then XX is hypergeometric with N=20N=20, K=8K=8, n=5n=5, and

P(X=3)=(83)(122)(205)=56⋅6615504=369615504≈0.2383.P(X=3)=\frac{\binom{8}{3}\binom{12}{2}}{\binom{20}{5}}=\frac{56\cdot 66}{15504}=\frac{3696}{15504}\approx 0.2383.

Hypergeometric PMF for five draws from eight red and twelve blue balls.

TheoremMean and variance of a hypergeometric random variable

If XX is hypergeometric with integer parameters N>1N>1, 0≤K≤N0\le K\le N, and 0≤n≤N0\le n\le N, then

E[X]=nKN,Var⁡(X)=nKN(1−KN)N−nN−1.\mathbb{E}[X] = n\frac{K}{N}, \qquad \operatorname{Var}(X) = n\frac{K}{N}\left(1-\frac{K}{N}\right)\frac{N-n}{N-1}.
Proof

Write X=∑i=1nIiX=\sum_{i=1}^{n} I_i, where IiI_i is 11 if the ii-th draw is a success. Each IiI_i has mean K/NK/N, so linearity gives E[X]=nK/N\mathbb{E}[X]=nK/N.

For the second moment, E[X(X−1)]=n(n−1)K(K−1)N(N−1)E[X(X-1)]=n(n-1)\frac{K(K-1)}{N(N-1)}, hence

E[X2]=n(n−1)K(K−1)N(N−1)+nKN.E[X^2]=n(n-1)\frac{K(K-1)}{N(N-1)}+n\frac{K}{N}.

Subtracting (E[X])2(E[X])^2 and rearranging yields the variance formula.

Let p=K/Np=K/N. Then Var⁡(X)=np(1−p)N−nN−1\operatorname{Var}(X)=np(1-p)\frac{N-n}{N-1}. When NN is large relative to nn, the extra factor is near 11, and the variance is close to the binomial variance np(1−p)np(1-p).

Applications: Sampling without replacement, quality control, ecological studies


Poisson Distribution

Models the number of events occurring in a fixed interval

Parameters: λ\lambda (rate parameter), where λ>0\lambda > 0

Support: x∈{0,1,2,...}x \in \{0, 1, 2, ...\}

PMF: P(X=x)=e−λλxx!P(X = x) = \frac{e^{-\lambda}\lambda^x}{x!}

Moment Calculations:

For the expected value:

E[X]=∑x=0∞x⋅e−λλxx!=e−λ∑x=1∞λx(x−1)!=e−λλ∑x=1∞λx−1(x−1)!\begin{aligned} \mathbb{E}[X] &= \sum_{x=0}^{\infty} x \cdot \frac{e^{-\lambda}\lambda^x}{x!} \\ &= e^{-\lambda}\sum_{x=1}^{\infty} \frac{\lambda^x}{(x-1)!} \\ &= e^{-\lambda}\lambda\sum_{x=1}^{\infty} \frac{\lambda^{x-1}}{(x-1)!} \end{aligned}

Let k=x−1k = x-1:

E[X]=e−λλ∑k=0∞λkk!=e−λλeλ=λ\begin{aligned} \mathbb{E}[X] &= e^{-\lambda}\lambda\sum_{k=0}^{\infty} \frac{\lambda^k}{k!} \\ &= e^{-\lambda}\lambda e^{\lambda} \\ &= \lambda \end{aligned}

For the second moment:

E[X2]=∑x=0∞x2⋅e−λλxx!=e−λ∑x=1∞x⋅λx(x−1)!\begin{aligned} \mathbb{E}[X^2] &= \sum_{x=0}^{\infty} x^2 \cdot \frac{e^{-\lambda}\lambda^x}{x!} \\ &= e^{-\lambda}\sum_{x=1}^{\infty} x \cdot \frac{\lambda^x}{(x-1)!} \end{aligned}

Let k=x−1k = x-1:

E[X2]=e−λ∑k=0∞(k+1)⋅λk+1k!=e−λλ∑k=0∞(k+1)⋅λkk!=e−λλ(∑k=0∞k⋅λkk!+∑k=0∞λkk!)=e−λλ(λeλ+eλ)=λ(λ+1)\begin{aligned} \mathbb{E}[X^2] &= e^{-\lambda}\sum_{k=0}^{\infty} (k+1) \cdot \frac{\lambda^{k+1}}{k!} \\ &= e^{-\lambda}\lambda\sum_{k=0}^{\infty} (k+1) \cdot \frac{\lambda^k}{k!} \\ &= e^{-\lambda}\lambda\left(\sum_{k=0}^{\infty} k \cdot \frac{\lambda^k}{k!} + \sum_{k=0}^{\infty} \frac{\lambda^k}{k!}\right) \\ &= e^{-\lambda}\lambda(\lambda e^{\lambda} + e^{\lambda}) \\ &= \lambda(\lambda + 1) \end{aligned}

Therefore:

V(X)=E[X2]−(E[X])2=λ(λ+1)−λ2=λ\begin{aligned} \mathbb{V}(X) &= \mathbb{E}[X^2] - (\mathbb{E}[X])^2 \\ &= \lambda(\lambda + 1) - \lambda^2 \\ &= \lambda \end{aligned}

Properties: The Poisson distribution is the limit of Binomial(nn, pp) as n→∞n \to \infty, p→0p \to 0 with np=λnp = \lambda.

Applications: Call centers, traffic flow, radioactive decay, rare events

The source chapter also plots Rutherford's α\alpha-particle counts. The histogram below uses those tabulated probabilities.

Poisson probabilities for the number of alpha particles in a fixed interval.

ProofPoisson limit

Fix a nonnegative integer kk and set p=λ/np=\lambda/n. For n≥kn\ge k and n>λn>\lambda,

(nk)(λ/n)k(1−λ/n)n−k=λkk!∏j=0k−1(1−j/n)(1−λ/n)n(1−λ/n)−k.\binom nk(\lambda/n)^k(1-\lambda/n)^{n-k} =\frac{\lambda^k}{k!}\prod_{j=0}^{k-1}(1-j/n) (1-\lambda/n)^n(1-\lambda/n)^{-k}.

The three factors after λk/k!\lambda^k/k! tend to 1,e−λ,11,e^{-\lambda},1, respectively. This proves convergence of each point probability to the Poisson PMF; summing finitely many points gives convergence of the CDF at each finite noninteger argument.

Geometric Distribution

The geometric distribution is the waiting time until the first success in independent Bernoulli trials with success probability pp.

DefinitionGeometric distribution

If YY is the number of trials up to and including the first success, then YY follows a geometric distribution with parameter pp. Its probability mass function is

P(Y=k)=(1−p)k−1p,k=1,2,3,…P(Y = k) = (1 - p)^{k-1} p, \quad k = 1, 2, 3, \ldots

The sequence is fixed: k−1k-1 failures followed by one success, so there is no binomial coefficient. The probabilities form a geometric sequence with first term pp and ratio 1−p1-p, and they sum to 1:

∑n=1∞P{X=n}=p∑n=1∞(1−p)n−1=p1−(1−p)=1.\sum_{n=1}^{\infty} P\{X=n\}=p \sum_{n=1}^{\infty}(1-p)^{n-1}=\frac{p}{1-(1-p)}=1.
ExampleFirst successful free throw

A basketball player makes a free throw with probability p=0.3p = 0.3. What is the probability that the first success occurs on the 4th attempt?

Solution

Let XX be the attempt of the first success. Then

P(X=4)=(0.7)3⋅0.3=0.1029.P(X = 4) = (0.7)^3 \cdot 0.3 = 0.1029.
TheoremMean and variance of a geometric random variable

If XX is geometric with parameter pp, then

E[X]=1p,Var⁡(X)=1−pp2.\mathbb{E}[X] = \frac{1}{p}, \qquad \operatorname{Var}(X) = \frac{1-p}{p^2}.
Proof

For 0<p<10<p<1, put q=1−pq=1-p. The sums of kqk−1kq^{k-1} and k2qk−1k^2q^{k-1} converge by the ratio test, so the following expectations are finite and can be rearranged. Condition on the first trial. On failure, X=1+X′X=1+X' where X′X' has the same geometric distribution. With m=E[X]m=E[X] and s=E[X2]s=E[X^2],

m=p+q(1+m)=1+qm,m=p+q(1+m)=1+qm,

so m=1/pm=1/p. Similarly,

s=p+q(1+2m+s)=1+2qm+qs.s=p+q(1+2m+s)=1+2qm+qs.

Thus s=(1+2q/p)/p=(2−p)/p2s=(1+2q/p)/p=(2-p)/p^2, and Var⁡(X)=s−m2=(1−p)/p2\operatorname{Var}(X)=s-m^2=(1-p)/p^2. If p=1p=1, X=1X=1 almost surely and the same formulas hold. The case p=0p=0 has no finite waiting time and is excluded.

Other Discrete Distributions

Section Pending Migration / Draft Placeholder

The source chapter lists several additional discrete distribution families as outline placeholders: Discrete Uniform, Negative Binomial, Zeta-Bernoulli, Logarithmic Series, and Zipf distributions. Full definitions, PMF formulations, and moment calculations will be authored in a future update.

Continuous Probability Distributions

Normal (Gaussian) Distribution

The most important continuous distribution in statistics

Parameters: μ\mu (mean), σ2\sigma^2 (variance)

Support: x∈(−∞,∞)x \in (-\infty, \infty)

PDF: f(x)=1σ2πe−(x−μ)22σ2f(x) = \frac{1}{\sigma\sqrt{2\pi}} e^{-\frac{(x-\mu)^2}{2\sigma^2}}

Moment Calculations:

For the standard normal distribution Z∼N(0,1)Z \sim N(0,1):

The expected value is:

E[Z]=∫−∞∞z⋅12πe−z2/2dz=0\begin{aligned} \mathbb{E}[Z] &= \int_{-\infty}^{\infty} z \cdot \frac{1}{\sqrt{2\pi}} e^{-z^2/2} dz \\ &= 0 \end{aligned}

This follows because the integrand is an odd function and the integral converges.

For the variance:

E[Z2]=∫−∞∞z2⋅12πe−z2/2dz\begin{aligned} \mathbb{E}[Z^2] &= \int_{-\infty}^{\infty} z^2 \cdot \frac{1}{\sqrt{2\pi}} e^{-z^2/2} dz \end{aligned}

Using integration by parts with u=zu = z, dv=ze−z2/2dzdv = z e^{-z^2/2} dz:

E[Z2]=12π[−ze−z2/2]−∞∞+12π∫−∞∞e−z2/2dz=0+1=1\begin{aligned} \mathbb{E}[Z^2] &= \frac{1}{\sqrt{2\pi}} \left[ -z e^{-z^2/2} \right]_{-\infty}^{\infty} + \frac{1}{\sqrt{2\pi}} \int_{-\infty}^{\infty} e^{-z^2/2} dz \\ &= 0 + 1 \\ &= 1 \end{aligned}

Therefore, V(Z)=E[Z2]−(E[Z])2=1−0=1\mathbb{V}(Z) = \mathbb{E}[Z^2] - (\mathbb{E}[Z])^2 = 1 - 0 = 1.

For the general normal distribution X=μ+σZX = \mu + \sigma Z:

E[X]=E[μ+σZ]=μ+σE[Z]=μ\begin{aligned} \mathbb{E}[X] &= \mathbb{E}[\mu + \sigma Z] \\ &= \mu + \sigma \mathbb{E}[Z] \\ &= \mu \end{aligned} V(X)=V[μ+σZ]=σ2V(Z)=σ2\begin{aligned} \mathbb{V}(X) &= \mathbb{V}[\mu + \sigma Z] \\ &= \sigma^2 \mathbb{V}(Z) \\ &= \sigma^2 \end{aligned}

Conditions: Linear combinations of mutually independent normal variables are normal, as follows by iterating the proof below and using affine transformations. Marginal normality alone is insufficient. The i.i.d. central limit theorem concerns centered sums divided by σn\sigma\sqrt n, with finite positive variance; its general proof belongs to the planned limit-theory chapter.

Additivity Property: If X∼N(μ1,σ12)X \sim N(\mu_1, \sigma_1^2) and Y∼N(μ2,σ22)Y \sim N(\mu_2, \sigma_2^2) are independent, then:

X+Y∼N(μ1+μ2,σ12+σ22)X + Y \sim N(\mu_1 + \mu_2, \sigma_1^2 + \sigma_2^2)
ProofAdditivity

Independence gives the convolution density fX+Y(z)=∫fX(x)fY(z−x) dxf_{X+Y}(z)=\int f_X(x)f_Y(z-x)\,dx. This follows by integrating the joint density over x+y≤zx+y\le z and differentiating in zz. Put v=σ12+σ22v=\sigma_1^2+\sigma_2^2 and

mz=σ22μ1+σ12(z−μ2)v.m_z=\frac{\sigma_2^2\mu_1+\sigma_1^2(z-\mu_2)}{v}.

Completing the square gives

(x−μ1)2σ12+(z−x−μ2)2σ22=v(x−mz)2σ12σ22+(z−μ1−μ2)2v.\frac{(x-\mu_1)^2}{\sigma_1^2} +\frac{(z-x-\mu_2)^2}{\sigma_2^2} =\frac{v(x-m_z)^2}{\sigma_1^2\sigma_2^2} +\frac{(z-\mu_1-\mu_2)^2}{v}.

The first Gaussian factor integrates to 2πσ1σ2/v\sqrt{2\pi}\sigma_1\sigma_2/\sqrt v. Including the two density normalizers leaves

fX+Y(z)=12πvexp⁡(−(z−μ1−μ2)22v),f_{X+Y}(z)=\frac1{\sqrt{2\pi v}} \exp\left(-\frac{(z-\mu_1-\mu_2)^2}{2v}\right),

which is the claimed normal density. If one variance is zero, the corresponding variable is constant and the result follows by translation.

ProofUsing Moment Generating Functions

The MGF of X∼N(μ,σ2)X \sim N(\mu, \sigma^2) is:

MX(t)=eμt+12σ2t2M_X(t) = e^{\mu t + \frac{1}{2}\sigma^2 t^2}

For independent XX and YY:

MX+Y(t)=MX(t)⋅MY(t)=eμ1t+12σ12t2⋅eμ2t+12σ22t2=e(μ1+μ2)t+12(σ12+σ22)t2M_{X+Y}(t) = M_X(t) \cdot M_Y(t) = e^{\mu_1 t + \frac{1}{2}\sigma_1^2 t^2} \cdot e^{\mu_2 t + \frac{1}{2}\sigma_2^2 t^2} = e^{(\mu_1 + \mu_2)t + \frac{1}{2}(\sigma_1^2 + \sigma_2^2)t^2}

This is the MGF of N(μ1+μ2,σ12+σ22)N(\mu_1 + \mu_2, \sigma_1^2 + \sigma_2^2), proving the result.

Applications: Natural phenomena, measurement errors, statistical inference


Exponential Distribution

Models time between events in a Poisson process

Parameters: λ\lambda (rate parameter), where λ>0\lambda > 0

Support: x∈[0,∞)x \in [0, \infty)

PDF: f(x)=λe−λxf(x) = \lambda e^{-\lambda x} for x≥0x \geq 0

Moment Calculations:

For the expected value:

E[X]=∫0∞xλe−λxdx\begin{aligned} \mathbb{E}[X] &= \int_{0}^{\infty} x \lambda e^{-\lambda x} dx \end{aligned}

Using integration by parts with u=xu = x, dv=λe−λxdxdv = \lambda e^{-\lambda x} dx:

E[X]=[−xe−λx]0∞+∫0∞e−λxdx=0+[−1λe−λx]0∞=1λ\begin{aligned} \mathbb{E}[X] &= \left[ -x e^{-\lambda x} \right]_{0}^{\infty} + \int_{0}^{\infty} e^{-\lambda x} dx \\ &= 0 + \left[ -\frac{1}{\lambda} e^{-\lambda x} \right]_{0}^{\infty} \\ &= \frac{1}{\lambda} \end{aligned}

For the second moment:

E[X2]=∫0∞x2λe−λxdx\begin{aligned} \mathbb{E}[X^2] &= \int_{0}^{\infty} x^2 \lambda e^{-\lambda x} dx \end{aligned}

Using integration by parts with u=x2u = x^2, dv=λe−λxdxdv = \lambda e^{-\lambda x} dx:

E[X2]=[−x2e−λx]0∞+∫0∞2xe−λxdx=0+2λ∫0∞xλe−λxdx=2λ⋅1λ=2λ2\begin{aligned} \mathbb{E}[X^2] &= \left[ -x^2 e^{-\lambda x} \right]_{0}^{\infty} + \int_{0}^{\infty} 2x e^{-\lambda x} dx \\ &= 0 + \frac{2}{\lambda} \int_{0}^{\infty} x \lambda e^{-\lambda x} dx \\ &= \frac{2}{\lambda} \cdot \frac{1}{\lambda} \\ &= \frac{2}{\lambda^2} \end{aligned}

Therefore:

V(X)=E[X2]−(E[X])2=2λ2−(1λ)2=1λ2\begin{aligned} \mathbb{V}(X) &= \mathbb{E}[X^2] - (\mathbb{E}[X])^2 \\ &= \frac{2}{\lambda^2} - \left(\frac{1}{\lambda}\right)^2 \\ &= \frac{1}{\lambda^2} \end{aligned}

Properties: Memoryless property: P(X>s+t∣X>s)=P(X>t)P(X > s+t | X > s) = P(X > t)

Applications: Reliability engineering, queuing theory, survival analysis


ProofExponential memorylessness

For u≥0u\ge0, integration gives P(X>u)=e−λuP(X>u)=e^{-\lambda u}. For s,t≥0s,t\ge0, the event X>s+tX>s+t is contained in X>sX>s, whose probability is positive. Thus

P(X>s+t∣X>s)=e−λ(s+t)e−λs=e−λt=P(X>t).P(X>s+t\mid X>s)=\frac{e^{-\lambda(s+t)}}{e^{-\lambda s}}=e^{-\lambda t}=P(X>t).

Gamma Distribution(optional)

Generalizes exponential distribution, models waiting times

Parameters: α\alpha (shape), β\beta (rate), both >0> 0

Support: x∈[0,∞)x \in [0, \infty)

PDF: f(x)=βαΓ(α)xα−1e−βxf(x) = \frac{\beta^\alpha}{\Gamma(\alpha)} x^{\alpha-1} e^{-\beta x} for x≥0x \geq 0

Moment Calculations:

The moment generating function is:

MX(t)=E[etX]=∫0∞etxβαΓ(α)xα−1e−βxdx=βαΓ(α)∫0∞xα−1e−(β−t)xdx=βαΓ(α)⋅Γ(α)(β−t)α=(ββ−t)α for t<β\begin{aligned} M_X(t) &= \mathbb{E}[e^{tX}] \\ &= \int_{0}^{\infty} e^{tx} \frac{\beta^\alpha}{\Gamma(\alpha)} x^{\alpha-1} e^{-\beta x} dx \\ &= \frac{\beta^\alpha}{\Gamma(\alpha)} \int_{0}^{\infty} x^{\alpha-1} e^{-(\beta-t)x} dx \\ &= \frac{\beta^\alpha}{\Gamma(\alpha)} \cdot \frac{\Gamma(\alpha)}{(\beta-t)^\alpha} \\ &= \left(\frac{\beta}{\beta-t}\right)^\alpha \text{ for } t < \beta \end{aligned}

Using the MGF to find moments:

E[X]=MX′(0)=αβα(β−t)−α−1∣t=0=αβαβ−α−1=αβ\begin{aligned} \mathbb{E}[X] &= M_X'(0) \\ &= \alpha \beta^{\alpha} (\beta-t)^{-\alpha-1} \Big|_{t=0} \\ &= \alpha \beta^{\alpha} \beta^{-\alpha-1} \\ &= \frac{\alpha}{\beta} \end{aligned} E[X2]=MX′′(0)=α(α+1)βα(β−t)−α−2∣t=0=α(α+1)β2\begin{aligned} \mathbb{E}[X^2] &= M_X''(0) \\ &= \alpha(\alpha+1)\beta^{\alpha} (\beta-t)^{-\alpha-2} \Big|_{t=0} \\ &= \frac{\alpha(\alpha+1)}{\beta^2} \end{aligned}

Therefore:

V(X)=E[X2]−(E[X])2=α(α+1)β2−α2β2=αβ2\begin{aligned} \mathbb{V}(X) &= \mathbb{E}[X^2] - (\mathbb{E}[X])^2 \\ &= \frac{\alpha(\alpha+1)}{\beta^2} - \frac{\alpha^2}{\beta^2} \\ &= \frac{\alpha}{\beta^2} \end{aligned}

Integer-shape case: For a positive integer α\alpha, this is the sum of α\alpha independent Exponential(β\beta) variables. Noninteger shapes are not a count of summands.

Applications: Bayesian statistics, rainfall modeling, insurance


For completeness define Γ(a)=∫0∞ua−1e−u du\Gamma(a)=\int_0^\infty u^{a-1}e^{-u}\,du for a>0a>0. Integration by parts gives Γ(a+1)=aΓ(a)\Gamma(a+1)=a\Gamma(a) and Γ(1)=1\Gamma(1)=1, hence Γ(k)=(k−1)!\Gamma(k)=(k-1)! for positive integers. Substitution u=βxu=\beta x proves that the gamma density integrates to one (its displayed formula is for x>0x>0; the value at zero does not affect the law).

The integer-shape sum property follows by convolution and induction. If the sum of kk exponential variables has density βkxk−1e−βx/(k−1)!\beta^kx^{k-1}e^{-\beta x}/(k-1)!, convolving with one more exponential gives

∫0xβkuk−1e−βu(k−1)!βe−β(x−u) du=βk+1xke−βxk!.\int_0^x\frac{\beta^ku^{k-1}e^{-\beta u}}{(k-1)!}\beta e^{-\beta(x-u)}\,du =\frac{\beta^{k+1}x^ke^{-\beta x}}{k!}.

The base case is the exponential density itself.

Logistic Distribution(optional)

Models growth curves and binary choice models

Parameters: μ\mu (location), ss (scale), where s>0s > 0

Support: x∈(−∞,∞)x \in (-\infty, \infty)

PDF: f(x)=e−(x−μ)/ss(1+e−(x−μ)/s)2f(x) = \frac{e^{-(x-\mu)/s}}{s(1+e^{-(x-\mu)/s})^2}

Moment Calculations:

The cumulative distribution function is:

F(x)=11+e−(x−μ)/sF(x) = \frac{1}{1+e^{-(x-\mu)/s}}

For the standard logistic distribution where μ=0\mu = 0 and s=1s = 1:

f(x)=e−x(1+e−x)2f(x) = \frac{e^{-x}}{(1+e^{-x})^2}

The expected value can be found using symmetry:

E[X]=∫−∞∞x⋅e−x(1+e−x)2dx\begin{aligned} \mathbb{E}[X] &= \int_{-\infty}^{\infty} x \cdot \frac{e^{-x}}{(1+e^{-x})^2} dx \end{aligned}

Let u=−xu = -x, then:

E[X]=∫∞−∞(−u)⋅eu(1+eu)2(−du)=∫−∞∞(−u)⋅eu(1+eu)2du\begin{aligned} \mathbb{E}[X] &= \int_{\infty}^{-\infty} (-u) \cdot \frac{e^{u}}{(1+e^{u})^2} (-du) \\ &= \int_{-\infty}^{\infty} (-u) \cdot \frac{e^{u}}{(1+e^{u})^2} du \end{aligned}

Using the identity eu(1+eu)2=e−u(1+e−u)2\frac{e^{u}}{(1+e^{u})^2} = \frac{e^{-u}}{(1+e^{-u})^2}:

E[X]=−∫−∞∞u⋅e−u(1+e−u)2du=−E[X]\begin{aligned} \mathbb{E}[X] &= -\int_{-\infty}^{\infty} u \cdot \frac{e^{-u}}{(1+e^{-u})^2} du \\ &= -\mathbb{E}[X] \end{aligned}

Therefore, E[X]=0\mathbb{E}[X] = 0.

For the variance:

E[X2]=∫−∞∞x2⋅e−x(1+e−x)2dx\begin{aligned} \mathbb{E}[X^2] &= \int_{-\infty}^{\infty} x^2 \cdot \frac{e^{-x}}{(1+e^{-x})^2} dx \end{aligned}

Using the substitution u=11+e−xu = \frac{1}{1+e^{-x}}, which gives x=ln⁡(u1−u)x = \ln\left(\frac{u}{1-u}\right) and dx=duu(1−u)dx = \frac{du}{u(1-u)}:

E[X2]=∫01[ln⁡(u1−u)]2du\begin{aligned} \mathbb{E}[X^2] &= \int_{0}^{1} \left[\ln\left(\frac{u}{1-u}\right)\right]^2 du \end{aligned}
ProofEvaluating the logistic second moment

For x≥0x\ge0, the standard logistic density is at most e−xe^{-x}, so the first two absolute moments are finite and the symmetry argument is justified. Integration by parts gives ∫01log⁡2u du=2\int_0^1\log^2u\,du=2. Expanding −log⁡(1−u)=∑n≥1un/n-\log(1-u)=\sum_{n\ge1}u^n/n and applying monotone convergence to the nonnegative summands gives

∫01log⁡ulog⁡(1−u) du=∑n≥11n(n+1)2=∑n≥1(1n−1n+1−1(n+1)2)=2−∑n≥11n2.\begin{aligned} \int_0^1\log u\log(1-u)\,du &=\sum_{n\ge1}\frac{1}{n(n+1)^2}\\ &=\sum_{n\ge1}\left(\frac1n-\frac1{n+1}-\frac1{(n+1)^2}\right) =2-\sum_{n\ge1}\frac1{n^2}. \end{aligned}

Here ∫01(−log⁡u)un du=1/(n+1)2\int_0^1(-\log u)u^n\,du=1/(n+1)^2 follows by integration by parts. Expanding the square of log⁡u−log⁡(1−u)\log u-\log(1-u) therefore yields

∫01log⁡2u1−u du=2∑n≥11n2=π23.\int_0^1\log^2\frac{u}{1-u}\,du=2\sum_{n\ge1}\frac1{n^2}=\frac{\pi^2}{3}.

The last equality uses the Basel identity ∑n−2=π2/6\sum n^{-2}=\pi^2/6, whose proof is not yet established in this collection. The reduction to that identity is complete; its numerical evaluation remains an explicit analysis dependency.

For the general logistic distribution X=μ+sZX = \mu + sZ where Z∼Logistic(0,1)Z \sim \text{Logistic}(0,1):

E[X]=μ+sE[Z]=μ\begin{aligned} \mathbb{E}[X] &= \mu + s\mathbb{E}[Z] \\ &= \mu \end{aligned} V(X)=s2V(Z)=s2π23\begin{aligned} \mathbb{V}(X) &= s^2\mathbb{V}(Z) \\ &= \frac{s^2\pi^2}{3} \end{aligned}

Gumbel difference. Let G1,G2G_1,G_2 be independent Gumbel variables with the same location μ\mu and scale s>0s>0, with CDF exp⁡(−e−(x−μ)/s)\exp(-e^{-(x-\mu)/s}). Set Ei=e−(Gi−μ)/sE_i=e^{-(G_i-\mu)/s}. Direct substitution gives P(Ei>t)=e−tP(E_i>t)=e^{-t} for t>0t>0, so they are independent rate-one exponential variables. Thus

P(G1−G2≤z)=P(E2≤ez/sE1)=∫0∞(1−e−ez/st)e−t dt=11+e−z/s.\begin{aligned} P(G_1-G_2\le z) &=P(E_2\le e^{z/s}E_1)\\ &=\int_0^\infty(1-e^{-e^{z/s}t})e^{-t}\,dt =\frac1{1+e^{-z/s}}. \end{aligned}

This is Logistic(0,s)(0,s). Different locations shift the difference by μ1−μ2\mu_1-\mu_2; independence and the common scale are essential assumptions.

Applications: Logistic regression, choice modeling, growth curves


For more details on random variables and their properties, see Random Variable.

For expectation and variance calculations, see Expectation and Variance.

Exercises

ExerciseBinomial Moments via Indicator Variables

Let X∼Binomial(n,p)X \sim \text{Binomial}(n, p). Using the representation of XX as a sum of independent and identically distributed Bernoulli random variables, derive the formulas for the expected value E[X]\mathbb{E}[X] and the variance V(X)\mathbb{V}(X).

Solution

Represent the binomial random variable XX as the sum of nn independent indicator random variables:

X=∑i=1nIiX = \sum_{i=1}^n I_i

where each Ii∼Bernoulli(p)I_i \sim \text{Bernoulli}(p) with:

P(Ii=1)=p,P(Ii=0)=1−p.P(I_i = 1) = p, \quad P(I_i = 0) = 1 - p.
  1. Expected Value of IiI_i:
E[Ii]=1⋅p+0⋅(1−p)=p.\mathbb{E}[I_i] = 1 \cdot p + 0 \cdot (1 - p) = p.

By the linearity of expectation:

E[X]=E[∑i=1nIi]=∑i=1nE[Ii]=∑i=1np=np.\mathbb{E}[X] = \mathbb{E}\left[\sum_{i=1}^n I_i\right] = \sum_{i=1}^n \mathbb{E}[I_i] = \sum_{i=1}^n p = np.
  1. Variance of IiI_i:
E[Ii2]=12⋅p+02⋅(1−p)=p.\mathbb{E}[I_i^2] = 1^2 \cdot p + 0^2 \cdot (1 - p) = p.V(Ii)=E[Ii2]−(E[Ii])2=p−p2=p(1−p).\mathbb{V}(I_i) = \mathbb{E}[I_i^2] - (\mathbb{E}[I_i])^2 = p - p^2 = p(1 - p).

Since the trials I1,I2,…,InI_1, I_2, \dots, I_n are mutually independent, variance is additive:

V(X)=V(∑i=1nIi)=∑i=1nV(Ii)=∑i=1np(1−p)=np(1−p).\mathbb{V}(X) = \mathbb{V}\left(\sum_{i=1}^n I_i\right) = \sum_{i=1}^n \mathbb{V}(I_i) = \sum_{i=1}^n p(1 - p) = np(1 - p).
ExerciseBinomial PMF Recurrence Relation
  1. Prove that the probability mass function of X∼Binomial(n,p)X \sim \text{Binomial}(n, p) satisfies the following recurrence relation for k∈{0,1,…,n−1}k \in \{0, 1, \dots, n-1\}:
P(X=k+1)=p1−p⋅n−kk+1⋅P(X=k).P(X = k+1) = \frac{p}{1-p} \cdot \frac{n-k}{k+1} \cdot P(X = k).
  1. Identify the base case of the recurrence.
  2. Explain the algorithmic advantage of computing binomial probabilities using this recurrence instead of calculating factorials directly.
Solution
  1. Derivation of the recurrence ratio: By the binomial PMF definition:
P(X=k)=n!k!(n−k)!pk(1−p)n−k,P(X = k) = \frac{n!}{k!(n-k)!} p^k (1-p)^{n-k},P(X=k+1)=n!(k+1)!(n−k−1)!pk+1(1−p)n−k−1.P(X = k+1) = \frac{n!}{(k+1)!(n-k-1)!} p^{k+1} (1-p)^{n-k-1}.

Taking the ratio of successive terms:

P(X=k+1)P(X=k)=n!(k+1)!(n−k−1)!pk+1(1−p)n−k−1n!k!(n−k)!pk(1−p)n−k=k!(n−k)!(k+1)!(n−k−1)!⋅pk+1(1−p)n−k−1pk(1−p)n−k.\frac{P(X = k+1)}{P(X = k)} = \frac{\frac{n!}{(k+1)!(n-k-1)!} p^{k+1} (1-p)^{n-k-1}}{\frac{n!}{k!(n-k)!} p^k (1-p)^{n-k}} = \frac{k!(n-k)!}{(k+1)!(n-k-1)!} \cdot \frac{p^{k+1}(1-p)^{n-k-1}}{p^k(1-p)^{n-k}}.

Simplifying the factorials and powers:

k!(k+1)!=1k+1,(n−k)!(n−k−1)!=n−k,pk+1(1−p)n−k−1pk(1−p)n−k=p1−p.\frac{k!}{(k+1)!} = \frac{1}{k+1}, \qquad \frac{(n-k)!}{(n-k-1)!} = n - k, \qquad \frac{p^{k+1}(1-p)^{n-k-1}}{p^k(1-p)^{n-k}} = \frac{p}{1-p}.

Therefore:

P(X=k+1)=p1−p⋅n−kk+1⋅P(X=k).P(X = k+1) = \frac{p}{1-p} \cdot \frac{n-k}{k+1} \cdot P(X = k).
  1. Base Case: The base case is the probability of zero successes:
P(X=0)=(1−p)n.P(X = 0) = (1 - p)^n.
  1. Algorithmic Advantage: Direct calculation of (nk)=n!k!(n−k)!\binom{n}{k} = \frac{n!}{k!(n-k)!} requires computing huge factorials, which rapidly cause numerical floating-point overflow for moderate values of nn (such as n>170n > 170 in standard 64-bit IEEE 754 float). The recurrence relation computes each subsequent probability P(X=k+1)P(X=k+1) from P(X=k)P(X=k) in O(1)O(1) multiplications without ever forming factorials, achieving O(n)O(n) time complexity. This avoids factorial overflow but does not guarantee numerical stability: the starting value (1−p)n(1-p)^n can underflow, and rounding errors can accumulate. The ratio derivation assumes 0<p<10<p<1; handle p=0,1p=0,1 as point masses.