Skip to the content

Topics Covered

MGF CGF Cumulants from Moments Characteristic Function PGF WLLN SLLN Convergence in Probability Convergence in Distribution Almost Sure Convergence Central Limit Theorem Lindeberg–Lévy CLT Liapounoff CLT
On this page
  1. 1. Moment Generating Function (MGF)
  2. 2. Cumulant Generating Function (CGF)
  3. 3. Characteristic Function (CF)
  4. 4. Probability Generating Function (PGF)
  5. 5. Modes of Convergence
  6. 6. Weak Law of Large Numbers (WLLN)
  7. 7. Strong Law of Large Numbers (SLLN)
  8. 8. Central Limit Theorem (CLT)
  9. Worked Problems on Generating Functions and Limit Theorems
  10. Key Take-aways from Unit 5

1. Moment Generating Function (MGF)

DEFINITION

The moment generating function of a random variable \(X\) is

\[ M_X(t) \;=\; E(e^{tX}), \quad t \in \mathbb{R}, \]

provided the expectation exists in some neighbourhood of \(t = 0\).

DISCRETE / CONTINUOUS \[ M_X(t) \;=\; \sum_x e^{tx}\, p(x) \quad \text{or} \quad \int_{-\infty}^{\infty} e^{tx}\, f(x)\,dx \]

Generating Moments

Expanding \(e^{tX} = 1 + tX + \dfrac{t^2 X^2}{2!} + \cdots\) and taking expectation:

\[ M_X(t) \;=\; \sum_{r=0}^{\infty} \dfrac{t^r}{r!}\, \mu'_r, \qquad \mu'_r \;=\; \dfrac{d^r}{dt^r} M_X(t)\Big|_{t=0} \]

Properties of MGF

  1. \(M_X(0) = 1\).
  2. Origin shift & scale: \(M_{aX+b}(t) = e^{bt} M_X(at)\).
  3. Sum of independent r.v.: \(M_{X+Y}(t) = M_X(t)\, M_Y(t)\).
  4. Uniqueness: If \(M_X(t) = M_Y(t)\) in a neighbourhood of 0, then \(X\) and \(Y\) have the same distribution.
  5. MGF may not exist for all distributions (e.g., Cauchy).
EXAMPLE 1 (Bernoulli)

\(X\) ~ Bernoulli(\(p\)): \(M_X(t) = q + p e^t\) where \(q = 1-p\).
\(M_X'(t) = pe^t,\; M_X'(0) = p = E(X)\).
\(M_X''(t) = pe^t,\; M_X''(0) = p = E(X^2);\) Var = \(p - p^2 = pq\).

EXAMPLE 2 (Exponential)

\(X\) ~ Exp(\(\lambda\)) with \(f(x) = \lambda e^{-\lambda x},\; x \ge 0\).
\(M_X(t) = \int_0^\infty \lambda e^{-(\lambda - t)x}\,dx = \dfrac{\lambda}{\lambda - t},\; t < \lambda\).
\(E(X) = M'(0) = 1/\lambda,\; E(X^2) = 2/\lambda^2;\) Var = \(1/\lambda^2\).

Proofs of the MGF Properties

A CONSTANT MULTIPLE

For a constant \(C\), \(M_{CX}(t) = E\big(e^{tCX}\big) = E\big(e^{(Ct)X}\big) = M_X(Ct)\): the MGF of \(CX\) is the MGF of \(X\) evaluated at \(Ct\).

ADDITIVE PROPERTY

If \(X_1, \ldots, X_n\) are independent, then so are \(e^{tX_1}, \ldots, e^{tX_n}\), and the multiplication theorem of expectation (Unit 4) turns the expectation of the product into the product of expectations:

\[ M_{X_1 + \cdots + X_n}(t) = E\big(e^{tX_1}e^{tX_2}\cdots e^{tX_n}\big) = E\big(e^{tX_1}\big)\cdots E\big(e^{tX_n}\big) = M_{X_1}(t)\cdots M_{X_n}(t) . \]

The MGF of a sum of independent variables is the product of their MGFs.

CHANGE OF ORIGIN AND SCALE

Let \(U = (X - a)/h\). Then \(e^{tU} = e^{-at/h}\,e^{(t/h)X}\), and \(e^{-at/h}\) is a constant:

\[ M_U(t) = E\big(e^{tU}\big) = e^{-at/h}\,E\big(e^{(t/h)X}\big) = e^{-at/h}\,M_X(t/h) . \]

So the MGF is not unchanged by a change of origin and scale: the origin brings in the factor \(e^{-at/h}\), and the scale rescales the argument.

Limitations of the MGF

WHAT CAN GO WRONG

Correction. The textbook lists, as its first limitation, that “the m.g.f. may exist but moments may not exist”, and as its second that the MGF may exist with some or all moments and yet “not generate” them. Under the definition used here, where \(M_X(t)\) is finite for all \(t\) in an interval \(-h < t < h\) around 0, both are false: such an MGF guarantees that every moment exists, and its Taylor coefficients are exactly \(\mu_r'/r!\). The two statements are true only if “exists” means finite on one side of 0. For example, \(f(x) = 1/x^2\), \(x \ge 1\), has \(E(e^{tX}) \le 1\) for every \(t \le 0\) (it is \(0.1485\) at \(t = -1\)) but \(E(X) = \infty\) (Unit 4, §1). The textbook's third limitation, that the MGF may not exist while moments do, is the lognormal case above, and is correct.

2. Cumulant Generating Function (CGF)

DEFINITION

The cumulant generating function is the natural logarithm of the MGF:

\[ K_X(t) \;=\; \ln M_X(t) \]
\[ K_X(t) \;=\; \sum_{r=1}^{\infty} \dfrac{t^r}{r!}\, k_r, \quad k_r = \dfrac{d^r}{dt^r}K_X(t)\Big|_{t=0} \]

The first few cumulants are:

Properties

  1. \(K_{aX+b}(t) = bt + K_X(at)\).
  2. For independent \(X, Y\): \(K_{X+Y}(t) = K_X(t) + K_Y(t)\) — cumulants are additive.
EXAMPLE 1 (Bernoulli)

\(K(t) = \ln(q + p e^t)\). \(K'(0) = p,\; K''(0) = pq\). Confirms mean = \(p\), var = \(pq\).

EXAMPLE 2 (Exponential)

\(K(t) = \ln \lambda - \ln(\lambda - t)\). \(K'(0) = 1/\lambda;\; K''(0) = 1/\lambda^2\).

Cumulants in Terms of Moments

DERIVATION

Write \(M_X(t) = 1 + u\) with \(u = \mu_1't + \mu_2'\dfrac{t^2}{2!} + \mu_3'\dfrac{t^3}{3!} + \mu_4'\dfrac{t^4}{4!} + \cdots\), and expand \(\log(1 + u) = u - \dfrac{u^2}{2} + \dfrac{u^3}{3} - \dfrac{u^4}{4} + \cdots\). The cumulant \(K_r\) is \(r!\) times the coefficient of \(t^r\). Collecting powers of \(t\):

\[ \begin{aligned} t:&\quad \mu_1' &&\Longrightarrow\; K_1 = \mu_1', \\ t^2:&\quad \tfrac12\mu_2' - \tfrac12\mu_1'^{\,2} &&\Longrightarrow\; K_2 = \mu_2' - \mu_1'^{\,2}, \\ t^3:&\quad \tfrac16\mu_3' - \tfrac12\mu_1'\mu_2' + \tfrac13\mu_1'^{\,3} &&\Longrightarrow\; K_3 = \mu_3' - 3\mu_2'\mu_1' + 2\mu_1'^{\,3} . \end{aligned} \]

For \(t^4\): \(u\) contributes \(\mu_4'/24\); \(u^2\) contributes \(\mu_2'^{\,2}/4 + \mu_1'\mu_3'/3\); \(u^3\) contributes \(\tfrac32\mu_1'^{\,2}\mu_2'\); \(u^4\) contributes \(\mu_1'^{\,4}\). With the signs and divisors of the log series, and multiplied by \(4! = 24\),

\[ K_4 = \mu_4' - 4\mu_3'\mu_1' - 3\mu_2'^{\,2} + 12\mu_2'\mu_1'^{\,2} - 6\mu_1'^{\,4} . \]

By Unit 4, §2, \(K_2 = \mu_2\) and \(K_3 = \mu_3\). Subtracting \(3\mu_2^2 = 3(\mu_2'^{\,2} - 2\mu_2'\mu_1'^{\,2} + \mu_1'^{\,4})\) from \(\mu_4 = \mu_4' - 4\mu_3'\mu_1' + 6\mu_2'\mu_1'^{\,2} - 3\mu_1'^{\,4}\) gives exactly the line above, so

\[ K_4 = \mu_4 - 3\mu_2^2, \qquad\text{that is}\qquad \mu_4 = K_4 + 3K_2^2 . \]

Printing note. The textbook's first line for \(K_4\) prints \(+\,4\mu_3'\mu_1'\); the sign is minus, as its next line (and the expansion of \(\mu_4\)) shows.

A shortcut. Cumulants after the first do not depend on the origin (below), so the origin may be put at the mean, \(\mu_1' = 0\). Then raw and central moments coincide and the formulae collapse at once to \(K_2 = \mu_2\), \(K_3 = \mu_3\), \(K_4 = \mu_4 - 3\mu_2^2\).

Cumulants under a Change of Origin and Scale

PROOF

With \(U = (X - a)/h\), take logarithms of the MGF result of §1:

\[ K_U(t) = \log\big[e^{-at/h}M_X(t/h)\big] = -\frac{at}{h} + K_X(t/h) . \]

Expand both sides in powers of \(t\), writing \(K_r'\) for the cumulants of \(U\):

\[ K_1't + K_2'\frac{t^2}{2!} + \cdots = -\frac{at}{h} + K_1\frac{t}{h} + K_2\frac{(t/h)^2}{2!} + \cdots + K_r\frac{(t/h)^r}{r!} + \cdots \]

and compare coefficients:

\[ K_1' = \frac{K_1 - a}{h}, \qquad K_r' = \frac{K_r}{h^r}, \quad r = 2, 3, \ldots \]

Every cumulant except the first is unaffected by a change of origin, but each depends on the scale. The first (the mean) depends on both.

ADDITIVE PROPERTY OF CUMULANTS

For independent \(X_1, \ldots, X_n\), the logarithm turns the product of MGFs into a sum: \(K_{X_1 + \cdots + X_n}(t) = K_{X_1}(t) + \cdots + K_{X_n}(t)\). Differentiating \(r\) times and putting \(t = 0\),

\[ K_r(X_1 + \cdots + X_n) = K_r(X_1) + \cdots + K_r(X_n) . \]

For \(r = 1\) and \(r = 2\) this is the familiar addition of means and of variances.

3. Characteristic Function (CF)

DEFINITION

The characteristic function of \(X\) is

\[ \phi_X(t) \;=\; E(e^{itX}), \quad t \in \mathbb{R},\; i = \sqrt{-1}. \]

Why use CF?

Properties

  1. \(\phi_X(0) = 1\) and \(|\phi_X(t)| \le 1\).
  2. \(\phi_X(t)\) is uniformly continuous on \(\mathbb{R}\).
  3. \(\phi_{aX+b}(t) = e^{ibt}\phi_X(at)\).
  4. For independent \(X,Y\): \(\phi_{X+Y}(t) = \phi_X(t)\, \phi_Y(t)\).
  5. Inversion theorem: distribution of \(X\) is recovered from \(\phi_X\).
EXAMPLE 1 (Bernoulli)

\(\phi(t) = q + p e^{it}\).

EXAMPLE 2 (Standard Normal)

\(\phi(t) = e^{-t^2/2}\) for \(X \sim N(0,1)\).

Proofs of the CF Properties

\(\phi_X(0) = 1\) AND \(|\phi_X(t)| \le 1\)

At \(t = 0\), \(\phi_X(0) = E(e^0) = 1\). For the bound, use \(|\int g| \le \int |g|\) and \(|e^{itx}| = \sqrt{\cos^2 tx + \sin^2 tx} = 1\):

\[ |\phi_X(t)| = \Big|\int_{-\infty}^{\infty} e^{itx} f(x)\,dx\Big| \le \int_{-\infty}^{\infty} |e^{itx}|\,f(x)\,dx = \int_{-\infty}^{\infty} f(x)\,dx = 1 . \]

Since \(\phi_X(t)\) is in general a complex number, \(|\phi_X(t)| \le 1\) says that it lies in the unit disc of the complex plane. (The textbook's gloss, “lies between \(-1\) and \(+1\)”, describes only the real case, such as a symmetric distribution.)

CONSTANT MULTIPLE, SUMS, ORIGIN AND SCALE

As for the MGF, with \(it\) in place of \(t\):

\[ \phi_{CX}(t) = E\big(e^{itCX}\big) = \phi_X(Ct), \] \[ \phi_{X_1 + \cdots + X_n}(t) = \phi_{X_1}(t)\cdots\phi_{X_n}(t) \quad\text{(independent } X_i), \] \[ U = \frac{X - a}{h}: \qquad \phi_U(t) = E\big(e^{-ita/h}e^{i(t/h)X}\big) = e^{-ita/h}\,\phi_X(t/h) . \]

Printing notes. The textbook ends the first proof with \(M_X(Ct)\); the result is \(\phi_X(Ct)\). In the third it writes \(\phi_X(it/h)\); the \(i\) is already inside the definition of \(\phi\), so the argument is \(t/h\).

MOMENTS FROM THE CF

If \(E|X|^r < \infty\), differentiating under the expectation gives \(\phi_X^{(r)}(0) = E\big[(iX)^r\big] = i^r\mu_r'\), so \(\mu_r' = \phi_X^{(r)}(0)/i^r\). The CF always exists, and gives every moment that exists. It cannot produce moments that do not exist: the Cauchy CF, \(e^{-|t|}\), is not differentiable at 0, matching the Cauchy distribution's lack of a mean. For such distributions the CF is used to identify the distribution (by uniqueness), not to compute moments.

4. Probability Generating Function (PGF)

DEFINITION

For a non-negative integer-valued r.v. \(X\) with PMF \(p_k = P(X=k)\), the probability generating function is

\[ P_X(s) \;=\; E(s^X) \;=\; \sum_{k=0}^{\infty} p_k\, s^k, \quad |s| \le 1. \]

Properties

  1. \(P_X(1) = 1\).
  2. \(p_k = \dfrac{1}{k!}\dfrac{d^k}{ds^k}P_X(s)\Big|_{s=0}\).
  3. \(E(X) = P_X'(1)\); Factorial moment of order 2: \(E[X(X-1)] = P_X''(1)\).
  4. For independent \(X,Y\): \(P_{X+Y}(s) = P_X(s)\, P_Y(s)\).
EXAMPLE 1 (Binomial)

\(X \sim B(n,p)\): \(P_X(s) = (q + ps)^n\).
\(P'(s) = np(q+ps)^{n-1};\; P'(1) = np = E(X)\) ✓.

EXAMPLE 2 (Poisson)

\(X \sim \) Poisson(\(\lambda\)): \(P_X(s) = e^{\lambda(s-1)}\).
\(P'(s) = \lambda e^{\lambda(s-1)};\; P'(1) = \lambda = E(X)\) ✓.

Proofs of the PGF Properties

SUMS AND A CHANGE OF ORIGIN

For independent \(X_1, \ldots, X_n\), the powers \(s^{X_i}\) are independent, so

\[ P_{X_1 + \cdots + X_n}(s) = E\big(s^{X_1}\cdots s^{X_n}\big) = P_{X_1}(s)\cdots P_{X_n}(s) . \]

For \(U = (X - a)/h\), \(s^U = s^{-a/h}\big(s^{1/h}\big)^X\), so

\[ P_U(s) = s^{-a/h}\,P_X\big(s^{1/h}\big) . \]

This identity is formal: a PGF is meant for a variable on \(0, 1, 2, \ldots\), and \(U\) is one only for special \(a, h\). The useful case is a whole-number shift, \(h = 1\): \(P_{X + a}(s) = s^a P_X(s)\).

Comparison Table

FunctionDefinitionAlways Exists?Best For
MGF\(E(e^{tX})\)No (e.g., Cauchy)Computing moments
CGF\(\ln M_X(t)\)If MGF existsCumulants & sums
CF\(E(e^{itX})\)Yes (always)Limit theorems / CLT
PGF\(E(s^X)\) for integer XYes for \(|s|\le 1\)Discrete distributions

5. Modes of Convergence

5.1 Convergence in Probability

DEFINITION

A sequence \(\{X_n\}\) converges in probability to \(X\) if for every \(\epsilon > 0\):

\[ \lim_{n \to \infty} P(|X_n - X| \ge \epsilon) \;=\; 0. \]

Notation: \(X_n \xrightarrow{P} X\).

5.2 Convergence in Distribution

DEFINITION

\(X_n \xrightarrow{d} X\) if the CDFs converge:

\[ F_{X_n}(x) \to F_X(x) \quad \text{at every continuity point of } F_X. \]

Relationship: Convergence in probability ⇒ convergence in distribution. (Reverse is true only when the limit is a constant.)

5.3 Almost Sure Convergence

DEFINITION

\(\{X_n\}\) converges to \(X\) almost surely (with probability one) if the set of outcomes on which the sequence of numbers \(X_n(\omega)\) converges to \(X(\omega)\) has probability 1:

\[ P\Big(\big\{\omega \in S : \lim_{n \to \infty} X_n(\omega) = X(\omega)\big\}\Big) = 1, \qquad\text{written}\qquad X_n \xrightarrow{a.s.} X . \]

The three modes are nested:

\[ X_n \xrightarrow{a.s.} X \;\Longrightarrow\; X_n \xrightarrow{P} X \;\Longrightarrow\; X_n \xrightarrow{d} X . \]

Convergence in probability to a constant \(a\) is the case used by the laws of large numbers: \(\lim_{n\to\infty} P(|X_n - a| < \epsilon) = 1\) for every \(\epsilon > 0\).

6. Weak Law of Large Numbers (WLLN)

STATEMENT

Let \(X_1, X_2, \ldots\) be i.i.d. random variables with mean \(\mu\) and finite variance \(\sigma^2\). Let \(\bar{X}_n = (X_1 + \cdots + X_n)/n\). Then for every \(\epsilon > 0\):

\[ P(|\bar X_n - \mu| \ge \epsilon) \;\to\; 0 \quad \text{as } n \to \infty, \]

i.e., \(\bar X_n \xrightarrow{P} \mu\).

Proof (using Chebyshev's inequality). Let \(X_1,\dots,X_n\) be independent with common mean \(\mu\) and variance \(\sigma^2\), and \(\bar X_n = \frac1n\sum_{i=1}^n X_i\).

  1. Mean of the sample mean. By linearity of expectation, \[ E(\bar X_n) = \frac1n\sum_{i=1}^n E(X_i) = \frac1n\,(n\mu) = \mu. \]
  2. Variance of the sample mean. Because the \(X_i\) are independent their variances add, and a constant factor comes out squared: \[ \operatorname{Var}(\bar X_n) = \frac{1}{n^2}\sum_{i=1}^n \operatorname{Var}(X_i) = \frac{1}{n^2}\,(n\sigma^2) = \frac{\sigma^2}{n}. \]
  3. Apply Chebyshev's inequality to \(\bar X_n\) (whose mean is \(\mu\) from step 1). For any \(\epsilon > 0\), \[ P\big(|\bar X_n - \mu| \ge \epsilon\big) \;\le\; \dfrac{\operatorname{Var}(\bar X_n)}{\epsilon^2} \;=\; \dfrac{\sigma^2/n}{\epsilon^2} = \dfrac{\sigma^2}{n\,\epsilon^2}. \]
  4. Take the limit. With \(\epsilon\) fixed, \(\dfrac{\sigma^2}{n\,\epsilon^2}\to 0\) as \(n\to\infty\). Since a probability cannot be negative, it is squeezed to \(0\): \[ P\big(|\bar X_n - \mu| \ge \epsilon\big) \;\to\; 0 \quad\text{as } n\to\infty, \] which is exactly the statement \(\bar X_n \xrightarrow{P} \mu\). \(\blacksquare\)
EXAMPLE 1

Toss a fair coin many times. Let \(X_i = 1\) if head, 0 otherwise. \(\bar X_n\) = proportion of heads. By WLLN, \(\bar X_n \to 0.5\) in probability — empirical frequency justifies the classical definition.

EXAMPLE 2

The sample mean of 1 000 dice rolls is very likely close to the true mean \(\mu = 3.5\): by Chebyshev with \(\sigma^2 = 35/12 \approx 2.92\), \(P(|\bar X_{1000} - 3.5| \ge 0.1) \le 2.92/(1000 \cdot 0.01) \approx 0.29\).

The WLLN without Identical Distributions

STATEMENT AND PROOF

Let \(X_1, X_2, \ldots\) have means \(\mu_1, \mu_2, \ldots\), and let \(B_n = \operatorname{Var}(X_1 + \cdots + X_n) < \infty\). If \(B_n/n^2 \to 0\), then for every \(\epsilon > 0\)

\[ P\left\{\left|\frac{X_1 + \cdots + X_n}{n} - \frac{\mu_1 + \cdots + \mu_n}{n}\right| < \epsilon\right\} \to 1 . \]

Proof. The average \(\bar X_n\) has mean \(\bar\mu_n = (\mu_1 + \cdots + \mu_n)/n\) and variance \(B_n/n^2\). Chebyshev's inequality with \(c = \epsilon\) gives

\[ P\big(|\bar X_n - \bar\mu_n| \ge \epsilon\big) \le \frac{B_n}{n^2\epsilon^2} \to 0 . \]

Neither independence nor a common distribution is assumed, only the variance condition. For i.i.d. variables \(B_n = n\sigma^2\) (the covariances vanish), \(B_n/n^2 = \sigma^2/n \to 0\), and this is the theorem above.

7. Strong Law of Large Numbers (SLLN)

STATEMENT (Kolmogorov)

Let \(X_1, X_2, \ldots\) be i.i.d. with mean \(\mu\). Then

\[ P\!\left(\lim_{n \to \infty} \bar X_n = \mu\right) \;=\; 1, \]

i.e., \(\bar X_n \to \mu\) almost surely.

WLLN vs. SLLN:

PropertyWLLNSLLN
Mode of convergenceIn probabilityAlmost sure (with prob 1)
StrengthWeakerStronger (implies WLLN)
ConditionsFinite variance (sufficient)Finite mean (Kolmogorov's)

8. Central Limit Theorem (CLT)

STATEMENT (Lindeberg–Lévy)

Let \(X_1, X_2, \ldots\) be i.i.d. with mean \(\mu\) and finite variance \(\sigma^2 > 0\). Define \(\bar X_n = \dfrac{1}{n}\sum_{i=1}^n X_i\). Then

\[ Z_n \;=\; \dfrac{\bar X_n - \mu}{\sigma / \sqrt{n}} \;\xrightarrow{d}\; N(0, 1) \quad \text{as } n \to \infty. \]

Equivalently, the standardized sum \(\dfrac{\sum X_i - n\mu}{\sigma\sqrt{n}}\) converges to standard normal.

Implications

Population n = 1 (skewed) Mean of n = 5 less skewed Mean of n = 30 ≈ Normal bell As n grows, the distribution of X̄ approaches Normal — the CLT
Fig 5.1 — Even a strongly skewed population (left) produces a sampling distribution of the mean that becomes ever more symmetric and bell-shaped as \(n\) increases, converging to \(N(\mu,\sigma^2/n)\). This is why normal-based inference works so widely.
EXAMPLE 1 (Quality control)

Diameter of bolts has mean 5 mm, SD 0.1 mm. For a sample of 100 bolts, the sample mean is approximately \(N(5,\; 0.1^2/100) = N(5, 0.0001)\), i.e., SE = 0.01 mm.

\(P(|\bar X - 5| < 0.02) \approx P(|Z| < 2) \approx 0.9544\).

EXAMPLE 2 (Polling)

Suppose true proportion of voters supporting party A is 0.6. From a sample of 400, the sample proportion \(\hat p \sim N(0.6,\; 0.6 \times 0.4/400) = N(0.6, 0.000\,6)\), SE ≈ 0.0245.

\(P(|\hat p - 0.6| < 0.05) \approx P(|Z| < 2.04) \approx 0.96\) — explains the typical "± 5 % margin of error".

History and the General Statement

FROM DE MOIVRE TO LIAPOUNOFF

De Moivre found the normal approximation to the binomial with \(p = \tfrac12\) in 1733; Laplace extended it and stated the central limit theorem in 1812; Liapounoff (Lyapunov) gave the first rigorous general proof in 1901. In its general form: if \(X_1, \ldots, X_n\) are independent with means \(\mu_i\) and variances \(\sigma_i^2\), then under very general conditions \(S_n = X_1 + \cdots + X_n\) is asymptotically normal with

\[ \text{mean } \mu = \sum_{i=1}^n \mu_i, \qquad \text{variance } \sigma^2 = \sum_{i=1}^n \sigma_i^2 . \]

Three Forms of the Lindeberg–Lévy Theorem

I.I.D. VARIABLES, MEAN \(\mu_1\), FINITE VARIANCE \(\sigma_1^2\)

\(S_n\) has mean \(n\mu_1\) and standard deviation \(\sigma_1\sqrt n\); \(\bar X_n = S_n/n\) has mean \(\mu_1\) and standard deviation \(\sigma_1/\sqrt n\). With \(\Phi\) the standard normal distribution function, for all \(-\infty < a < b < \infty\):

\[ \text{(i)}\quad \lim_{n\to\infty} P\left\{a \le \frac{S_n - n\mu_1}{\sigma_1\sqrt n} \le b\right\} = \Phi(b) - \Phi(a) = \int_a^b \frac{1}{\sqrt{2\pi}}e^{-x^2/2}\,dx, \] \[ \text{(ii)}\quad \lim_{n\to\infty} P\left\{a \le \frac{S_n - E(S_n)}{\sqrt{\operatorname{Var}(S_n)}} \le b\right\} = \Phi(b) - \Phi(a), \] \[ \text{(iii)}\quad \lim_{n\to\infty} P\left\{a \le \frac{\bar X_n - \mu_1}{\sigma_1/\sqrt n} \le b\right\} = \Phi(b) - \Phi(a) . \]

All three standardise the same quantity: dividing the numerator and denominator of (i) by \(n\) gives (iii).

Correction. The textbook's form (i) divides \(S_n - n\mu_1\) by \(\sigma_1/\sqrt n\). That is the standard deviation of \(\bar X_n\), not of \(S_n\); the denominator must be \(\sigma_1\sqrt n\). (With \(n = 100\), \(\sigma_1 = 2\): the two are 20 and 0.2, a hundredfold apart.)

Liapounoff's Central Limit Theorem

INDEPENDENT, NOT IDENTICALLY DISTRIBUTED

Let \(X_1, \ldots, X_n\) be independent with means \(\mu_i\), variances \(\sigma_i^2\), and finite third absolute central moments \(\rho_i^3 = E|X_i - \mu_i|^3\). Put

\[ \rho^3 = \sum_{i=1}^n \rho_i^3, \qquad \sigma^2 = \sum_{i=1}^n \sigma_i^2 . \]

If \(\lim_{n\to\infty} \rho/\sigma = 0\), then \(S_n = X_1 + \cdots + X_n\) is asymptotically \(N\big(\sum\mu_i, \sum\sigma_i^2\big)\).

The condition says that no single variable dominates the sum. For i.i.d. variables \(\rho/\sigma = (n\rho_1^3)^{1/3}/(n\sigma_1^2)^{1/2} = (\rho_1/\sigma_1)\,n^{-1/6} \to 0\), so Lindeberg–Lévy (with a finite third moment) is a special case.

The CLT and the WLLN Compared

WHAT EACH ONE SAYS

For i.i.d. variables with finite \(\mu\) and \(\sigma^2\), both laws hold. The WLLN says only that \(P(|\bar X_n - \mu| \ge \epsilon) \to 0\). The CLT says how fast, by giving the probability itself, approximately, for large \(n\):

\[ P\big(|\bar X_n - \mu| \ge \epsilon\big) = P\left(\left|\frac{\bar X_n - \mu}{\sigma/\sqrt n}\right| \ge \frac{\epsilon\sqrt n}{\sigma}\right) \approx 1 - \left[\Phi\Big(\frac{\epsilon\sqrt n}{\sigma}\Big) - \Phi\Big(-\frac{\epsilon\sqrt n}{\sigma}\Big)\right] . \]

The textbook writes “\(=\)” in the last step; it is an approximation, which the CLT makes good as \(n \to \infty\). Worked Problem 5 puts numbers on it.

For sequences that are independent but not identically distributed, the two laws can part company:

  1. For independent, uniformly bounded variables the WLLN always holds (since \(B_n \le nc^2\), \(B_n/n^2 \to 0\)), and the CLT holds provided \(B_n = \sigma_1^2 + \cdots + \sigma_n^2 \to \infty\).
  2. The CLT may hold while the WLLN fails. Take independent \(X_i \sim N(0, i)\). Every sum is exactly normal, so the CLT holds trivially. But \(B_n = n(n + 1)/2\), so \(\operatorname{Var}(\bar X_n) = (n + 1)/(2n) \to \tfrac12\), not 0: \(\bar X_n\) never settles at 0, and \(P(|\bar X_n| \ge 0.5) \to 0.4795\).

Worked Problems on Generating Functions and Limit Theorems

The textbook's generating-function problem, then its three applications of the central limit theorem, then one more problem that puts numbers on its comparison of the CLT with the weak law. The method in the CLT problems is always the same: write the variable as a sum of i.i.d. pieces, find the mean and variance of the sum, and standardise.

Source note. Each result was checked symbolically (the MGF by differentiation, the cumulants by series expansion, the sums by convolution) and the limits numerically. Worked Problem 2 corrects the range of \(p\); Worked Problem 5 is added here and is not in the textbook.

A. A Generating Function Computed

WORKED PROBLEM 1 — the MGF of the number of heads in two tosses

A coin is tossed twice and \(X\) is the number of heads, with distribution \(P(X = 0, 1, 2) = \tfrac14, \tfrac12, \tfrac14\). Find the MGF of \(X\).

\[ M_X(t) = \sum_x e^{tx}P(X = x) = e^{0}\cdot\tfrac14 + e^{t}\cdot\tfrac12 + e^{2t}\cdot\tfrac14 = \frac{e^{2t} + 2e^t + 1}{4} = \Big(\frac{1 + e^t}{2}\Big)^2 . \]

The square is no accident: \(X\) is the sum of two independent Bernoulli(\(\tfrac12\)) variables, each with MGF \(\tfrac12 + \tfrac12e^t\), and by the additive property the MGF of the sum is the product.

Moments from it. Differentiate and put \(t = 0\):

\[ M_X'(t) = \frac{e^t + e^{2t}}{2} \Rightarrow \mu_1' = M_X'(0) = 1, \qquad M_X''(t) = \frac{e^t + 2e^{2t}}{2} \Rightarrow \mu_2' = \frac32, \] \[ \operatorname{Var}(X) = \frac32 - 1^2 = \frac12 . \]

Cumulants from it. \(K_X(t) = 2\log\frac{1 + e^t}{2}\), so \(K_X'(t) = \dfrac{2e^t}{1 + e^t}\) and \(K_X''(t) = \dfrac{2e^t}{(1 + e^t)^2}\), giving \(K_1 = 1\) and \(K_2 = \tfrac24 = \tfrac12\), the mean and variance again. Carrying on, \(K_3 = 0\) (the distribution is symmetric) and \(K_4 = -\tfrac14\).

PGF. \(P_X(s) = \tfrac14 + \tfrac12s + \tfrac14s^2 = \big(\tfrac{1 + s}{2}\big)^2\), and \(P_X'(1) = 1\), the mean once more.

B. Applications of the Central Limit Theorem

WORKED PROBLEM 2 — the sum of \(n\) i.i.d. binomial variates

Discuss the distribution of the sum of \(n\) i.i.d. binomial variates.

Let \(X_1, \ldots, X_n\) be i.i.d. \(B(r, p)\), so \(E(X_i) = rp\) and \(\operatorname{Var}(X_i) = rpq\). For \(S_n = X_1 + \cdots + X_n\), the addition theorem and independence give

\[ E(S_n) = \sum_{i=1}^n rp = nrp, \qquad \operatorname{Var}(S_n) = \sum_{i=1}^n rpq = nrpq . \]

By the Lindeberg–Lévy theorem, form (ii),

\[ \lim_{n\to\infty} P\left\{a \le \frac{S_n - nrp}{\sqrt{nrpq}} \le b\right\} = \Phi(b) - \Phi(a), \qquad 0 < p < 1 . \]

Correction. The textbook gives the range as \(0 \le p < 1\). At \(p = 0\) every \(X_i\) is 0, the variance \(nrpq\) is 0, and the standardised variable is not defined; the theorem needs \(0 < \sigma^2 < \infty\), so \(0 < p < 1\).

An exact check. Here the limit is not even needed for the shape: by the PGF, \(P_{S_n}(s) = \big[(q + ps)^r\big]^n = (q + ps)^{nr}\), so \(S_n \sim B(nr, p)\) exactly (for \(r = 3\), \(n = 4\), \(p = \tfrac14\): the convolution of four \(B(3, \tfrac14)\) is \(B(12, \tfrac14)\), with mean 3 and variance \(\tfrac94\)). The CLT then says that a binomial with many trials is close to normal.

WORKED PROBLEM 3 — the binomial tends to the normal

Prove that if \(Y_n \sim B(n, p)\), then \(\displaystyle\lim_{n\to\infty} P\left\{a \le \frac{Y_n - np}{\sqrt{np(1 - p)}} \le b\right\} = \Phi(b) - \Phi(a)\), \(0 < p < 1\).

Let \(X_1, \ldots, X_n\) be i.i.d. Bernoulli, \(B(1, p)\): each has mean \(p\) and variance \(pq\). Their sum counts the successes in \(n\) independent trials, so it has the same distribution as \(Y_n\), with

\[ E(Y_n) = np, \qquad \operatorname{Var}(Y_n) = npq = np(1 - p) . \]

The Lindeberg–Lévy theorem applied to this sum gives the result. This is the De Moivre–Laplace theorem, the oldest case of the CLT.

WORKED PROBLEM 4 — a Poisson limit: \(P(Y_n \le n) \to \tfrac12\)

If \(Y_n \sim\) Poisson(\(n\)), show that \(\displaystyle\lim_{n\to\infty} P(Y_n \le n) = \tfrac12\), i.e. \(\displaystyle\sum_{k=0}^{n} \frac{e^{-n}n^k}{k!} \to \frac12\).

Write \(Y_n\) as a sum. Let \(X_1, \ldots, X_n\) be i.i.d. Poisson(1). By the PGF, \(P_{X_1 + \cdots + X_n}(s) = \big[e^{(s - 1)}\big]^n = e^{n(s - 1)}\), the PGF of Poisson(\(n\)). So \(Y_n\) is such a sum, with \(E(Y_n) = \operatorname{Var}(Y_n) = n\).

Apply the CLT. \((Y_n - n)/\sqrt n\) tends in distribution to \(N(0, 1)\). Since \(\Phi\) is continuous, the distribution functions converge at every point, in particular at 0:

\[ P(Y_n \le n) = P\left\{\frac{Y_n - n}{\sqrt n} \le 0\right\} \to \Phi(0) = \frac12 . \]

(This is the textbook's \(a = -\infty\), \(b = 0\): \(\Phi(0) - \Phi(-\infty) = \tfrac12\).) Writing out the Poisson probabilities, \(\sum_{k=0}^{n} e^{-n}n^k/k! \to \tfrac12\).

How fast? Exactly computed:

\(n\)11010010005000
\(P(Y_n \le n)\)0.73580.58300.52660.50840.5038

(At \(n = 1\) it is \(2e^{-1}\).) The approach is slow, and from above, for two reasons that push the same way: the event \(Y_n \le n\) counts the whole probability at the single value \(n\), about \(1/\sqrt{2\pi n}\), and the Poisson is skewed to the right. Together they leave an excess of about \(2/(3\sqrt{2\pi n})\) over \(\tfrac12\) (\(0.5266\) predicted and found at \(n = 100\)).

limit 1/2 0.7358 0.6160 0.5830 0.5375 0.5266 0.5119 0.5084 0.5038 0.5 0.6 0.7 1 10 100 1000 n (logarithmic scale) P(Yₙ ≤ n), Yₙ ~ Poisson(n)
Fig 5.2 — Worked Problem 4. The exact values of \(\sum_{k=0}^{n} e^{-n} n^k / k!\) approach \(\tfrac12\), as the central limit theorem says they must, but slowly and always from above: the excess is close to \(2/(3\sqrt{2\pi n})\), which halves only when \(n\) is multiplied by four.
WORKED PROBLEM 5 — how much the CLT adds to the WLLN

A fair coin is tossed \(n = 400\) times and \(\bar X\) is the proportion of heads. Bound \(P(|\bar X - \tfrac12| \ge 0.05)\) by Chebyshev's inequality, approximate it by the CLT, and compare with the exact value.

Each toss has \(\mu = \tfrac12\), \(\sigma^2 = \tfrac14\), so \(\operatorname{Var}(\bar X) = \tfrac{1}{4 \cdot 400}\) and \(\sigma/\sqrt n = \tfrac{0.5}{20} = 0.025\).

Chebyshev (the WLLN's tool):

\[ P\big(|\bar X - \tfrac12| \ge 0.05\big) \le \frac{\sigma^2}{n\epsilon^2} = \frac{0.25}{400 \times 0.0025} = 0.25 . \]

CLT: \(\epsilon\sqrt n/\sigma = 0.05 \times 20/0.5 = 2\), so

\[ P\big(|\bar X - \tfrac12| \ge 0.05\big) \approx P(|Z| \ge 2) = 2\big[1 - \Phi(2)\big] = 0.0455 . \]

Exact (binomial, 180 heads or fewer, or 220 or more): \(0.0510\).

The WLLN guarantees only that this probability goes to 0; Chebyshev caps it at 0.25. The CLT gives its actual size, within half a percentage point here.

Key Take-aways from Unit 5