The moment generating function of a random variable \(X\) is
\[ M_X(t) \;=\; E(e^{tX}), \quad t \in \mathbb{R}, \]provided the expectation exists in some neighbourhood of \(t = 0\).
Expanding \(e^{tX} = 1 + tX + \dfrac{t^2 X^2}{2!} + \cdots\) and taking expectation:
\(X\) ~ Bernoulli(\(p\)): \(M_X(t) = q + p e^t\) where \(q = 1-p\).
\(M_X'(t) = pe^t,\; M_X'(0) = p = E(X)\).
\(M_X''(t) = pe^t,\; M_X''(0) = p = E(X^2);\) Var = \(p - p^2 = pq\).
\(X\) ~ Exp(\(\lambda\)) with \(f(x) = \lambda e^{-\lambda x},\; x \ge 0\).
\(M_X(t) = \int_0^\infty \lambda e^{-(\lambda - t)x}\,dx = \dfrac{\lambda}{\lambda - t},\; t < \lambda\).
\(E(X) = M'(0) = 1/\lambda,\; E(X^2) = 2/\lambda^2;\) Var = \(1/\lambda^2\).
For a constant \(C\), \(M_{CX}(t) = E\big(e^{tCX}\big) = E\big(e^{(Ct)X}\big) = M_X(Ct)\): the MGF of \(CX\) is the MGF of \(X\) evaluated at \(Ct\).
If \(X_1, \ldots, X_n\) are independent, then so are \(e^{tX_1}, \ldots, e^{tX_n}\), and the multiplication theorem of expectation (Unit 4) turns the expectation of the product into the product of expectations:
\[ M_{X_1 + \cdots + X_n}(t) = E\big(e^{tX_1}e^{tX_2}\cdots e^{tX_n}\big) = E\big(e^{tX_1}\big)\cdots E\big(e^{tX_n}\big) = M_{X_1}(t)\cdots M_{X_n}(t) . \]The MGF of a sum of independent variables is the product of their MGFs.
Let \(U = (X - a)/h\). Then \(e^{tU} = e^{-at/h}\,e^{(t/h)X}\), and \(e^{-at/h}\) is a constant:
\[ M_U(t) = E\big(e^{tU}\big) = e^{-at/h}\,E\big(e^{(t/h)X}\big) = e^{-at/h}\,M_X(t/h) . \]So the MGF is not unchanged by a change of origin and scale: the origin brings in the factor \(e^{-at/h}\), and the scale rescales the argument.
Correction. The textbook lists, as its first limitation, that “the m.g.f. may exist but moments may not exist”, and as its second that the MGF may exist with some or all moments and yet “not generate” them. Under the definition used here, where \(M_X(t)\) is finite for all \(t\) in an interval \(-h < t < h\) around 0, both are false: such an MGF guarantees that every moment exists, and its Taylor coefficients are exactly \(\mu_r'/r!\). The two statements are true only if “exists” means finite on one side of 0. For example, \(f(x) = 1/x^2\), \(x \ge 1\), has \(E(e^{tX}) \le 1\) for every \(t \le 0\) (it is \(0.1485\) at \(t = -1\)) but \(E(X) = \infty\) (Unit 4, §1). The textbook's third limitation, that the MGF may not exist while moments do, is the lognormal case above, and is correct.
The cumulant generating function is the natural logarithm of the MGF:
\[ K_X(t) \;=\; \ln M_X(t) \]The first few cumulants are:
\(K(t) = \ln(q + p e^t)\). \(K'(0) = p,\; K''(0) = pq\). Confirms mean = \(p\), var = \(pq\).
\(K(t) = \ln \lambda - \ln(\lambda - t)\). \(K'(0) = 1/\lambda;\; K''(0) = 1/\lambda^2\).
Write \(M_X(t) = 1 + u\) with \(u = \mu_1't + \mu_2'\dfrac{t^2}{2!} + \mu_3'\dfrac{t^3}{3!} + \mu_4'\dfrac{t^4}{4!} + \cdots\), and expand \(\log(1 + u) = u - \dfrac{u^2}{2} + \dfrac{u^3}{3} - \dfrac{u^4}{4} + \cdots\). The cumulant \(K_r\) is \(r!\) times the coefficient of \(t^r\). Collecting powers of \(t\):
\[ \begin{aligned} t:&\quad \mu_1' &&\Longrightarrow\; K_1 = \mu_1', \\ t^2:&\quad \tfrac12\mu_2' - \tfrac12\mu_1'^{\,2} &&\Longrightarrow\; K_2 = \mu_2' - \mu_1'^{\,2}, \\ t^3:&\quad \tfrac16\mu_3' - \tfrac12\mu_1'\mu_2' + \tfrac13\mu_1'^{\,3} &&\Longrightarrow\; K_3 = \mu_3' - 3\mu_2'\mu_1' + 2\mu_1'^{\,3} . \end{aligned} \]For \(t^4\): \(u\) contributes \(\mu_4'/24\); \(u^2\) contributes \(\mu_2'^{\,2}/4 + \mu_1'\mu_3'/3\); \(u^3\) contributes \(\tfrac32\mu_1'^{\,2}\mu_2'\); \(u^4\) contributes \(\mu_1'^{\,4}\). With the signs and divisors of the log series, and multiplied by \(4! = 24\),
\[ K_4 = \mu_4' - 4\mu_3'\mu_1' - 3\mu_2'^{\,2} + 12\mu_2'\mu_1'^{\,2} - 6\mu_1'^{\,4} . \]By Unit 4, §2, \(K_2 = \mu_2\) and \(K_3 = \mu_3\). Subtracting \(3\mu_2^2 = 3(\mu_2'^{\,2} - 2\mu_2'\mu_1'^{\,2} + \mu_1'^{\,4})\) from \(\mu_4 = \mu_4' - 4\mu_3'\mu_1' + 6\mu_2'\mu_1'^{\,2} - 3\mu_1'^{\,4}\) gives exactly the line above, so
\[ K_4 = \mu_4 - 3\mu_2^2, \qquad\text{that is}\qquad \mu_4 = K_4 + 3K_2^2 . \]Printing note. The textbook's first line for \(K_4\) prints \(+\,4\mu_3'\mu_1'\); the sign is minus, as its next line (and the expansion of \(\mu_4\)) shows.
A shortcut. Cumulants after the first do not depend on the origin (below), so the origin may be put at the mean, \(\mu_1' = 0\). Then raw and central moments coincide and the formulae collapse at once to \(K_2 = \mu_2\), \(K_3 = \mu_3\), \(K_4 = \mu_4 - 3\mu_2^2\).
With \(U = (X - a)/h\), take logarithms of the MGF result of §1:
\[ K_U(t) = \log\big[e^{-at/h}M_X(t/h)\big] = -\frac{at}{h} + K_X(t/h) . \]Expand both sides in powers of \(t\), writing \(K_r'\) for the cumulants of \(U\):
\[ K_1't + K_2'\frac{t^2}{2!} + \cdots = -\frac{at}{h} + K_1\frac{t}{h} + K_2\frac{(t/h)^2}{2!} + \cdots + K_r\frac{(t/h)^r}{r!} + \cdots \]and compare coefficients:
\[ K_1' = \frac{K_1 - a}{h}, \qquad K_r' = \frac{K_r}{h^r}, \quad r = 2, 3, \ldots \]Every cumulant except the first is unaffected by a change of origin, but each depends on the scale. The first (the mean) depends on both.
For independent \(X_1, \ldots, X_n\), the logarithm turns the product of MGFs into a sum: \(K_{X_1 + \cdots + X_n}(t) = K_{X_1}(t) + \cdots + K_{X_n}(t)\). Differentiating \(r\) times and putting \(t = 0\),
\[ K_r(X_1 + \cdots + X_n) = K_r(X_1) + \cdots + K_r(X_n) . \]For \(r = 1\) and \(r = 2\) this is the familiar addition of means and of variances.
The characteristic function of \(X\) is
\[ \phi_X(t) \;=\; E(e^{itX}), \quad t \in \mathbb{R},\; i = \sqrt{-1}. \]\(\phi(t) = q + p e^{it}\).
\(\phi(t) = e^{-t^2/2}\) for \(X \sim N(0,1)\).
At \(t = 0\), \(\phi_X(0) = E(e^0) = 1\). For the bound, use \(|\int g| \le \int |g|\) and \(|e^{itx}| = \sqrt{\cos^2 tx + \sin^2 tx} = 1\):
\[ |\phi_X(t)| = \Big|\int_{-\infty}^{\infty} e^{itx} f(x)\,dx\Big| \le \int_{-\infty}^{\infty} |e^{itx}|\,f(x)\,dx = \int_{-\infty}^{\infty} f(x)\,dx = 1 . \]Since \(\phi_X(t)\) is in general a complex number, \(|\phi_X(t)| \le 1\) says that it lies in the unit disc of the complex plane. (The textbook's gloss, “lies between \(-1\) and \(+1\)”, describes only the real case, such as a symmetric distribution.)
As for the MGF, with \(it\) in place of \(t\):
\[ \phi_{CX}(t) = E\big(e^{itCX}\big) = \phi_X(Ct), \] \[ \phi_{X_1 + \cdots + X_n}(t) = \phi_{X_1}(t)\cdots\phi_{X_n}(t) \quad\text{(independent } X_i), \] \[ U = \frac{X - a}{h}: \qquad \phi_U(t) = E\big(e^{-ita/h}e^{i(t/h)X}\big) = e^{-ita/h}\,\phi_X(t/h) . \]Printing notes. The textbook ends the first proof with \(M_X(Ct)\); the result is \(\phi_X(Ct)\). In the third it writes \(\phi_X(it/h)\); the \(i\) is already inside the definition of \(\phi\), so the argument is \(t/h\).
If \(E|X|^r < \infty\), differentiating under the expectation gives \(\phi_X^{(r)}(0) = E\big[(iX)^r\big] = i^r\mu_r'\), so \(\mu_r' = \phi_X^{(r)}(0)/i^r\). The CF always exists, and gives every moment that exists. It cannot produce moments that do not exist: the Cauchy CF, \(e^{-|t|}\), is not differentiable at 0, matching the Cauchy distribution's lack of a mean. For such distributions the CF is used to identify the distribution (by uniqueness), not to compute moments.
For a non-negative integer-valued r.v. \(X\) with PMF \(p_k = P(X=k)\), the probability generating function is
\[ P_X(s) \;=\; E(s^X) \;=\; \sum_{k=0}^{\infty} p_k\, s^k, \quad |s| \le 1. \]\(X \sim B(n,p)\): \(P_X(s) = (q + ps)^n\).
\(P'(s) = np(q+ps)^{n-1};\; P'(1) = np = E(X)\) ✓.
\(X \sim \) Poisson(\(\lambda\)): \(P_X(s) = e^{\lambda(s-1)}\).
\(P'(s) = \lambda e^{\lambda(s-1)};\; P'(1) = \lambda = E(X)\) ✓.
For independent \(X_1, \ldots, X_n\), the powers \(s^{X_i}\) are independent, so
\[ P_{X_1 + \cdots + X_n}(s) = E\big(s^{X_1}\cdots s^{X_n}\big) = P_{X_1}(s)\cdots P_{X_n}(s) . \]For \(U = (X - a)/h\), \(s^U = s^{-a/h}\big(s^{1/h}\big)^X\), so
\[ P_U(s) = s^{-a/h}\,P_X\big(s^{1/h}\big) . \]This identity is formal: a PGF is meant for a variable on \(0, 1, 2, \ldots\), and \(U\) is one only for special \(a, h\). The useful case is a whole-number shift, \(h = 1\): \(P_{X + a}(s) = s^a P_X(s)\).
| Function | Definition | Always Exists? | Best For |
|---|---|---|---|
| MGF | \(E(e^{tX})\) | No (e.g., Cauchy) | Computing moments |
| CGF | \(\ln M_X(t)\) | If MGF exists | Cumulants & sums |
| CF | \(E(e^{itX})\) | Yes (always) | Limit theorems / CLT |
| PGF | \(E(s^X)\) for integer X | Yes for \(|s|\le 1\) | Discrete distributions |
A sequence \(\{X_n\}\) converges in probability to \(X\) if for every \(\epsilon > 0\):
\[ \lim_{n \to \infty} P(|X_n - X| \ge \epsilon) \;=\; 0. \]Notation: \(X_n \xrightarrow{P} X\).
\(X_n \xrightarrow{d} X\) if the CDFs converge:
\[ F_{X_n}(x) \to F_X(x) \quad \text{at every continuity point of } F_X. \]Relationship: Convergence in probability ⇒ convergence in distribution. (Reverse is true only when the limit is a constant.)
\(\{X_n\}\) converges to \(X\) almost surely (with probability one) if the set of outcomes on which the sequence of numbers \(X_n(\omega)\) converges to \(X(\omega)\) has probability 1:
\[ P\Big(\big\{\omega \in S : \lim_{n \to \infty} X_n(\omega) = X(\omega)\big\}\Big) = 1, \qquad\text{written}\qquad X_n \xrightarrow{a.s.} X . \]The three modes are nested:
\[ X_n \xrightarrow{a.s.} X \;\Longrightarrow\; X_n \xrightarrow{P} X \;\Longrightarrow\; X_n \xrightarrow{d} X . \]Convergence in probability to a constant \(a\) is the case used by the laws of large numbers: \(\lim_{n\to\infty} P(|X_n - a| < \epsilon) = 1\) for every \(\epsilon > 0\).
Let \(X_1, X_2, \ldots\) be i.i.d. random variables with mean \(\mu\) and finite variance \(\sigma^2\). Let \(\bar{X}_n = (X_1 + \cdots + X_n)/n\). Then for every \(\epsilon > 0\):
i.e., \(\bar X_n \xrightarrow{P} \mu\).
Proof (using Chebyshev's inequality). Let \(X_1,\dots,X_n\) be independent with common mean \(\mu\) and variance \(\sigma^2\), and \(\bar X_n = \frac1n\sum_{i=1}^n X_i\).
Toss a fair coin many times. Let \(X_i = 1\) if head, 0 otherwise. \(\bar X_n\) = proportion of heads. By WLLN, \(\bar X_n \to 0.5\) in probability — empirical frequency justifies the classical definition.
The sample mean of 1 000 dice rolls is very likely close to the true mean \(\mu = 3.5\): by Chebyshev with \(\sigma^2 = 35/12 \approx 2.92\), \(P(|\bar X_{1000} - 3.5| \ge 0.1) \le 2.92/(1000 \cdot 0.01) \approx 0.29\).
Let \(X_1, X_2, \ldots\) have means \(\mu_1, \mu_2, \ldots\), and let \(B_n = \operatorname{Var}(X_1 + \cdots + X_n) < \infty\). If \(B_n/n^2 \to 0\), then for every \(\epsilon > 0\)
\[ P\left\{\left|\frac{X_1 + \cdots + X_n}{n} - \frac{\mu_1 + \cdots + \mu_n}{n}\right| < \epsilon\right\} \to 1 . \]Proof. The average \(\bar X_n\) has mean \(\bar\mu_n = (\mu_1 + \cdots + \mu_n)/n\) and variance \(B_n/n^2\). Chebyshev's inequality with \(c = \epsilon\) gives
\[ P\big(|\bar X_n - \bar\mu_n| \ge \epsilon\big) \le \frac{B_n}{n^2\epsilon^2} \to 0 . \]Neither independence nor a common distribution is assumed, only the variance condition. For i.i.d. variables \(B_n = n\sigma^2\) (the covariances vanish), \(B_n/n^2 = \sigma^2/n \to 0\), and this is the theorem above.
Let \(X_1, X_2, \ldots\) be i.i.d. with mean \(\mu\). Then
\[ P\!\left(\lim_{n \to \infty} \bar X_n = \mu\right) \;=\; 1, \]i.e., \(\bar X_n \to \mu\) almost surely.
WLLN vs. SLLN:
| Property | WLLN | SLLN |
|---|---|---|
| Mode of convergence | In probability | Almost sure (with prob 1) |
| Strength | Weaker | Stronger (implies WLLN) |
| Conditions | Finite variance (sufficient) | Finite mean (Kolmogorov's) |
Let \(X_1, X_2, \ldots\) be i.i.d. with mean \(\mu\) and finite variance \(\sigma^2 > 0\). Define \(\bar X_n = \dfrac{1}{n}\sum_{i=1}^n X_i\). Then
\[ Z_n \;=\; \dfrac{\bar X_n - \mu}{\sigma / \sqrt{n}} \;\xrightarrow{d}\; N(0, 1) \quad \text{as } n \to \infty. \]Equivalently, the standardized sum \(\dfrac{\sum X_i - n\mu}{\sigma\sqrt{n}}\) converges to standard normal.
Diameter of bolts has mean 5 mm, SD 0.1 mm. For a sample of 100 bolts, the sample mean is approximately \(N(5,\; 0.1^2/100) = N(5, 0.0001)\), i.e., SE = 0.01 mm.
\(P(|\bar X - 5| < 0.02) \approx P(|Z| < 2) \approx 0.9544\).
Suppose true proportion of voters supporting party A is 0.6. From a sample of 400, the sample proportion \(\hat p \sim N(0.6,\; 0.6 \times 0.4/400) = N(0.6, 0.000\,6)\), SE ≈ 0.0245.
\(P(|\hat p - 0.6| < 0.05) \approx P(|Z| < 2.04) \approx 0.96\) — explains the typical "± 5 % margin of error".
De Moivre found the normal approximation to the binomial with \(p = \tfrac12\) in 1733; Laplace extended it and stated the central limit theorem in 1812; Liapounoff (Lyapunov) gave the first rigorous general proof in 1901. In its general form: if \(X_1, \ldots, X_n\) are independent with means \(\mu_i\) and variances \(\sigma_i^2\), then under very general conditions \(S_n = X_1 + \cdots + X_n\) is asymptotically normal with
\[ \text{mean } \mu = \sum_{i=1}^n \mu_i, \qquad \text{variance } \sigma^2 = \sum_{i=1}^n \sigma_i^2 . \]\(S_n\) has mean \(n\mu_1\) and standard deviation \(\sigma_1\sqrt n\); \(\bar X_n = S_n/n\) has mean \(\mu_1\) and standard deviation \(\sigma_1/\sqrt n\). With \(\Phi\) the standard normal distribution function, for all \(-\infty < a < b < \infty\):
\[ \text{(i)}\quad \lim_{n\to\infty} P\left\{a \le \frac{S_n - n\mu_1}{\sigma_1\sqrt n} \le b\right\} = \Phi(b) - \Phi(a) = \int_a^b \frac{1}{\sqrt{2\pi}}e^{-x^2/2}\,dx, \] \[ \text{(ii)}\quad \lim_{n\to\infty} P\left\{a \le \frac{S_n - E(S_n)}{\sqrt{\operatorname{Var}(S_n)}} \le b\right\} = \Phi(b) - \Phi(a), \] \[ \text{(iii)}\quad \lim_{n\to\infty} P\left\{a \le \frac{\bar X_n - \mu_1}{\sigma_1/\sqrt n} \le b\right\} = \Phi(b) - \Phi(a) . \]All three standardise the same quantity: dividing the numerator and denominator of (i) by \(n\) gives (iii).
Correction. The textbook's form (i) divides \(S_n - n\mu_1\) by \(\sigma_1/\sqrt n\). That is the standard deviation of \(\bar X_n\), not of \(S_n\); the denominator must be \(\sigma_1\sqrt n\). (With \(n = 100\), \(\sigma_1 = 2\): the two are 20 and 0.2, a hundredfold apart.)
Let \(X_1, \ldots, X_n\) be independent with means \(\mu_i\), variances \(\sigma_i^2\), and finite third absolute central moments \(\rho_i^3 = E|X_i - \mu_i|^3\). Put
\[ \rho^3 = \sum_{i=1}^n \rho_i^3, \qquad \sigma^2 = \sum_{i=1}^n \sigma_i^2 . \]If \(\lim_{n\to\infty} \rho/\sigma = 0\), then \(S_n = X_1 + \cdots + X_n\) is asymptotically \(N\big(\sum\mu_i, \sum\sigma_i^2\big)\).
The condition says that no single variable dominates the sum. For i.i.d. variables \(\rho/\sigma = (n\rho_1^3)^{1/3}/(n\sigma_1^2)^{1/2} = (\rho_1/\sigma_1)\,n^{-1/6} \to 0\), so Lindeberg–Lévy (with a finite third moment) is a special case.
For i.i.d. variables with finite \(\mu\) and \(\sigma^2\), both laws hold. The WLLN says only that \(P(|\bar X_n - \mu| \ge \epsilon) \to 0\). The CLT says how fast, by giving the probability itself, approximately, for large \(n\):
\[ P\big(|\bar X_n - \mu| \ge \epsilon\big) = P\left(\left|\frac{\bar X_n - \mu}{\sigma/\sqrt n}\right| \ge \frac{\epsilon\sqrt n}{\sigma}\right) \approx 1 - \left[\Phi\Big(\frac{\epsilon\sqrt n}{\sigma}\Big) - \Phi\Big(-\frac{\epsilon\sqrt n}{\sigma}\Big)\right] . \]The textbook writes “\(=\)” in the last step; it is an approximation, which the CLT makes good as \(n \to \infty\). Worked Problem 5 puts numbers on it.
For sequences that are independent but not identically distributed, the two laws can part company:
The textbook's generating-function problem, then its three applications of the central limit theorem, then one more problem that puts numbers on its comparison of the CLT with the weak law. The method in the CLT problems is always the same: write the variable as a sum of i.i.d. pieces, find the mean and variance of the sum, and standardise.
Source note. Each result was checked symbolically (the MGF by differentiation, the cumulants by series expansion, the sums by convolution) and the limits numerically. Worked Problem 2 corrects the range of \(p\); Worked Problem 5 is added here and is not in the textbook.
A coin is tossed twice and \(X\) is the number of heads, with distribution \(P(X = 0, 1, 2) = \tfrac14, \tfrac12, \tfrac14\). Find the MGF of \(X\).
\[ M_X(t) = \sum_x e^{tx}P(X = x) = e^{0}\cdot\tfrac14 + e^{t}\cdot\tfrac12 + e^{2t}\cdot\tfrac14 = \frac{e^{2t} + 2e^t + 1}{4} = \Big(\frac{1 + e^t}{2}\Big)^2 . \]The square is no accident: \(X\) is the sum of two independent Bernoulli(\(\tfrac12\)) variables, each with MGF \(\tfrac12 + \tfrac12e^t\), and by the additive property the MGF of the sum is the product.
Moments from it. Differentiate and put \(t = 0\):
\[ M_X'(t) = \frac{e^t + e^{2t}}{2} \Rightarrow \mu_1' = M_X'(0) = 1, \qquad M_X''(t) = \frac{e^t + 2e^{2t}}{2} \Rightarrow \mu_2' = \frac32, \] \[ \operatorname{Var}(X) = \frac32 - 1^2 = \frac12 . \]Cumulants from it. \(K_X(t) = 2\log\frac{1 + e^t}{2}\), so \(K_X'(t) = \dfrac{2e^t}{1 + e^t}\) and \(K_X''(t) = \dfrac{2e^t}{(1 + e^t)^2}\), giving \(K_1 = 1\) and \(K_2 = \tfrac24 = \tfrac12\), the mean and variance again. Carrying on, \(K_3 = 0\) (the distribution is symmetric) and \(K_4 = -\tfrac14\).
PGF. \(P_X(s) = \tfrac14 + \tfrac12s + \tfrac14s^2 = \big(\tfrac{1 + s}{2}\big)^2\), and \(P_X'(1) = 1\), the mean once more.
Discuss the distribution of the sum of \(n\) i.i.d. binomial variates.
Let \(X_1, \ldots, X_n\) be i.i.d. \(B(r, p)\), so \(E(X_i) = rp\) and \(\operatorname{Var}(X_i) = rpq\). For \(S_n = X_1 + \cdots + X_n\), the addition theorem and independence give
\[ E(S_n) = \sum_{i=1}^n rp = nrp, \qquad \operatorname{Var}(S_n) = \sum_{i=1}^n rpq = nrpq . \]By the Lindeberg–Lévy theorem, form (ii),
\[ \lim_{n\to\infty} P\left\{a \le \frac{S_n - nrp}{\sqrt{nrpq}} \le b\right\} = \Phi(b) - \Phi(a), \qquad 0 < p < 1 . \]Correction. The textbook gives the range as \(0 \le p < 1\). At \(p = 0\) every \(X_i\) is 0, the variance \(nrpq\) is 0, and the standardised variable is not defined; the theorem needs \(0 < \sigma^2 < \infty\), so \(0 < p < 1\).
An exact check. Here the limit is not even needed for the shape: by the PGF, \(P_{S_n}(s) = \big[(q + ps)^r\big]^n = (q + ps)^{nr}\), so \(S_n \sim B(nr, p)\) exactly (for \(r = 3\), \(n = 4\), \(p = \tfrac14\): the convolution of four \(B(3, \tfrac14)\) is \(B(12, \tfrac14)\), with mean 3 and variance \(\tfrac94\)). The CLT then says that a binomial with many trials is close to normal.
Prove that if \(Y_n \sim B(n, p)\), then \(\displaystyle\lim_{n\to\infty} P\left\{a \le \frac{Y_n - np}{\sqrt{np(1 - p)}} \le b\right\} = \Phi(b) - \Phi(a)\), \(0 < p < 1\).
Let \(X_1, \ldots, X_n\) be i.i.d. Bernoulli, \(B(1, p)\): each has mean \(p\) and variance \(pq\). Their sum counts the successes in \(n\) independent trials, so it has the same distribution as \(Y_n\), with
\[ E(Y_n) = np, \qquad \operatorname{Var}(Y_n) = npq = np(1 - p) . \]The Lindeberg–Lévy theorem applied to this sum gives the result. This is the De Moivre–Laplace theorem, the oldest case of the CLT.
If \(Y_n \sim\) Poisson(\(n\)), show that \(\displaystyle\lim_{n\to\infty} P(Y_n \le n) = \tfrac12\), i.e. \(\displaystyle\sum_{k=0}^{n} \frac{e^{-n}n^k}{k!} \to \frac12\).
Write \(Y_n\) as a sum. Let \(X_1, \ldots, X_n\) be i.i.d. Poisson(1). By the PGF, \(P_{X_1 + \cdots + X_n}(s) = \big[e^{(s - 1)}\big]^n = e^{n(s - 1)}\), the PGF of Poisson(\(n\)). So \(Y_n\) is such a sum, with \(E(Y_n) = \operatorname{Var}(Y_n) = n\).
Apply the CLT. \((Y_n - n)/\sqrt n\) tends in distribution to \(N(0, 1)\). Since \(\Phi\) is continuous, the distribution functions converge at every point, in particular at 0:
\[ P(Y_n \le n) = P\left\{\frac{Y_n - n}{\sqrt n} \le 0\right\} \to \Phi(0) = \frac12 . \](This is the textbook's \(a = -\infty\), \(b = 0\): \(\Phi(0) - \Phi(-\infty) = \tfrac12\).) Writing out the Poisson probabilities, \(\sum_{k=0}^{n} e^{-n}n^k/k! \to \tfrac12\).
How fast? Exactly computed:
| \(n\) | 1 | 10 | 100 | 1000 | 5000 |
|---|---|---|---|---|---|
| \(P(Y_n \le n)\) | 0.7358 | 0.5830 | 0.5266 | 0.5084 | 0.5038 |
(At \(n = 1\) it is \(2e^{-1}\).) The approach is slow, and from above, for two reasons that push the same way: the event \(Y_n \le n\) counts the whole probability at the single value \(n\), about \(1/\sqrt{2\pi n}\), and the Poisson is skewed to the right. Together they leave an excess of about \(2/(3\sqrt{2\pi n})\) over \(\tfrac12\) (\(0.5266\) predicted and found at \(n = 100\)).
A fair coin is tossed \(n = 400\) times and \(\bar X\) is the proportion of heads. Bound \(P(|\bar X - \tfrac12| \ge 0.05)\) by Chebyshev's inequality, approximate it by the CLT, and compare with the exact value.
Each toss has \(\mu = \tfrac12\), \(\sigma^2 = \tfrac14\), so \(\operatorname{Var}(\bar X) = \tfrac{1}{4 \cdot 400}\) and \(\sigma/\sqrt n = \tfrac{0.5}{20} = 0.025\).
Chebyshev (the WLLN's tool):
\[ P\big(|\bar X - \tfrac12| \ge 0.05\big) \le \frac{\sigma^2}{n\epsilon^2} = \frac{0.25}{400 \times 0.0025} = 0.25 . \]CLT: \(\epsilon\sqrt n/\sigma = 0.05 \times 20/0.5 = 2\), so
\[ P\big(|\bar X - \tfrac12| \ge 0.05\big) \approx P(|Z| \ge 2) = 2\big[1 - \Phi(2)\big] = 0.0455 . \]Exact (binomial, 180 heads or fewer, or 220 or more): \(0.0510\).
The WLLN guarantees only that this probability goes to 0; Chebyshev caps it at 0.25. The CLT gives its actual size, within half a percentage point here.