Let \(X_1, X_2, \ldots\) be random variables, \(S_n = \sum_{i=1}^{n} X_i\) and \(\bar{X}_n = S_n / n\).
A weak law of large numbers (WLLN) asserts convergence in probability:
\[ \bar{X}_n - E\left(\bar{X}_n\right) \xrightarrow{P} 0. \]A strong law (SLLN) asserts the same thing almost surely:
\[ \bar{X}_n - E\left(\bar{X}_n\right) \xrightarrow{a.s.} 0. \]By the implication chain of Unit 3, every strong law contains the corresponding weak law; the converse fails, and Example 3.2 of that unit is the reason such a distinction has to be drawn at all.
Bernoulli's is the special case of Chebyshev's for indicator variables. Khintchine's is the important one: it asks only that the mean exist — no variance, finite or otherwise. It cannot be proved by Chebyshev's inequality, and is proved instead with characteristic functions.
Statement. If \(X_1, X_2, \ldots\) are independent with \(\operatorname{Var}(\bar{X}_n) \to 0\), then \(\bar{X}_n - E(\bar{X}_n) \xrightarrow{P} 0\).
Step 1 — write the variance of the mean. Independence makes the variance of a sum the sum of the variances, and \(\operatorname{Var}(aY) = a^{2}\operatorname{Var}(Y)\), so
\[ \operatorname{Var}\left(\bar{X}_n\right) = \operatorname{Var}\left(\frac{1}{n}\sum_{i=1}^{n} X_i\right) = \frac{1}{n^{2}} \sum_{i=1}^{n} \operatorname{Var}(X_i) = \frac{1}{n^{2}} \sum_{i=1}^{n} \sigma_i^{2}. \]Step 2 — apply Chebyshev's inequality to \(\bar{X}_n\). From Unit 2, for any \(\varepsilon > 0\),
\[ P\left(\left|\bar{X}_n - E\left(\bar{X}_n\right)\right| \ge \varepsilon\right) \;\le\; \frac{\operatorname{Var}\left(\bar{X}_n\right)}{\varepsilon^{2}} \;=\; \frac{1}{n^{2}\varepsilon^{2}} \sum_{i=1}^{n} \sigma_i^{2}. \]Step 3 — let \(n \to \infty\). By hypothesis the numerator of the middle expression tends to \(0\) while \(\varepsilon^{2}\) stays fixed, so the bound tends to \(0\), and the probability it bounds is squeezed to \(0\). \(\blacksquare\)
The commonest sufficient condition. If all the variances are bounded by one constant, \(\sigma_i^{2} \le C\), then \(\operatorname{Var}(\bar{X}_n) \le nC/n^{2} = C/n \to 0\), so the hypothesis holds automatically. But boundedness is not necessary, as the next example shows.
Given. Independent \(X_k\) taking the values \(\pm k^{1/4}\) with probability \(\tfrac12\) each.
Asked. Show that Chebyshev's WLLN applies even though \(\operatorname{Var}(X_k) \to \infty\).
Step 1 — mean and variance of one term. By symmetry \(E(X_k) = \tfrac12 k^{1/4} + \tfrac12(-k^{1/4}) = 0\), so
\[ \operatorname{Var}(X_k) = E\left(X_k^{2}\right) = \tfrac12 \left(k^{1/4}\right)^{2} + \tfrac12 \left(-k^{1/4}\right)^{2} = \left(k^{1/4}\right)^{2} = \sqrt{k}. \]This grows without bound, so the "bounded variances" shortcut does not apply.
Step 2 — the variance of the mean. By Step 1 of the proof above,
\[ \operatorname{Var}\left(\bar{X}_n\right) = \frac{1}{n^{2}} \sum_{k=1}^{n} \sqrt{k}. \]Step 3 — evaluate it.
| \(n\) | \(\dfrac{1}{n^{2}}\sum_{k \le n} \sqrt{k}\) |
|---|---|
| 100 | 0.0671463 |
| 10,000 | 0.0066672 |
| 1,000,000 | 0.0006667 |
Step 4 — identify the rate. A tenfold rise in \(\sqrt{n}\) divides the value by ten, which says \(\operatorname{Var}(\bar{X}_n) \asymp n^{-1/2}\). The reason is that \(\sum_{k \le n} \sqrt{k} \approx \tfrac23 n^{3/2}\), so the ratio is about \(\tfrac23 n^{3/2}/n^{2} = \tfrac23 n^{-1/2}\); at \(n = 10^{6}\) that is \(\tfrac23 \times 10^{-3} = 0.000667\), matching the table.
Step 5 — conclude. \(\operatorname{Var}(\bar{X}_n) \to 0\), so Chebyshev's WLLN applies and \(\bar{X}_n \xrightarrow{P} 0\).
Interpretation. The condition that matters is on the average, not on the individual terms. Variances may grow, provided they grow more slowly than \(n\) so that dividing by \(n^{2}\) still wins.
For an i.i.d. sample from a distribution with a mean but infinite variance — a Pareto with shape parameter \(\alpha\) between 1 and 2, for instance — Chebyshev's law is unusable, because the bound in its proof is \(\infty/\varepsilon^{2}\). Khintchine's law still gives \(\bar{X}_n \xrightarrow{P} \mu\).
For the Cauchy distribution even Khintchine fails, because no mean exists. The characteristic function shows what happens instead. With \(\phi_X(t) = e^{-|t|}\) for the standard Cauchy, property (4) of Unit 2 gives
\[ \phi_{\bar{X}_n}(t) = \left[\phi_X\!\left(\frac{t}{n}\right)\right]^{n} = \left(e^{-|t|/n}\right)^{n} = e^{-|t|} = \phi_X(t). \]So \(\bar{X}_n\) has exactly the same distribution as a single observation, for every \(n\). Averaging a Cauchy sample achieves nothing at all, however large the sample.
Statement. Let \(X_1, \ldots, X_n\) be independent with \(E(X_i) = 0\) and finite variances, and \(S_k = X_1 + \cdots + X_k\). Then for every \(\lambda > 0\),
\[ P\left(\max_{1 \le k \le n} |S_k| \ge \lambda\right) \;\le\; \frac{\operatorname{Var}(S_n)}{\lambda^{2}}. \]Why it is stronger than Chebyshev. Chebyshev applied to \(S_n\) alone bounds \(P(|S_n| \ge \lambda)\). Kolmogorov bounds the probability that the partial sum is ever large — a much bigger event — by the same quantity. Controlling a whole path rather than one endpoint is exactly what an almost sure statement needs, which is why the strong laws are proved from it.
Given. \(X_1, X_2, X_3\) independent, each \(+1\) or \(-1\) with probability \(\tfrac12\). Take \(\lambda = 2\).
Asked. Compute the bound and the exact probability.
Step 1 — the bound. Each \(X_i\) has mean \(0\) and variance \(E(X_i^{2}) = 1\). By independence \(\operatorname{Var}(S_3) = 1 + 1 + 1 = 3\), so
\[ P\left(\max_{1 \le k \le 3} |S_k| \ge 2\right) \le \frac{3}{2^{2}} = \frac{3}{4} = 0.75. \]Step 2 — enumerate. There are \(2^{3} = 8\) equally likely sign patterns. For each, list \(S_1, S_2, S_3\) and the largest absolute value reached:
| \((X_1,X_2,X_3)\) | \(S_1, S_2, S_3\) | \(\max_k |S_k|\) | \(\ge 2\)? |
|---|---|---|---|
| \((+,+,+)\) | 1, 2, 3 | 3 | yes |
| \((+,+,-)\) | 1, 2, 1 | 2 | yes |
| \((+,-,+)\) | 1, 0, 1 | 1 | no |
| \((+,-,-)\) | 1, 0, −1 | 1 | no |
| \((-,+,+)\) | −1, 0, 1 | 1 | no |
| \((-,+,-)\) | −1, 0, −1 | 1 | no |
| \((-,-,+)\) | −1, −2, −1 | 2 | yes |
| \((-,-,-)\) | −1, −2, −3 | 3 | yes |
Step 3 — count. Four of the eight patterns reach \(2\), so
\[ P\left(\max_{1 \le k \le 3} |S_k| \ge 2\right) = \frac{4}{8} = 0.5. \]Step 4 — compare. \(0.5 \le 0.75\), so the inequality holds, with the bound exceeding the truth by half again.
Interpretation. Notice rows 2 and 7: the path touches \(\pm 2\) at \(k = 2\) and comes back, so \(|S_3| = 1 < 2\). Chebyshev on \(S_3\) alone would miss those two patterns entirely and bound only \(P(|S_3| \ge 2) = 2/8 = 0.25\). Kolmogorov catches the excursion, and that is precisely the extra strength being paid for.
The i.i.d. form is remarkable for being an equivalence. A finite mean is not merely sufficient for the sample mean to converge almost surely — it is necessary. If \(E|X_1| = \infty\) then \(\bar{X}_n\) almost surely fails to converge to anything finite. It is the sharpest statement in this unit, and it is the theorem that justifies calling \(\bar{X}\) an estimator of \(\mu\) at all.
where \(s_n^{2} = \sum_{i=1}^{n} \operatorname{Var}(X_i)\). In the last two the conclusion is the same, \(\left(S_n - \sum \mu_i\right)/s_n \xrightarrow{d} N(0,1)\); what changes is the condition under which it holds.
How they relate. De Moivre–Laplace is Lindeberg–Lévy for Bernoulli summands. Liapunov's condition implies Lindeberg's, so Lindeberg–Feller is the more general; but Liapunov's is easier to check, needing only a \((2+\delta)\)-th absolute moment, and \(\delta = 1\) usually suffices. Lindeberg's condition is in fact necessary as well as sufficient, under the mild extra requirement that no single term dominates the sum.
What they all say. A sum of many independent contributions, none of them dominant, is approximately normal whatever the shape of the individual terms. That is the reason the normal distribution appears so widely in applied work, and the reason the tests of Inferential Statistics, Unit 3 may use normal critical values for statistics built from non-normal data.
Given. \(X \sim \text{Bin}(100, 0.5)\).
Asked. Find \(P(X \ge 60)\) exactly and by normal approximation, with and without the continuity correction, and compare.
Step 1 — the exact value. Summing the binomial probabilities,
\[ P(X \ge 60) = \sum_{k=60}^{100} \binom{100}{k} (0.5)^{100} = 0.0284440. \]Step 2 — the normal parameters.
\[ \mu = np = 100 \times 0.5 = 50, \qquad \sigma = \sqrt{npq} = \sqrt{100 \times 0.5 \times 0.5} = \sqrt{25} = 5. \]Step 3 — the naive approximation. Standardising at the value \(60\) itself,
\[ z = \frac{60 - 50}{5} = 2.00, \qquad P(X \ge 60) \approx 1 - \Phi(2.00) = 1 - 0.977250 = 0.022750. \]Step 4 — the correction for continuity. \(X\) is integer-valued while the normal is continuous. The event \(X \ge 60\) corresponds to the bar centred at \(60\) and every bar to its right, and those bars begin at \(59.5\), not at \(60\). So standardise at \(59.5\):
\[ z = \frac{59.5 - 50}{5} = \frac{9.5}{5} = 1.90, \qquad P(X \ge 60) \approx 1 - \Phi(1.90) = 1 - 0.971283 = 0.028717. \]Step 5 — compare the two errors.
| method | value | error against 0.0284440 |
|---|---|---|
| exact binomial | 0.0284440 | — |
| normal, no correction (\(z = 2.00\)) | 0.0227501 | 0.0056938 |
| normal, with correction (\(z = 1.90\)) | 0.0287166 | 0.0002726 |
Step 6 — read the comparison. The correction reduces the error by a factor of \(0.0056938 / 0.0002726 = 20.9\). Without it the approximation understates the tail by \(20\%\) of its own size; with it the error is under \(1\%\).
Interpretation. Half a unit sounds like a detail and is not. It matters most exactly where these calculations are used — in the tail, where the probability is small and the relative error is therefore large. The figure below shows why: the shaded bars start at \(59.5\), so cutting the normal curve at \(60\) discards half of the bar at \(60\), which is the tallest bar in the whole region being summed.
Given. A population with mean \(\mu = 50\) and standard deviation \(\sigma = 8\), shape unknown. A random sample of \(n = 64\) is drawn.
Asked. Find \(P(\bar{X} > 52)\) and \(P(48.5 < \bar{X} < 52)\).
Step 1 — check the theorem applies. The observations are i.i.d. with a finite variance. That is the whole hypothesis of Lindeberg–Lévy; the shape of the population is not needed and is not assumed.
Step 2 — the standard error.
\[ SE = \frac{\sigma}{\sqrt{n}} = \frac{8}{\sqrt{64}} = \frac{8}{8} = 1. \]So \(\bar{X}\) is approximately \(N(50, 1^{2})\).
Step 3 — the first probability.
\[ z = \frac{52 - 50}{1} = 2.00, \qquad P(\bar{X} > 52) \approx 1 - \Phi(2.00) = 1 - 0.977250 = 0.022750. \]Step 4 — the second probability. Standardise both endpoints:
\[ z_1 = \frac{48.5 - 50}{1} = -1.50, \qquad z_2 = \frac{52 - 50}{1} = 2.00, \] \[ P(48.5 < \bar{X} < 52) \approx \Phi(2.00) - \Phi(-1.50) = 0.977250 - 0.066807 = 0.910443. \]Step 5 — a consistency check. The three pieces of the line must account for the whole probability:
\[ \underbrace{0.066807}_{\bar{X} < 48.5} + \underbrace{0.910443}_{\text{middle}} + \underbrace{0.022750}_{\bar{X} > 52} = 1.000000. \checkmark \]Interpretation. No distributional assumption about the population was used anywhere. Dividing \(\sigma\) by \(\sqrt{n}\) shrank the spread from \(8\) to \(1\), and it is that shrinkage — the \(\sqrt{n}\) rate — that makes a sample of \(64\) informative about a population with a standard deviation eight times larger.