Let \(X_1, X_2, \ldots\) and \(X\) be random variables on a common probability space.
(a) Convergence in probability, written \(X_n \xrightarrow{P} X\): for every \(\varepsilon > 0\),
\[ \lim_{n \to \infty} P\big(|X_n - X| > \varepsilon\big) = 0. \](b) Almost sure convergence, written \(X_n \xrightarrow{a.s.} X\):
\[ P\left(\left\{\omega : \lim_{n \to \infty} X_n(\omega) = X(\omega)\right\}\right) = 1. \](c) Convergence in quadratic mean (the case \(r = 2\) of convergence in \(r\)-th mean), written \(X_n \xrightarrow{q.m.} X\):
\[ \lim_{n \to \infty} E\left(X_n - X\right)^{2} = 0. \](d) Convergence in law (in distribution), written \(X_n \xrightarrow{d} X\): with \(F_n\) and \(F\) the distribution functions,
\[ \lim_{n \to \infty} F_n(x) = F(x) \quad \text{at every } x \text{ where } F \text{ is continuous.} \]Both say "\(X_n\) gets close to \(X\)", and the difference is where the limit sits.
In probability, the limit is taken outside: for each fixed \(n\) we measure the set on which \(X_n\) is still far from \(X\), and ask that its probability shrink. Nothing forbids the bad set from moving around, so a given \(\omega\) may be inside it infinitely often.
Almost surely, the limit is taken inside: we look at one \(\omega\) at a time, ask whether the ordinary numerical sequence \(X_n(\omega)\) converges, and require the set of \(\omega\) for which it does to have probability 1. A given \(\omega\) may be in the bad set only finitely often.
That is exactly why almost sure convergence is the stronger of the two, and why the counterexample of Example 3.2 is possible.
None of the first three reverses, and almost sure convergence and quadratic mean convergence do not imply one another in either direction.
Statement. If \(E(X_n - X)^{2} \to 0\) then \(X_n \xrightarrow{P} X\).
Step 1 — fix the tolerance. Let \(\varepsilon > 0\) be arbitrary and held fixed throughout.
Step 2 — apply Markov's inequality. From Unit 2, \(P(|Z| \ge a) \le E|Z|^{r}/a^{r}\). Take \(Z = X_n - X\), \(a = \varepsilon\), \(r = 2\):
\[ P\big(|X_n - X| > \varepsilon\big) \;\le\; P\big(|X_n - X| \ge \varepsilon\big) \;\le\; \frac{E\left(X_n - X\right)^{2}}{\varepsilon^{2}}. \]The first inequality is monotonicity of \(P\), because \(\{|X_n - X| > \varepsilon\} \subseteq \{|X_n - X| \ge \varepsilon\}\).
Step 3 — let \(n \to \infty\). The numerator tends to \(0\) by hypothesis and \(\varepsilon^{2}\) is a fixed positive constant, so the whole right-hand side tends to \(0\). A non-negative quantity squeezed below something tending to \(0\) tends to \(0\), so
\[ P\big(|X_n - X| > \varepsilon\big) \longrightarrow 0. \]Since \(\varepsilon\) was arbitrary, this is convergence in probability. \(\blacksquare\)
Given. A sequence with
\[ P\left(X_n = n\right) = \frac{1}{n}, \qquad P\left(X_n = 0\right) = 1 - \frac{1}{n}. \]Asked. Show \(X_n \xrightarrow{P} 0\) but that \(X_n\) converges neither in quadratic mean nor in mean to \(0\).
Step 1 — convergence in probability. Fix \(\varepsilon > 0\). For all \(n\) large enough that \(n > \varepsilon\), the only way \(|X_n - 0| > \varepsilon\) can happen is \(X_n = n\). Hence
\[ P\big(|X_n| > \varepsilon\big) = P(X_n = n) = \frac{1}{n} \longrightarrow 0. \]Reading the values: at \(n = 10\) this is \(0.1\); at \(n = 100\), \(0.01\); at \(n = 10{,}000\), \(0.0001\). So \(X_n \xrightarrow{P} 0\).
Step 2 — the first moment.
\[ E(X_n) = n \cdot \frac{1}{n} + 0 \cdot \left(1 - \frac{1}{n}\right) = 1 \quad \text{for every } n. \]So \(E|X_n - 0| = 1 \not\to 0\): there is no convergence in mean.
Step 3 — the second moment.
\[ E\left(X_n - 0\right)^{2} = n^{2} \cdot \frac{1}{n} + 0 = n \longrightarrow \infty. \]Values: \(1, 10, 100\) at \(n = 1, 10, 100\). Far from tending to \(0\), it diverges, so there is no convergence in quadratic mean either.
Interpretation. The probability of the bad outcome vanishes, but the size of the bad outcome grows faster than its probability shrinks, and expectation multiplies the two. This is the same mechanism as the spike sequence of Unit 1, Example 1.4, now in probabilistic dress. It is why a consistent estimator need not be asymptotically unbiased, and why the two properties must be checked separately.
Given. Take \(\Omega = (0,1]\) with \(P\) = Lebesgue measure (the uniform distribution). Lay out the intervals in blocks: block \(k\) consists of the \(k\) intervals
\[ \left(\frac{j-1}{k}, \frac{j}{k}\right], \qquad j = 1, 2, \ldots, k, \]and list all of them in one sequence \(A_1; A_2, A_3; A_4, A_5, A_6; \ldots\), so that block 1 contributes 1 interval, block 2 contributes 2, block 3 contributes 3, and so on. Put \(X_n = \mathbf{1}_{A_n}\).
Asked. Show \(X_n \xrightarrow{P} 0\) but that \(X_n(\omega)\) converges for no \(\omega\) at all.
Step 1 — convergence in probability. If \(A_n\) sits in block \(k\), its length is \(1/k\), so for \(0 < \varepsilon < 1\),
\[ P\big(|X_n - 0| > \varepsilon\big) = P(A_n) = \frac{1}{k}. \]As \(n \to \infty\) the block index \(k\) also \(\to \infty\), because each block is finite, so this probability tends to \(0\). Hence \(X_n \xrightarrow{P} 0\).
Step 2 — no pointwise convergence. Fix any \(\omega \in (0,1]\). The \(k\) intervals of block \(k\) partition \((0,1]\), so exactly one of them contains \(\omega\). Therefore, in every block, \(X_n(\omega) = 1\) for exactly one \(n\) and \(X_n(\omega) = 0\) for the other \(k - 1\).
Step 3 — read off the consequence. Since there are infinitely many blocks, the numerical sequence \(X_n(\omega)\) contains infinitely many \(1\)s and (from block 2 onwards) infinitely many \(0\)s. A sequence taking both values infinitely often does not converge. This is true for every \(\omega\), so the set of \(\omega\) where \(X_n(\omega) \to 0\) is empty and has probability \(0\), not \(1\).
Interpretation. The bad set shrinks, but it sweeps across \(\Omega\) rather than settling down, so no single point is eventually spared. This is exactly the distinction drawn in words in section 1, made concrete. Note also that \(E(X_n - 0)^{2} = P(A_n) = 1/k \to 0\), so this sequence does converge in quadratic mean — which shows quadratic mean convergence does not imply almost sure convergence either.
Given. Let \(X \sim N(0,1)\) and define \(X_n = -X\) for every \(n\).
Asked. Show \(X_n \xrightarrow{d} X\) but \(X_n \not\xrightarrow{P} X\).
Step 1 — convergence in distribution. The standard normal is symmetric about \(0\), so \(-X\) has the same distribution as \(X\). Hence \(F_n = F\) for every \(n\), and a constant sequence trivially converges: \(F_n(x) \to F(x)\) at every \(x\). So \(X_n \xrightarrow{d} X\).
Step 2 — failure in probability. Here \(X_n - X = -X - X = -2X\), so for \(\varepsilon = 1\),
\[ P\big(|X_n - X| > 1\big) = P\big(|{-2X}| > 1\big) = P\left(|X| > \tfrac12\right) = 2\big[1 - \Phi(0.5)\big] = 2(1 - 0.691462) = 0.617075. \]This does not depend on \(n\) at all, so it certainly does not tend to \(0\).
Interpretation. Convergence in distribution is a statement about the law of \(X_n\) and says nothing about how close \(X_n\) and \(X\) are as functions on \(\Omega\). Here they are as far apart as a symmetric variable can be from its own negative, while having identical distributions. The one case where the converse does hold is when the limit is a constant \(c\): there is then no room for the two to differ in sign or anywhere else, and \(X_n \xrightarrow{d} c\) does give \(X_n \xrightarrow{P} c\).
Suppose \(X_n \xrightarrow{d} X\) and \(Y_n \xrightarrow{P} c\), where \(c\) is a constant. Then
\[ X_n + Y_n \xrightarrow{d} X + c, \qquad X_n Y_n \xrightarrow{d} cX, \qquad \frac{X_n}{Y_n} \xrightarrow{d} \frac{X}{c} \ \ (c \ne 0). \]The constancy of the limit of \(Y_n\) is essential; the theorem is false if \(Y_n\) converges in distribution to a non-degenerate variable.
Given. \(X_1, \ldots, X_n\) i.i.d. with mean \(\mu\) and finite variance \(\sigma^{2}\); \(\bar{X}\) the sample mean and \(S\) the sample standard deviation. Two facts are available: the central limit theorem of Unit 4 gives
\[ Z_n = \frac{\bar{X} - \mu}{\sigma / \sqrt{n}} \xrightarrow{d} N(0,1), \]and the weak law of large numbers gives \(S^{2} \xrightarrow{P} \sigma^{2}\), hence \(\sigma / S \xrightarrow{P} 1\) by the continuous mapping theorem.
Asked. Deduce the limiting distribution of \(T_n = (\bar{X} - \mu)/(S/\sqrt{n})\), and use it on a sample with \(n = 100\), \(\bar{x} = 52.4\), \(s = 8\).
Step 1 — write \(T_n\) as a product. Multiply and divide by \(\sigma\):
\[ T_n = \frac{\bar{X} - \mu}{S/\sqrt{n}} = \frac{\bar{X} - \mu}{\sigma/\sqrt{n}} \times \frac{\sigma}{S} = Z_n \cdot \frac{\sigma}{S}. \]Step 2 — identify the two factors. \(Z_n \xrightarrow{d} N(0,1)\) and \(\sigma/S \xrightarrow{P} 1\), a constant. Slutzky's product form applies with \(c = 1\).
Step 3 — conclude.
\[ T_n \xrightarrow{d} 1 \cdot N(0,1) = N(0,1). \]Step 4 — use it. An approximate \(95\%\) confidence interval for \(\mu\) is \(\bar{x} \pm 1.96 \, s/\sqrt{n}\). Here
\[ \frac{s}{\sqrt{n}} = \frac{8}{\sqrt{100}} = \frac{8}{10} = 0.8, \qquad 1.96 \times 0.8 = 1.568, \] \[ \text{interval} = 52.4 \pm 1.568 = (50.832,\ 53.968). \]Interpretation. Replacing the unknown \(\sigma\) by the estimate \(S\) costs nothing asymptotically, and Slutzky is the theorem that says so. It is why the large-sample tests of Inferential Statistics, Unit 3 may use \(s\) in place of \(\sigma\) and still refer to the normal table.
For a sequence of events \(A_1, A_2, \ldots\), define
\[ \limsup_{n \to \infty} A_n \;=\; \bigcap_{m=1}^{\infty} \bigcup_{n=m}^{\infty} A_n \;=\; \{A_n \text{ occurs infinitely often}\}, \]usually written \(\{A_n \text{ i.o.}\}\). Reading the formula: \(\omega\) belongs to it exactly when, for every starting index \(m\), some \(A_n\) with \(n \ge m\) contains \(\omega\) — that is, the \(A_n\) never stop catching \(\omega\).
Statement. If \(\sum_{n=1}^{\infty} P(A_n) < \infty\) then \(P(A_n \text{ i.o.}) = 0\). No independence is required.
Step 1 — bound by a tail union. For every \(m\),
\[ \{A_n \text{ i.o.}\} = \bigcap_{k=1}^{\infty} \bigcup_{n=k}^{\infty} A_n \;\subseteq\; \bigcup_{n=m}^{\infty} A_n, \]since an intersection is contained in each of the sets being intersected. Monotonicity of \(P\) gives \(P(A_n \text{ i.o.}) \le P\left(\bigcup_{n \ge m} A_n\right)\).
Step 2 — apply countable subadditivity. Property (P2) of Unit 1 gives
\[ P\left(\bigcup_{n=m}^{\infty} A_n\right) \le \sum_{n=m}^{\infty} P(A_n). \]Step 3 — let the tail vanish. A convergent series has tails tending to zero, so \(\sum_{n \ge m} P(A_n) \to 0\) as \(m \to \infty\). Combining Steps 1 and 2, the fixed number \(P(A_n \text{ i.o.})\) is below a quantity that can be made arbitrarily small, so it is \(0\). \(\blacksquare\)
Statement. If the \(A_n\) are independent and \(\sum_{n=1}^{\infty} P(A_n) = \infty\), then \(P(A_n \text{ i.o.}) = 1\).
Independence is indispensable here. Without it the conclusion fails: take a single event \(A\) with \(P(A) = \tfrac12\) and set \(A_n = A\) for every \(n\). Then \(\sum_n P(A_n) = \infty\), but \(\{A_n \text{ i.o.}\} = A\), whose probability is \(\tfrac12\), not \(1\).
Putting the two lemmas together: for independent events, \(P(A_n \text{ i.o.})\) is either \(0\) or \(1\), and which one is decided entirely by whether \(\sum_n P(A_n)\) converges or diverges. There is no middle ground and no third case to check.
Given. Independent events with (a) \(P(A_n) = 1/n^{2}\) and (b) \(P(B_n) = 1/n\).
Asked. Find \(P(A_n \text{ i.o.})\) and \(P(B_n \text{ i.o.})\).
Step 1 — case (a), test the series. \(\sum_{n \ge 1} 1/n^{2}\) is the Basel series and converges to \(\pi^{2}/6 = 1.644934\). Partial sums confirm the convergence: summing to \(n = 4000\) gives \(1.644684\), already within \(0.00025\) of the limit.
Step 2 — apply the first lemma. The series converges, so \(P(A_n \text{ i.o.}) = 0\). Only finitely many \(A_n\) occur, with probability 1.
Step 3 — put a number on "finitely many". The proof's own bound is usable. The probability that any \(A_n\) with \(n \ge m\) occurs at all is at most \(\sum_{n \ge m} 1/n^{2}\), and that tail is
| \(m\) | \(\sum_{n \ge m} 1/n^{2}\) | reading |
|---|---|---|
| 10 | 0.10517 | about a 1-in-10 chance of any event from \(A_{10}\) onwards |
| 100 | 0.01005 | about 1 in 100 |
| 1000 | 0.00100 | about 1 in 1000 |
Step 4 — case (b), test the series. \(\sum_{n \ge 1} 1/n\) is the harmonic series and diverges. Its growth is slow but unbounded: the partial sum to \(n = 4000\) is \(8.871390\), and to \(n = 40{,}000\) it is \(11.173863\) — a tenfold increase in \(n\) adds only about \(2.30 \approx \ln 10\), which is the signature of logarithmic divergence.
Step 5 — apply the second lemma. The events are independent and the series diverges, so \(P(B_n \text{ i.o.}) = 1\). Infinitely many \(B_n\) occur, with certainty.
Interpretation. The two cases sit either side of the borderline \(\sum 1/n^{p}\), which converges for \(p > 1\) and diverges for \(p \le 1\). Note how little it takes to cross: \(P(A_n) = 1/n^{2}\) gives "finitely many, almost surely", while \(P(B_n) = 1/n\) — larger, but still tending to zero — gives "infinitely many, almost surely". Individual probabilities tending to zero say nothing on their own; only the series does.
Let \(X_1, \ldots, X_n\) be i.i.d. with distribution function \(F\), and let
\[ F_n(x) = \frac{1}{n} \sum_{i=1}^{n} \mathbf{1}_{\{X_i \le x\}} \]be the empirical distribution function — the proportion of the sample at or below \(x\). Then
\[ \sup_{x \in \mathbb{R}} \left| F_n(x) - F(x) \right| \xrightarrow{a.s.} 0. \]Why it matters. For each fixed \(x\), \(F_n(x) \xrightarrow{a.s.} F(x)\) is just the strong law of large numbers applied to the indicators. Glivenko–Cantelli says far more: the convergence is uniform in \(x\), so the worst discrepancy anywhere on the line goes to zero. That is the theorem behind the Kolmogorov–Smirnov goodness-of-fit test, whose statistic is exactly this supremum, and behind the bootstrap of Estimation Theory (STS-201), Unit 2, which resamples from \(F_n\) on the grounds that it is uniformly close to \(F\). It is sometimes called the fundamental theorem of mathematical statistics, because it is what justifies learning about a population from a sample at all.