Skip to the content

Topics Covered

Four Modes Implications Counterexamples Slutzky's Theorem Borel–Cantelli Zero-One Law Glivenko–Cantelli
On this page
  1. 1. The Four Modes of Convergence
  2. 2. The Implications Between the Modes
  3. 3. Slutzky's Theorem
  4. 4. The Borel–Cantelli Lemmas
  5. 5. The Glivenko–Cantelli Lemma
  6. Key Take-aways
Where this unit starts. Theory of Probability, Unit 5 names the modes of convergence and states the laws of large numbers. It does not prove the implications between the modes, does not give the counterexamples that show the arrows run one way only, and does not reach Slutzky or Borel–Cantelli. Those are the content of this unit. The characteristic function and Lévy's continuity theorem from Unit 2 are used freely from here on.

1. The Four Modes of Convergence

DEFINITIONS

Let \(X_1, X_2, \ldots\) and \(X\) be random variables on a common probability space.

(a) Convergence in probability, written \(X_n \xrightarrow{P} X\): for every \(\varepsilon > 0\),

\[ \lim_{n \to \infty} P\big(|X_n - X| > \varepsilon\big) = 0. \]

(b) Almost sure convergence, written \(X_n \xrightarrow{a.s.} X\):

\[ P\left(\left\{\omega : \lim_{n \to \infty} X_n(\omega) = X(\omega)\right\}\right) = 1. \]

(c) Convergence in quadratic mean (the case \(r = 2\) of convergence in \(r\)-th mean), written \(X_n \xrightarrow{q.m.} X\):

\[ \lim_{n \to \infty} E\left(X_n - X\right)^{2} = 0. \]

(d) Convergence in law (in distribution), written \(X_n \xrightarrow{d} X\): with \(F_n\) and \(F\) the distribution functions,

\[ \lim_{n \to \infty} F_n(x) = F(x) \quad \text{at every } x \text{ where } F \text{ is continuous.} \]
THE DIFFERENCE BETWEEN (a) AND (b), IN WORDS

Both say "\(X_n\) gets close to \(X\)", and the difference is where the limit sits.

In probability, the limit is taken outside: for each fixed \(n\) we measure the set on which \(X_n\) is still far from \(X\), and ask that its probability shrink. Nothing forbids the bad set from moving around, so a given \(\omega\) may be inside it infinitely often.

Almost surely, the limit is taken inside: we look at one \(\omega\) at a time, ask whether the ordinary numerical sequence \(X_n(\omega)\) converges, and require the set of \(\omega\) for which it does to have probability 1. A given \(\omega\) may be in the bad set only finitely often.

That is exactly why almost sure convergence is the stronger of the two, and why the counterexample of Example 3.2 is possible.

2. The Implications Between the Modes

WHAT IMPLIES WHAT \[ \begin{aligned} X_n \xrightarrow{a.s.} X \;&\Longrightarrow\; X_n \xrightarrow{P} X \\ X_n \xrightarrow{q.m.} X \;&\Longrightarrow\; X_n \xrightarrow{P} X \\ X_n \xrightarrow{P} X \;&\Longrightarrow\; X_n \xrightarrow{d} X \\ X_n \xrightarrow{d} c \;&\Longrightarrow\; X_n \xrightarrow{P} c \quad \text{(only when the limit is a constant)} \end{aligned} \]

None of the first three reverses, and almost sure convergence and quadratic mean convergence do not imply one another in either direction.

The arrows run one way only almost sure P(lim Xn = X) = 1 quadratic mean E(Xn - X)² → 0 in probability P(|Xn - X| > ε) → 0 in distribution Fn(x) → F(x) only if the limit is a constant Ex. 3.2 blocks the reverse Ex. 3.1 blocks the reverse Ex. 3.3 blocks the reverse of the vertical arrow
Fig 3.1 — The four modes and the implications between them, with the example on this page that blocks each converse.
PROOF: QUADRATIC MEAN IMPLIES PROBABILITY

Statement. If \(E(X_n - X)^{2} \to 0\) then \(X_n \xrightarrow{P} X\).

Step 1 — fix the tolerance. Let \(\varepsilon > 0\) be arbitrary and held fixed throughout.

Step 2 — apply Markov's inequality. From Unit 2, \(P(|Z| \ge a) \le E|Z|^{r}/a^{r}\). Take \(Z = X_n - X\), \(a = \varepsilon\), \(r = 2\):

\[ P\big(|X_n - X| > \varepsilon\big) \;\le\; P\big(|X_n - X| \ge \varepsilon\big) \;\le\; \frac{E\left(X_n - X\right)^{2}}{\varepsilon^{2}}. \]

The first inequality is monotonicity of \(P\), because \(\{|X_n - X| > \varepsilon\} \subseteq \{|X_n - X| \ge \varepsilon\}\).

Step 3 — let \(n \to \infty\). The numerator tends to \(0\) by hypothesis and \(\varepsilon^{2}\) is a fixed positive constant, so the whole right-hand side tends to \(0\). A non-negative quantity squeezed below something tending to \(0\) tends to \(0\), so

\[ P\big(|X_n - X| > \varepsilon\big) \longrightarrow 0. \]

Since \(\varepsilon\) was arbitrary, this is convergence in probability. \(\blacksquare\)

EXAMPLE 3.1 — IN PROBABILITY BUT NOT IN QUADRATIC MEAN

Given. A sequence with

\[ P\left(X_n = n\right) = \frac{1}{n}, \qquad P\left(X_n = 0\right) = 1 - \frac{1}{n}. \]

Asked. Show \(X_n \xrightarrow{P} 0\) but that \(X_n\) converges neither in quadratic mean nor in mean to \(0\).

Step 1 — convergence in probability. Fix \(\varepsilon > 0\). For all \(n\) large enough that \(n > \varepsilon\), the only way \(|X_n - 0| > \varepsilon\) can happen is \(X_n = n\). Hence

\[ P\big(|X_n| > \varepsilon\big) = P(X_n = n) = \frac{1}{n} \longrightarrow 0. \]

Reading the values: at \(n = 10\) this is \(0.1\); at \(n = 100\), \(0.01\); at \(n = 10{,}000\), \(0.0001\). So \(X_n \xrightarrow{P} 0\).

Step 2 — the first moment.

\[ E(X_n) = n \cdot \frac{1}{n} + 0 \cdot \left(1 - \frac{1}{n}\right) = 1 \quad \text{for every } n. \]

So \(E|X_n - 0| = 1 \not\to 0\): there is no convergence in mean.

Step 3 — the second moment.

\[ E\left(X_n - 0\right)^{2} = n^{2} \cdot \frac{1}{n} + 0 = n \longrightarrow \infty. \]

Values: \(1, 10, 100\) at \(n = 1, 10, 100\). Far from tending to \(0\), it diverges, so there is no convergence in quadratic mean either.

Interpretation. The probability of the bad outcome vanishes, but the size of the bad outcome grows faster than its probability shrinks, and expectation multiplies the two. This is the same mechanism as the spike sequence of Unit 1, Example 1.4, now in probabilistic dress. It is why a consistent estimator need not be asymptotically unbiased, and why the two properties must be checked separately.

EXAMPLE 3.2 — IN PROBABILITY BUT NOT ALMOST SURELY

Given. Take \(\Omega = (0,1]\) with \(P\) = Lebesgue measure (the uniform distribution). Lay out the intervals in blocks: block \(k\) consists of the \(k\) intervals

\[ \left(\frac{j-1}{k}, \frac{j}{k}\right], \qquad j = 1, 2, \ldots, k, \]

and list all of them in one sequence \(A_1; A_2, A_3; A_4, A_5, A_6; \ldots\), so that block 1 contributes 1 interval, block 2 contributes 2, block 3 contributes 3, and so on. Put \(X_n = \mathbf{1}_{A_n}\).

Asked. Show \(X_n \xrightarrow{P} 0\) but that \(X_n(\omega)\) converges for no \(\omega\) at all.

Step 1 — convergence in probability. If \(A_n\) sits in block \(k\), its length is \(1/k\), so for \(0 < \varepsilon < 1\),

\[ P\big(|X_n - 0| > \varepsilon\big) = P(A_n) = \frac{1}{k}. \]

As \(n \to \infty\) the block index \(k\) also \(\to \infty\), because each block is finite, so this probability tends to \(0\). Hence \(X_n \xrightarrow{P} 0\).

Step 2 — no pointwise convergence. Fix any \(\omega \in (0,1]\). The \(k\) intervals of block \(k\) partition \((0,1]\), so exactly one of them contains \(\omega\). Therefore, in every block, \(X_n(\omega) = 1\) for exactly one \(n\) and \(X_n(\omega) = 0\) for the other \(k - 1\).

Step 3 — read off the consequence. Since there are infinitely many blocks, the numerical sequence \(X_n(\omega)\) contains infinitely many \(1\)s and (from block 2 onwards) infinitely many \(0\)s. A sequence taking both values infinitely often does not converge. This is true for every \(\omega\), so the set of \(\omega\) where \(X_n(\omega) \to 0\) is empty and has probability \(0\), not \(1\).

Interpretation. The bad set shrinks, but it sweeps across \(\Omega\) rather than settling down, so no single point is eventually spared. This is exactly the distinction drawn in words in section 1, made concrete. Note also that \(E(X_n - 0)^{2} = P(A_n) = 1/k \to 0\), so this sequence does converge in quadratic mean — which shows quadratic mean convergence does not imply almost sure convergence either.

EXAMPLE 3.3 — IN DISTRIBUTION BUT NOT IN PROBABILITY

Given. Let \(X \sim N(0,1)\) and define \(X_n = -X\) for every \(n\).

Asked. Show \(X_n \xrightarrow{d} X\) but \(X_n \not\xrightarrow{P} X\).

Step 1 — convergence in distribution. The standard normal is symmetric about \(0\), so \(-X\) has the same distribution as \(X\). Hence \(F_n = F\) for every \(n\), and a constant sequence trivially converges: \(F_n(x) \to F(x)\) at every \(x\). So \(X_n \xrightarrow{d} X\).

Step 2 — failure in probability. Here \(X_n - X = -X - X = -2X\), so for \(\varepsilon = 1\),

\[ P\big(|X_n - X| > 1\big) = P\big(|{-2X}| > 1\big) = P\left(|X| > \tfrac12\right) = 2\big[1 - \Phi(0.5)\big] = 2(1 - 0.691462) = 0.617075. \]

This does not depend on \(n\) at all, so it certainly does not tend to \(0\).

Interpretation. Convergence in distribution is a statement about the law of \(X_n\) and says nothing about how close \(X_n\) and \(X\) are as functions on \(\Omega\). Here they are as far apart as a symmetric variable can be from its own negative, while having identical distributions. The one case where the converse does hold is when the limit is a constant \(c\): there is then no room for the two to differ in sign or anywhere else, and \(X_n \xrightarrow{d} c\) does give \(X_n \xrightarrow{P} c\).

3. Slutzky's Theorem

STATEMENT

Suppose \(X_n \xrightarrow{d} X\) and \(Y_n \xrightarrow{P} c\), where \(c\) is a constant. Then

\[ X_n + Y_n \xrightarrow{d} X + c, \qquad X_n Y_n \xrightarrow{d} cX, \qquad \frac{X_n}{Y_n} \xrightarrow{d} \frac{X}{c} \ \ (c \ne 0). \]

The constancy of the limit of \(Y_n\) is essential; the theorem is false if \(Y_n\) converges in distribution to a non-degenerate variable.

EXAMPLE 3.4 — WHY t TESTS WORK IN LARGE SAMPLES

Given. \(X_1, \ldots, X_n\) i.i.d. with mean \(\mu\) and finite variance \(\sigma^{2}\); \(\bar{X}\) the sample mean and \(S\) the sample standard deviation. Two facts are available: the central limit theorem of Unit 4 gives

\[ Z_n = \frac{\bar{X} - \mu}{\sigma / \sqrt{n}} \xrightarrow{d} N(0,1), \]

and the weak law of large numbers gives \(S^{2} \xrightarrow{P} \sigma^{2}\), hence \(\sigma / S \xrightarrow{P} 1\) by the continuous mapping theorem.

Asked. Deduce the limiting distribution of \(T_n = (\bar{X} - \mu)/(S/\sqrt{n})\), and use it on a sample with \(n = 100\), \(\bar{x} = 52.4\), \(s = 8\).

Step 1 — write \(T_n\) as a product. Multiply and divide by \(\sigma\):

\[ T_n = \frac{\bar{X} - \mu}{S/\sqrt{n}} = \frac{\bar{X} - \mu}{\sigma/\sqrt{n}} \times \frac{\sigma}{S} = Z_n \cdot \frac{\sigma}{S}. \]

Step 2 — identify the two factors. \(Z_n \xrightarrow{d} N(0,1)\) and \(\sigma/S \xrightarrow{P} 1\), a constant. Slutzky's product form applies with \(c = 1\).

Step 3 — conclude.

\[ T_n \xrightarrow{d} 1 \cdot N(0,1) = N(0,1). \]

Step 4 — use it. An approximate \(95\%\) confidence interval for \(\mu\) is \(\bar{x} \pm 1.96 \, s/\sqrt{n}\). Here

\[ \frac{s}{\sqrt{n}} = \frac{8}{\sqrt{100}} = \frac{8}{10} = 0.8, \qquad 1.96 \times 0.8 = 1.568, \] \[ \text{interval} = 52.4 \pm 1.568 = (50.832,\ 53.968). \]

Interpretation. Replacing the unknown \(\sigma\) by the estimate \(S\) costs nothing asymptotically, and Slutzky is the theorem that says so. It is why the large-sample tests of Inferential Statistics, Unit 3 may use \(s\) in place of \(\sigma\) and still refer to the normal table.

4. The Borel–Cantelli Lemmas

THE EVENT "INFINITELY OFTEN"

For a sequence of events \(A_1, A_2, \ldots\), define

\[ \limsup_{n \to \infty} A_n \;=\; \bigcap_{m=1}^{\infty} \bigcup_{n=m}^{\infty} A_n \;=\; \{A_n \text{ occurs infinitely often}\}, \]

usually written \(\{A_n \text{ i.o.}\}\). Reading the formula: \(\omega\) belongs to it exactly when, for every starting index \(m\), some \(A_n\) with \(n \ge m\) contains \(\omega\) — that is, the \(A_n\) never stop catching \(\omega\).

FIRST LEMMA, WITH PROOF

Statement. If \(\sum_{n=1}^{\infty} P(A_n) < \infty\) then \(P(A_n \text{ i.o.}) = 0\). No independence is required.

Step 1 — bound by a tail union. For every \(m\),

\[ \{A_n \text{ i.o.}\} = \bigcap_{k=1}^{\infty} \bigcup_{n=k}^{\infty} A_n \;\subseteq\; \bigcup_{n=m}^{\infty} A_n, \]

since an intersection is contained in each of the sets being intersected. Monotonicity of \(P\) gives \(P(A_n \text{ i.o.}) \le P\left(\bigcup_{n \ge m} A_n\right)\).

Step 2 — apply countable subadditivity. Property (P2) of Unit 1 gives

\[ P\left(\bigcup_{n=m}^{\infty} A_n\right) \le \sum_{n=m}^{\infty} P(A_n). \]

Step 3 — let the tail vanish. A convergent series has tails tending to zero, so \(\sum_{n \ge m} P(A_n) \to 0\) as \(m \to \infty\). Combining Steps 1 and 2, the fixed number \(P(A_n \text{ i.o.})\) is below a quantity that can be made arbitrarily small, so it is \(0\). \(\blacksquare\)

SECOND LEMMA

Statement. If the \(A_n\) are independent and \(\sum_{n=1}^{\infty} P(A_n) = \infty\), then \(P(A_n \text{ i.o.}) = 1\).

Independence is indispensable here. Without it the conclusion fails: take a single event \(A\) with \(P(A) = \tfrac12\) and set \(A_n = A\) for every \(n\). Then \(\sum_n P(A_n) = \infty\), but \(\{A_n \text{ i.o.}\} = A\), whose probability is \(\tfrac12\), not \(1\).

BOREL'S ZERO-ONE LAW

Putting the two lemmas together: for independent events, \(P(A_n \text{ i.o.})\) is either \(0\) or \(1\), and which one is decided entirely by whether \(\sum_n P(A_n)\) converges or diverges. There is no middle ground and no third case to check.

EXAMPLE 3.5 — THE TWO SERIES, SIDE BY SIDE

Given. Independent events with (a) \(P(A_n) = 1/n^{2}\) and (b) \(P(B_n) = 1/n\).

Asked. Find \(P(A_n \text{ i.o.})\) and \(P(B_n \text{ i.o.})\).

Step 1 — case (a), test the series. \(\sum_{n \ge 1} 1/n^{2}\) is the Basel series and converges to \(\pi^{2}/6 = 1.644934\). Partial sums confirm the convergence: summing to \(n = 4000\) gives \(1.644684\), already within \(0.00025\) of the limit.

Step 2 — apply the first lemma. The series converges, so \(P(A_n \text{ i.o.}) = 0\). Only finitely many \(A_n\) occur, with probability 1.

Step 3 — put a number on "finitely many". The proof's own bound is usable. The probability that any \(A_n\) with \(n \ge m\) occurs at all is at most \(\sum_{n \ge m} 1/n^{2}\), and that tail is

\(m\)\(\sum_{n \ge m} 1/n^{2}\)reading
100.10517about a 1-in-10 chance of any event from \(A_{10}\) onwards
1000.01005about 1 in 100
10000.00100about 1 in 1000

Step 4 — case (b), test the series. \(\sum_{n \ge 1} 1/n\) is the harmonic series and diverges. Its growth is slow but unbounded: the partial sum to \(n = 4000\) is \(8.871390\), and to \(n = 40{,}000\) it is \(11.173863\) — a tenfold increase in \(n\) adds only about \(2.30 \approx \ln 10\), which is the signature of logarithmic divergence.

Step 5 — apply the second lemma. The events are independent and the series diverges, so \(P(B_n \text{ i.o.}) = 1\). Infinitely many \(B_n\) occur, with certainty.

Interpretation. The two cases sit either side of the borderline \(\sum 1/n^{p}\), which converges for \(p > 1\) and diverges for \(p \le 1\). Note how little it takes to cross: \(P(A_n) = 1/n^{2}\) gives "finitely many, almost surely", while \(P(B_n) = 1/n\) — larger, but still tending to zero — gives "infinitely many, almost surely". Individual probabilities tending to zero say nothing on their own; only the series does.

5. The Glivenko–Cantelli Lemma

STATEMENT

Let \(X_1, \ldots, X_n\) be i.i.d. with distribution function \(F\), and let

\[ F_n(x) = \frac{1}{n} \sum_{i=1}^{n} \mathbf{1}_{\{X_i \le x\}} \]

be the empirical distribution function — the proportion of the sample at or below \(x\). Then

\[ \sup_{x \in \mathbb{R}} \left| F_n(x) - F(x) \right| \xrightarrow{a.s.} 0. \]

Why it matters. For each fixed \(x\), \(F_n(x) \xrightarrow{a.s.} F(x)\) is just the strong law of large numbers applied to the indicators. Glivenko–Cantelli says far more: the convergence is uniform in \(x\), so the worst discrepancy anywhere on the line goes to zero. That is the theorem behind the Kolmogorov–Smirnov goodness-of-fit test, whose statistic is exactly this supremum, and behind the bootstrap of Estimation Theory (STS-201), Unit 2, which resamples from \(F_n\) on the grounds that it is uniformly close to \(F\). It is sometimes called the fundamental theorem of mathematical statistics, because it is what justifies learning about a population from a sample at all.

Key Take-aways