At Foundation level a probability is a number attached to an event, and an event is "any subset of the sample space". That works perfectly while \(\Omega\) is finite or countable. It breaks the moment \(\Omega\) is uncountable.
Take \(\Omega = (0,1]\) and ask for a \(P\) that (i) assigns to every interval its length, (ii) is countably additive, and (iii) is defined on every subset of \((0,1]\). No such \(P\) exists. Sets can be built — not by accident, but deliberately — for which no consistent length can be assigned at all.
The repair is not to weaken additivity. It is to shrink the domain: keep countable additivity, and only promise \(P\) on a collection of subsets that is rich enough to hold every event we actually need, and closed under the operations we actually perform. That collection is a sigma-field.
A non-empty class \(\mathcal{F}\) of subsets of \(\Omega\) is a field if
Closure under finite intersection follows and need not be assumed: by De Morgan's law \(A \cap B = (A^{c} \cup B^{c})^{c}\), and each operation on the right keeps us inside \(\mathcal{F}\) by rules 2 and 3.
A class \(\mathcal{A}\) of subsets of \(\Omega\) is a sigma-field (\(\sigma\)-field, or \(\sigma\)-algebra) if it satisfies 1 and 2 above and, in place of 3, the stronger
3'. \(A_1, A_2, A_3, \ldots \in \mathcal{A} \Rightarrow \bigcup_{n=1}^{\infty} A_n \in \mathcal{A}\) (closed under countable union).
Every sigma-field is a field: take \(A_3 = A_4 = \cdots = \varnothing\) in 3' and a countable union collapses to a finite one. The converse is false, and the next example shows why that matters.
Given. Let \(\Omega = \{1, 2, 3, \ldots\}\) and let
\[ \mathcal{F} = \{A \subseteq \Omega : A \text{ is finite, or } A^{c} \text{ is finite}\}. \]Asked. Show that \(\mathcal{F}\) is a field but not a sigma-field.
Step 1 — \(\Omega \in \mathcal{F}\). \(\Omega^{c} = \varnothing\), which is finite. So \(\Omega\) qualifies under the second half of the definition.
Step 2 — closed under complementation. The defining condition is symmetric in \(A\) and \(A^{c}\): it says "one of the two is finite". Swapping \(A\) for \(A^{c}\) leaves that statement unchanged, so \(A \in \mathcal{F} \Rightarrow A^{c} \in \mathcal{F}\).
Step 3 — closed under finite union. Take \(A, B \in \mathcal{F}\) and consider the two cases.
Both cases land inside \(\mathcal{F}\), so rule 3 holds and \(\mathcal{F}\) is a field.
Step 4 — it fails rule 3'. Put \(A_n = \{2n\}\) for \(n = 1, 2, 3, \ldots\). Each \(A_n\) is a singleton, hence finite, hence in \(\mathcal{F}\). Their countable union is
\[ \bigcup_{n=1}^{\infty} A_n = \{2, 4, 6, 8, \ldots\}, \]the set of even positive integers. This set is infinite, and its complement — the odd positive integers — is also infinite. So neither half of the defining condition holds, and the union is not in \(\mathcal{F}\).
Interpretation. A field can carry every event you can write down in finitely many steps. It cannot carry limits, and every limit theorem in this course — convergence, the laws of large numbers, the central limit theorems — is a statement about a countable sequence of events. That is exactly why probability is built on sigma-fields and not on fields.
Given. \(\Omega = \{a, b, c\}\).
Asked. Write down the smallest and the largest sigma-field on \(\Omega\), and one that is neither.
Smallest. \(\mathcal{A}_0 = \{\varnothing, \Omega\}\). Check: \(\Omega\) is in it; the complement of each member is the other member; every union of members is again a member. It is called the trivial sigma-field, and it is contained in every other sigma-field on \(\Omega\), because rules 1 and 2 force both of its elements to be present in any sigma-field at all.
Largest. The power set \(\mathcal{P}(\Omega)\), all \(2^{3} = 8\) subsets. It is closed under everything, being everything.
In between. \(\mathcal{A}_1 = \{\varnothing, \{a\}, \{b, c\}, \Omega\}\), with \(2^{2} = 4\) members. Complements: \(\varnothing^{c} = \Omega\) and \(\{a\}^{c} = \{b,c\}\), both present. Unions: \(\{a\} \cup \{b,c\} = \Omega\), present; every other union repeats a member. So \(\mathcal{A}_1\) is a sigma-field.
Interpretation. \(\mathcal{A}_1\) is the information carried by the single question "did \(a\) occur?". A sigma-field is not just a technical domain for \(P\); it is a precise record of how much can be observed. That reading is what makes conditional expectation in Unit 2 mean something.
Let \(\mathcal{C}\) be any class of subsets of \(\Omega\). Among all sigma-fields containing \(\mathcal{C}\) there is a smallest one, written \(\sigma(\mathcal{C})\) and called the sigma-field generated by \(\mathcal{C}\), or the minimal sigma-field over \(\mathcal{C}\).
Statement. For any class \(\mathcal{C}\) of subsets of \(\Omega\), there is a sigma-field \(\sigma(\mathcal{C}) \supseteq \mathcal{C}\) contained in every sigma-field that contains \(\mathcal{C}\).
Step 1 — the collection we intersect is not empty. The power set \(\mathcal{P}(\Omega)\) is a sigma-field and contains \(\mathcal{C}\). So at least one sigma-field containing \(\mathcal{C}\) exists, and the intersection in Step 2 is over a non-empty collection.
Step 2 — define the candidate. Put
\[ \sigma(\mathcal{C}) \;=\; \bigcap \{\mathcal{A} : \mathcal{A} \text{ is a sigma-field on } \Omega \text{ and } \mathcal{C} \subseteq \mathcal{A}\}. \]Step 3 — an intersection of sigma-fields is a sigma-field. Verify the three rules one at a time. \(\Omega\) lies in every \(\mathcal{A}\) in the collection, so it lies in the intersection. If \(A\) lies in the intersection then \(A \in \mathcal{A}\) for every \(\mathcal{A}\), so \(A^{c} \in \mathcal{A}\) for every \(\mathcal{A}\) by rule 2 applied inside each one, so \(A^{c}\) lies in the intersection. If \(A_1, A_2, \ldots\) lie in the intersection, the same argument with rule 3' puts \(\bigcup_n A_n\) in every \(\mathcal{A}\), hence in the intersection.
Step 4 — it contains \(\mathcal{C}\). Every \(\mathcal{A}\) in the collection contains \(\mathcal{C}\) by the way the collection was chosen, so the intersection does too.
Step 5 — it is the smallest. Let \(\mathcal{B}\) be any sigma-field with \(\mathcal{C} \subseteq \mathcal{B}\). Then \(\mathcal{B}\) is one of the sets being intersected in Step 2, and an intersection is contained in each of its members. So \(\sigma(\mathcal{C}) \subseteq \mathcal{B}\). \(\blacksquare\)
The Borel sigma-field \(\mathcal{B}\) on \(\mathbb{R}\) is the sigma-field generated by the open intervals:
\[ \mathcal{B} \;=\; \sigma\big(\{(a,b) : -\infty < a < b < \infty\}\big). \]Its members are called Borel sets. The same \(\mathcal{B}\) is obtained from the half-lines \((-\infty, x]\), and that generating class is the convenient one for probability, because it is exactly the class of sets appearing in a distribution function.
Given. \(\mathcal{B}\) generated by the open intervals.
Asked. Show \(\{x\} \in \mathcal{B}\) for every real \(x\), and deduce that every countable subset of \(\mathbb{R}\) is a Borel set.
Step 1 — write the singleton as a countable intersection of open intervals.
\[ \{x\} \;=\; \bigcap_{n=1}^{\infty} \left(x - \tfrac{1}{n},\; x + \tfrac{1}{n}\right). \]Justification of the equality: \(x\) belongs to every one of these intervals, so it is in the intersection. Conversely if \(y \ne x\) then \(|y - x| > 0\), and choosing \(n\) with \(1/n < |y - x|\) — possible because \(1/n \to 0\) — puts \(y\) outside that interval, hence outside the intersection.
Step 2 — a sigma-field is closed under countable intersection. Rule 3' gives countable unions; De Morgan converts:
\[ \bigcap_{n=1}^{\infty} A_n \;=\; \left( \bigcup_{n=1}^{\infty} A_n^{c} \right)^{c}, \]and the right-hand side uses only complementation (rule 2) and countable union (rule 3').
Step 3 — conclude. Each \((x - 1/n, x + 1/n)\) is an open interval, hence in \(\mathcal{B}\) by definition. By Step 2 the intersection is in \(\mathcal{B}\), so \(\{x\} \in \mathcal{B}\).
Step 4 — countable sets. If \(C = \{x_1, x_2, \ldots\}\) is countable then \(C = \bigcup_{k} \{x_k\}\), a countable union of Borel sets, so \(C \in \mathcal{B}\) by rule 3'.
Interpretation. This is the technical fact behind a statement made without proof at Foundation level: for a continuous random variable \(P(X = x) = 0\), yet \(P(a < X < b)\) can be positive. The event \(\{X = x\}\) is a perfectly legitimate Borel set — it simply carries measure zero. The set \(\mathbb{Q}\) of rationals, being countable, likewise has Lebesgue measure zero while being dense in \(\mathbb{R}\).
Let \(\mathcal{A}\) be a sigma-field on \(\Omega\). A set function \(\mu : \mathcal{A} \to [0, \infty]\) is a measure if
The triple \((\Omega, \mathcal{A}, \mu)\) is a measure space.
Statement. If \(A, B \in \mathcal{A}\) and \(A \subseteq B\) then \(\mu(A) \le \mu(B)\).
Step 1 — split \(B\) into disjoint pieces. Because \(A \subseteq B\),
\[ B = A \cup (B \cap A^{c}), \qquad A \cap (B \cap A^{c}) = \varnothing. \]Both pieces are in \(\mathcal{A}\): \(A\) is given, and \(B \cap A^{c}\) is built from \(B\) and \(A\) by complementation and intersection, which a sigma-field permits.
Step 2 — apply countable additivity. Pad the two-set union with empty sets, \(A_1 = A\), \(A_2 = B \cap A^{c}\), \(A_3 = A_4 = \cdots = \varnothing\). These are pairwise disjoint, so rule 3 of the definition gives
\[ \mu(B) = \mu(A) + \mu(B \cap A^{c}) + 0 + 0 + \cdots = \mu(A) + \mu(B \cap A^{c}). \]Step 3 — drop a non-negative term. By rule 1, \(\mu(B \cap A^{c}) \ge 0\). Removing it from the right-hand side can only decrease it, so \(\mu(B) \ge \mu(A)\). \(\blacksquare\)
Statement. If \(A_1 \subseteq A_2 \subseteq A_3 \subseteq \cdots\) and \(A = \bigcup_{n=1}^{\infty} A_n\), then \(\mu(A_n) \to \mu(A)\) as \(n \to \infty\).
Step 1 — disjointify the increasing sequence. Define
\[ B_1 = A_1, \qquad B_n = A_n \cap A_{n-1}^{c} \quad (n \ge 2). \]Each \(B_n \in \mathcal{A}\). They are pairwise disjoint: for \(m < n\) we have \(B_m \subseteq A_m \subseteq A_{n-1}\) by the nesting, while \(B_n \subseteq A_{n-1}^{c}\), and a set and its complement share no point.
Step 2 — the partial unions agree. By induction, \(\bigcup_{k=1}^{n} B_k = A_n\). It holds at \(n=1\). If it holds at \(n-1\) then
\[ \bigcup_{k=1}^{n} B_k = A_{n-1} \cup (A_n \cap A_{n-1}^{c}) = A_n, \]the last equality because \(A_{n-1} \subseteq A_n\). Letting \(n \to \infty\) gives \(\bigcup_{k=1}^{\infty} B_k = A\) as well.
Step 3 — apply countable additivity to the disjoint \(B_k\).
\[ \mu(A) = \mu\left(\bigcup_{k=1}^{\infty} B_k\right) = \sum_{k=1}^{\infty} \mu(B_k). \]Step 4 — recognise the partial sum. An infinite series is by definition the limit of its partial sums, and by Step 2 with finite additivity,
\[ \sum_{k=1}^{n} \mu(B_k) = \mu\left(\bigcup_{k=1}^{n} B_k\right) = \mu(A_n). \]Therefore \(\mu(A) = \lim_{n \to \infty} \mu(A_n)\). \(\blacksquare\)
Let \((\Omega, \mathcal{A})\) be a measurable space. A function \(f : \Omega \to \mathbb{R}\) is measurable (with respect to \(\mathcal{A}\)) if the inverse image of every Borel set is in \(\mathcal{A}\):
\[ f^{-1}(B) = \{\omega \in \Omega : f(\omega) \in B\} \in \mathcal{A} \quad \text{for every } B \in \mathcal{B}. \]It is enough to check this on a generating class. In particular \(f\) is measurable if and only if \(\{\omega : f(\omega) \le x\} \in \mathcal{A}\) for every real \(x\), which is the form used in the next section.
If \(f\) and \(g\) are measurable and \(c\) is a constant, then so are
\[ cf, \quad f + g, \quad fg, \quad |f|, \quad \max(f,g), \quad \min(f,g), \]and if \(f_1, f_2, \ldots\) are measurable then so are
\[ \sup_n f_n, \quad \inf_n f_n, \quad \limsup_n f_n, \quad \liminf_n f_n, \]with \(\lim_n f_n\) measurable whenever it exists. Closure under limits is the property the convergence theorems of Unit 3 rest on: the limit of a sequence of random variables is again a random variable, and needs no separate justification each time.
Let \(\mathcal{F}_0\) be a field of subsets of \(\Omega\) and let \(\mu_0 : \mathcal{F}_0 \to [0, \infty]\) be countably additive on \(\mathcal{F}_0\) with \(\mu_0(\varnothing) = 0\). Then \(\mu_0\) extends to a measure \(\mu\) on \(\sigma(\mathcal{F}_0)\), agreeing with \(\mu_0\) on \(\mathcal{F}_0\). If \(\mu_0\) is \(\sigma\)-finite — that is, \(\Omega = \bigcup_n E_n\) with \(E_n \in \mathcal{F}_0\) and \(\mu_0(E_n) < \infty\) — the extension is unique.
The proof is long and is not reproduced here; the syllabus asks for the statement and its applications, which is what follows.
(a) Lebesgue measure exists. On \(\Omega = (0,1]\) take \(\mathcal{F}_0\) to be the finite disjoint unions of half-open intervals \((a,b]\), and define \(\mu_0\big(\bigcup_{i=1}^{k}(a_i,b_i]\big) = \sum_{i=1}^{k}(b_i - a_i)\). This \(\mathcal{F}_0\) is a field and \(\mu_0\) is countably additive on it — both checkable directly. Carathéodory then produces a unique measure on \(\sigma(\mathcal{F}_0) = \mathcal{B} \cap (0,1]\) assigning each interval its length. That measure is Lebesgue measure, and the uniform distribution on \((0,1]\) is it.
(b) A distribution function determines a probability measure. Given any \(F\) with the four properties of section 8, define \(\mu_0((a,b]) = F(b) - F(a)\) on the same field. Carathéodory extends it uniquely to \(\mathcal{B}\). This is the theorem that lets a statistician specify a model by writing down one function of one real variable, and be certain a consistent probability measure stands behind it.
Let \((\Omega, \mathcal{A}, \mu)\) be a measure space and \(f_n\) measurable.
Monotone Convergence Theorem (MCT). If \(0 \le f_1 \le f_2 \le \cdots\) and \(f_n \to f\) pointwise, then
\[ \int f_n \, d\mu \;\longrightarrow\; \int f \, d\mu. \]Fatou's Lemma. If \(f_n \ge 0\), then
\[ \int \liminf_{n \to \infty} f_n \, d\mu \;\le\; \liminf_{n \to \infty} \int f_n \, d\mu. \]Dominated Convergence Theorem (DCT). If \(f_n \to f\) pointwise and there is a single integrable \(g\) with \(|f_n| \le g\) for every \(n\), then \(f\) is integrable and
\[ \int f_n \, d\mu \;\longrightarrow\; \int f \, d\mu. \]All three answer one question: when may a limit be taken inside an integral? MCT allows it when the sequence only rises; DCT allows it when the sequence is trapped under a fixed integrable ceiling; Fatou gives an inequality with no extra hypothesis at all.
Given. Lebesgue measure on \((0,1)\) and the functions
\[ f_n(x) = n \cdot \mathbf{1}_{(0,\,1/n)}(x), \qquad n = 1, 2, 3, \ldots \](a spike of height \(n\) on a base of width \(1/n\), and zero elsewhere).
Asked. Compute \(\int f_n\), find \(\lim f_n\), and say which of the three theorems applies.
Step 1 — the integral of each \(f_n\). The function is a constant \(n\) on a set of Lebesgue measure \(1/n\) and zero off it, so
\[ \int_0^1 f_n(x)\,dx = n \times \frac{1}{n} = 1 \qquad \text{for every } n. \]Check at three values: \(n=1\) gives \(1 \times 1 = 1\); \(n=10\) gives \(10 \times 0.1 = 1\); \(n=1000\) gives \(1000 \times 0.001 = 1\). The value never changes.
Step 2 — the pointwise limit. Fix any \(x\) with \(0 < x < 1\). Choose \(N\) with \(1/N < x\); this is possible because \(1/n \to 0\). For every \(n \ge N\) the point \(x\) lies outside \((0, 1/n)\), so \(f_n(x) = 0\). Hence \(f_n(x) \to 0\), and this holds at every \(x\), so \(f = 0\) and \(\int f = 0\).
Step 3 — compare.
\[ \lim_{n \to \infty} \int f_n \,d\mu = 1 \qquad \text{but} \qquad \int \lim_{n \to \infty} f_n \,d\mu = 0. \]Step 4 — which hypothesis failed? MCT does not apply: the sequence is not increasing (at \(x = 0.3\), \(f_1(x) = 1\) but \(f_5(x) = 0\)). DCT does not apply either: any \(g\) dominating all the \(f_n\) must satisfy \(g(x) \ge n\) for every \(n\) with \(x < 1/n\), which forces \(g(x) \ge 1/x\) near \(0\), and \(\int_0^1 (1/x)\,dx = \infty\), so no integrable dominating function exists.
Step 5 — Fatou still holds. The functions are non-negative, so Fatou applies with no further condition and asserts \(\int \liminf f_n \le \liminf \int f_n\), that is \(0 \le 1\). True, and strictly so.
Interpretation. Probability mass can "escape to infinity" in height while the area stays fixed. This is not a curiosity: it is the same mechanism by which a sequence of estimators can converge to a constant while their expectations do not converge to that constant — the distinction between convergence in probability and convergence in mean, taken up in Unit 3.
A probability measure is a measure \(P\) on \((\Omega, \mathcal{A})\) with the single extra requirement \(P(\Omega) = 1\). The triple \((\Omega, \mathcal{A}, P)\) is a probability space:
Because \(P(\Omega) = 1 < \infty\), the finiteness proviso in (P4) is automatic, so probability measures are continuous from above as well as from below without further conditions.
These are Kolmogorov's axioms. The Foundation course states them for finitely many events; the only change here is that the third runs over a countable sequence, and that the domain is stated explicitly as a sigma-field rather than left as "all subsets".
A random variable on \((\Omega, \mathcal{A}, P)\) is a measurable function \(X : \Omega \to \mathbb{R}\). Its distribution function is
\[ F(x) = P(X \le x) = P\big(\{\omega : X(\omega) \le x\}\big), \qquad x \in \mathbb{R}, \]which is well defined precisely because measurability puts \(\{X \le x\}\) in \(\mathcal{A}\), the domain of \(P\).
Statement. Every distribution function satisfies
(i) \(F\) is non-decreasing; (ii) \(F(-\infty) = 0\) and \(F(+\infty) = 1\); (iii) \(F\) is right-continuous.
Proof of (i). Let \(x < y\). Then \(\{X \le x\} \subseteq \{X \le y\}\), because any \(\omega\) with \(X(\omega) \le x\) also has \(X(\omega) \le y\). Monotonicity (P1) of the measure \(P\) gives \(P(X \le x) \le P(X \le y)\), i.e. \(F(x) \le F(y)\).
Proof of (ii). Take any \(x_n \uparrow \infty\). The events \(A_n = \{X \le x_n\}\) increase and \(\bigcup_n A_n = \Omega\), since every real value \(X(\omega)\) is eventually below some \(x_n\). Continuity from below (P3) gives \(F(x_n) = P(A_n) \to P(\Omega) = 1\). Similarly with \(x_n \downarrow -\infty\) the events \(\{X \le x_n\}\) decrease to \(\varnothing\), and continuity from above (P4), legitimate because \(P(\Omega) = 1 < \infty\), gives \(F(x_n) \to P(\varnothing) = 0\).
Proof of (iii). Let \(h_n \downarrow 0\). The events \(B_n = \{X \le x + h_n\}\) decrease, and their intersection is \(\{X \le x\}\): any \(\omega\) in every \(B_n\) has \(X(\omega) \le x + h_n\) for all \(n\), hence \(X(\omega) \le x\) on letting \(n \to \infty\). Continuity from above then gives \(F(x + h_n) \to F(x)\), which is right-continuity. \(\blacksquare\)
Note on left limits. The same argument with \(B_n = \{X \le x - h_n\}\) increasing gives \(F(x-) = P(X < x)\), so the jump of \(F\) at \(x\) is
\[ F(x) - F(x-) = P(X \le x) - P(X < x) = P(X = x). \]A distribution function is right-continuous but need not be left-continuous, and each discontinuity is an atom whose height is exactly its probability.
Given. A random variable \(X\) with
\[ F(x) = \begin{cases} 0, & x < 0, \\ \tfrac{1}{4} + \tfrac{1}{2}x, & 0 \le x < 1, \\ 1, & x \ge 1. \end{cases} \]Asked. Verify that \(F\) is a distribution function, and find \(P(X = 0)\), \(P(X = 1)\), \(P(0 < X < 1)\) and \(P(X \le \tfrac12)\).
Step 1 — non-decreasing. On \([0,1)\) the slope is \(\tfrac12 > 0\), so \(F\) rises there; elsewhere it is constant. At \(x = 0\) it jumps up from \(0\) to \(\tfrac14\), and at \(x = 1\) it jumps up from \(\tfrac14 + \tfrac12 = \tfrac34\) to \(1\). Every movement is upward, so (i) holds.
Step 2 — limits. \(F(x) = 0\) for all \(x < 0\), so \(F(-\infty) = 0\); \(F(x) = 1\) for all \(x \ge 1\), so \(F(+\infty) = 1\). (ii) holds.
Step 3 — right-continuity. At \(x = 0\): \(F(0) = \tfrac14 + \tfrac12(0) = \tfrac14\), and \(F(0+) = \lim_{h \downarrow 0} (\tfrac14 + \tfrac12 h) = \tfrac14\). They agree. At \(x = 1\): \(F(1) = 1\) and \(F(1+) = 1\). They agree. (iii) holds, so \(F\) is a genuine distribution function.
Step 4 — the atom at 0. Using the jump formula,
\[ P(X = 0) = F(0) - F(0-) = \tfrac14 - 0 = \tfrac14 = 0.25. \]Step 5 — the atom at 1.
\[ P(X = 1) = F(1) - F(1-) = 1 - \left(\tfrac14 + \tfrac12 \times 1\right) = 1 - \tfrac34 = \tfrac14 = 0.25. \]Step 6 — the continuous part.
\[ P(0 < X < 1) = F(1-) - F(0) = \tfrac34 - \tfrac14 = \tfrac12 = 0.50. \]Step 7 — check the three pieces exhaust the probability.
\[ 0.25 + 0.50 + 0.25 = 1.00. \checkmark \]Step 8 — the last requested value.
\[ P\left(X \le \tfrac12\right) = F\left(\tfrac12\right) = \tfrac14 + \tfrac12 \times \tfrac12 = \tfrac14 + \tfrac14 = \tfrac12 = 0.50. \]Interpretation. \(X\) is neither discrete nor continuous. It places a quarter of its mass on the single point \(0\), a quarter on the single point \(1\), and spreads the remaining half uniformly across \((0,1)\). At Foundation level such a variable is awkward to describe at all; in the measure-theoretic framework it is simply a probability measure on \(\mathcal{B}\), no more exceptional than a normal one. That generality is the whole return on the machinery of this unit.