Skip to the content

Topics Covered

Metric Spaces Compact Sets Heine–Borel Perfect Sets Connected Sets Continuity Uniform Continuity Monotonic Functions Mean Value Theorems
On this page
  1. 1. Metric Spaces
  2. 2. Compact Sets
  3. 3. Perfect and Connected Sets
  4. 4. Limits and Continuity
  5. 5. Continuity, Compactness and Connectedness
  6. 6. Discontinuities and Monotonic Functions
  7. 7. Differentiation
  8. Key Take-aways
Where this unit starts. Sequences, series, limits, continuity, differentiability and Riemann integration on the real line are covered at exam level in UGC NET Statistics, Unit 2. This unit does not repeat them. It replaces the interval by a metric space, which is what makes the same theorems apply to spaces of functions, to random variables and to \(\mathbb{R}^{n}\) at once — and that generality is what the courses that follow need.

1. Metric Spaces

DEFINITION

A metric space is a set \(X\) together with a function \(d : X \times X \to \mathbb{R}\) such that for all \(x, y, z \in X\),

  1. \(d(x,y) \ge 0\), with \(d(x,y) = 0\) if and only if \(x = y\);
  2. \(d(x,y) = d(y,x)\) (symmetry);
  3. \(d(x,z) \le d(x,y) + d(y,z)\) (the triangle inequality).

The only structure assumed is a notion of distance. Every idea below — open, closed, compact, connected, continuous, convergent — is built from \(d\) alone.

spacemetricwhere it is used
\(\mathbb{R}\)\(|x - y|\)ordinary calculus
\(\mathbb{R}^{n}\)\(\sqrt{\sum_i (x_i - y_i)^{2}}\)multivariate analysis; least squares
\(C[a,b]\)\(\sup_{t}|f(t) - g(t)|\)uniform convergence, Unit 4
random variables\(\left[E|X - Y|^{2}\right]^{1/2}\)convergence in quadratic mean

The last row is the one that matters here: the four modes of convergence in Probability Theory, Unit 3 are convergence in four different metrics, and the reason they do not agree is that the metrics are genuinely different.

THE BASIC VOCABULARY

Let \(E \subseteq X\). The open ball is \(B(x, r) = \{y : d(x,y) < r\}\), and then

\[ \begin{aligned} &x \text{ is an \textbf{interior point} of } E \text{ if } B(x,r) \subseteq E \text{ for some } r > 0 \\ &E \text{ is \textbf{open} if every point of } E \text{ is interior} \\ &x \text{ is a \textbf{limit point} of } E \text{ if every } B(x,r) \text{ meets } E \text{ in a point other than } x \\ &E \text{ is \textbf{closed} if it contains all its limit points} \\ &E \text{ is \textbf{bounded} if } E \subseteq B(x, M) \text{ for some } x \text{ and finite } M \end{aligned} \]

A set may be both open and closed, or neither: in \(\mathbb{R}\), \(\varnothing\) and \(\mathbb{R}\) are both, while \([0,1)\) is neither. "Not open" does not mean "closed".

2. Compact Sets

DEFINITION AND THE THREE CHARACTERISATIONS

An open cover of \(E\) is a collection of open sets whose union contains \(E\). \(E\) is compact if every open cover of \(E\) has a finite subcover.

The definition looks abstract and is the right one, because in a general metric space it is equivalent to two more familiar conditions:

\[ E \text{ compact} \iff E \text{ sequentially compact} \iff E \text{ complete and totally bounded}, \]

where sequentially compact means every sequence in \(E\) has a subsequence converging to a point of \(E\).

Heine–Borel, in \(\mathbb{R}^{n}\) only: \(E \subseteq \mathbb{R}^{n}\) is compact if and only if it is closed and bounded. This is a theorem about \(\mathbb{R}^{n}\), not a definition, and it fails in general metric spaces — the next example shows how.

EXAMPLE 1.1 — CLOSED AND BOUNDED BUT NOT COMPACT

Given. Let \(X\) be the space of square-summable sequences with the metric \(d(x,y) = \left(\sum_i (x_i - y_i)^{2}\right)^{1/2}\), and let \(E\) be its closed unit ball, \(\{x : d(x, 0) \le 1\}\). Let \(e_n\) be the sequence with 1 in position \(n\) and 0 elsewhere.

Step 1 — \(E\) is bounded. By construction every point is within distance 1 of the origin.

Step 2 — \(E\) is closed. It is the inverse image of the closed set \([0,1]\) under the continuous map \(x \mapsto d(x,0)\).

Step 3 — compute the distance between two of the \(e_n\). For \(m \ne n\), the difference \(e_m - e_n\) has a \(+1\) and a \(-1\) and zeros elsewhere, so

\[ d(e_m, e_n) = \sqrt{1^{2} + (-1)^{2}} = \sqrt{2} \approx 1.414214. \]

Step 4 — conclude. Every pair of terms of the sequence \((e_n)\) is at distance \(\sqrt2\), so no subsequence can be Cauchy and none can converge. \(E\) is therefore not sequentially compact, hence not compact, although it is closed and bounded.

Interpretation. Heine–Borel is a statement about finite dimension. In infinite dimensions — which is where every space of functions lives — "closed and bounded" is much weaker than "compact". That gap is the reason the Stone–Weierstrass theorem of Unit 4 is a substantial result rather than an obvious one.

WHAT COMPACTNESS BUYS \[ \begin{aligned} &\text{a closed subset of a compact set is compact} \\ &\text{a compact set is closed and bounded (in any metric space)} \\ &\text{nested non-empty compact sets have non-empty intersection} \\ &\text{an infinite subset of a compact set has a limit point in it} \end{aligned} \]

The last is the Bolzano–Weierstrass property, and it is what makes every existence proof in optimisation work.

3. Perfect and Connected Sets

PERFECT SETS

\(E\) is perfect if it is closed and every point of \(E\) is a limit point of \(E\) — that is, \(E\) has no isolated points.

Theorem. A non-empty perfect subset of \(\mathbb{R}^{n}\) is uncountable.

The standard example, and why it is standard. The Cantor set is built by removing the open middle third of \([0,1]\), then the middle thirds of what remains, and so on. The total length removed is

\[ \frac{1}{3} + \frac{2}{9} + \frac{4}{27} + \cdots = \frac{1}{3}\sum_{k=0}^{\infty}\left(\frac{2}{3}\right)^{k} = \frac{1}{3} \cdot \frac{1}{1 - 2/3} = \frac{1}{3} \times 3 = 1, \]

so what is left has measure zero. Yet it is perfect, hence uncountable. A set can therefore be as numerous as \(\mathbb{R}\) and still be negligible for integration. That is precisely the distinction between cardinality and measure which Probability Theory, Unit 1 depends on, and it is why "\(P(X = x) = 0\) for every \(x\)" is consistent with "\(X\) takes some value".

CONNECTED SETS

\(E\) is separated if it can be written as \(A \cup B\) with \(A, B\) non-empty and neither containing a limit point of the other. \(E\) is connected if it is not separated — it cannot be split into two pieces that stay clear of each other.

Theorem (the connected subsets of \(\mathbb{R}\)). \(E \subseteq \mathbb{R}\) is connected if and only if it is an interval: whenever \(x, y \in E\) and \(x < z < y\), then \(z \in E\).

Why this is the intermediate value theorem. Section 5 shows that a continuous image of a connected set is connected. Apply it to \(f : [a,b] \to \mathbb{R}\): \([a,b]\) is connected, so \(f([a,b])\) is a connected subset of \(\mathbb{R}\), hence an interval, hence it contains every value between \(f(a)\) and \(f(b)\). The intermediate value theorem is a one-line corollary once connectedness is available — which is the whole case for introducing it.

4. Limits and Continuity

THE TWO DEFINITIONS, AND THE THIRD

Let \(f : X \to Y\) between metric spaces. Then \(f\) is continuous at \(p\) if

\[ \forall \varepsilon > 0 \ \exists \delta > 0 : \ d_X(x,p) < \delta \Rightarrow d_Y\big(f(x), f(p)\big) < \varepsilon. \]

Equivalently, by sequences: \(x_n \to p \Rightarrow f(x_n) \to f(p)\).

And globally: \(f\) is continuous on \(X\) if and only if \(f^{-1}(V)\) is open in \(X\) for every open \(V \subseteq Y\).

The third form is the useful one. It mentions no \(\varepsilon\), no \(\delta\) and no point, which is why the proofs in section 5 are three lines each. It is also exactly the definition of a measurable function in Probability Theory, Unit 1, with "open" replaced by "Borel" — measurability is continuity's measure-theoretic cousin, and the resemblance is not an accident.

UNIFORM CONTINUITY

\(f\) is uniformly continuous on \(X\) if the same \(\delta\) works at every point:

\[ \forall \varepsilon > 0 \ \exists \delta > 0 : \ d_X(x,y) < \delta \Rightarrow d_Y\big(f(x), f(y)\big) < \varepsilon \ \text{ for all } x, y. \]

Ordinary continuity lets \(\delta\) depend on the point; uniform continuity does not. The difference is real: \(f(x) = 1/x\) is continuous on \((0,1)\) but not uniformly so, because near \(0\) the required \(\delta\) shrinks without limit. On \([a,1]\) with \(a > 0\) it is uniformly continuous — and \([a,1]\) is compact, which is the theorem of the next section.

5. Continuity, Compactness and Connectedness

THEOREM 1 — A CONTINUOUS IMAGE OF A COMPACT SET IS COMPACT

Statement. If \(f : X \to Y\) is continuous and \(K \subseteq X\) is compact, then \(f(K)\) is compact.

Step 1 — start from an arbitrary cover of the image. Let \(\{V_\alpha\}\) be an open cover of \(f(K)\).

Step 2 — pull it back. Each \(f^{-1}(V_\alpha)\) is open, by the global characterisation of continuity in section 4. And these sets cover \(K\): if \(x \in K\) then \(f(x) \in f(K)\), so \(f(x) \in V_\alpha\) for some \(\alpha\), and then \(x \in f^{-1}(V_\alpha)\).

Step 3 — use compactness of \(K\). Finitely many suffice: \(K \subseteq f^{-1}(V_{\alpha_1}) \cup \cdots \cup f^{-1}(V_{\alpha_n})\).

Step 4 — push forward. Applying \(f\),

\[ f(K) \subseteq V_{\alpha_1} \cup \cdots \cup V_{\alpha_n}, \]

a finite subcover of the original cover. Since the cover was arbitrary, \(f(K)\) is compact. \(\blacksquare\)

THE THREE CONSEQUENCES

Let \(f\) be continuous and real-valued on a compact \(K\). Then

\[ \begin{aligned} &\textbf{(a) } f \text{ is bounded on } K \\ &\textbf{(b) } f \text{ attains its supremum and its infimum on } K \\ &\textbf{(c) } f \text{ is uniformly continuous on } K \end{aligned} \]

(a) and (b) follow at once: \(f(K)\) is compact, hence closed and bounded, and a non-empty closed bounded subset of \(\mathbb{R}\) contains its own supremum and infimum.

What (b) is for. Every maximum-likelihood argument that says "the maximum exists" is invoking it. A likelihood continuous on a compact parameter space attains its maximum; on an open or unbounded one it need not, and the estimator may fail to exist. That is not a technicality — it is why the uniform \(U(0,\theta)\), whose likelihood is maximised at the boundary, needs separate treatment in Estimation Theory (STS-201).

THEOREM 2 — A CONTINUOUS IMAGE OF A CONNECTED SET IS CONNECTED

Statement. If \(f : X \to Y\) is continuous and \(E \subseteq X\) is connected, then \(f(E)\) is connected.

Proof, by contradiction. Suppose \(f(E) = A \cup B\) is a separation. Put \(G = E \cap f^{-1}(A)\) and \(H = E \cap f^{-1}(B)\). Then \(E = G \cup H\), both are non-empty, and continuity carries the separation back: any limit point of \(G\) lying in \(H\) would be mapped into a limit point of \(A\) lying in \(B\), which a separation forbids. So \(E\) is separated, contradicting the hypothesis. \(\blacksquare\)

As shown in section 3, the intermediate value theorem is the special case \(X = [a,b]\), \(Y = \mathbb{R}\).

6. Discontinuities and Monotonic Functions

THE CLASSIFICATION

For \(f : (a,b) \to \mathbb{R}\) and \(x \in (a,b)\), write \(f(x+)\) and \(f(x-)\) for the one-sided limits where they exist. Then

\(f(x) = \sin(1/x)\) at \(x = 0\) has a discontinuity of the second kind: the values oscillate through \([-1,1]\) without settling, so neither one-sided limit exists.

MONOTONIC FUNCTIONS, AND WHY THIS IS A DISTRIBUTION-FUNCTION THEOREM

Theorem. Let \(f\) be monotonic on \((a,b)\). Then

  1. \(f(x+)\) and \(f(x-)\) exist at every \(x\), so every discontinuity is of the first kind — a monotonic function cannot oscillate;
  2. the set of discontinuities is at most countable.

Why (2) holds. At each discontinuity \(x\) the interval \((f(x-), f(x+))\) is non-empty, and for different \(x\) these intervals are disjoint by monotonicity. Each contains a rational number, and distinct intervals contain distinct rationals, so the discontinuities inject into \(\mathbb{Q}\). A set injecting into a countable set is countable. \(\blacksquare\)

What it says about distribution functions. Every \(F\) is non-decreasing (Probability Theory, Unit 1), so this theorem applies verbatim: a distribution function has at most countably many jumps, and each jump has height \(P(X = x)\). Hence a random variable has at most countably many atoms — a fact used constantly and almost never proved. This is the proof, and it is a theorem of real analysis rather than of probability.

7. Differentiation

THE MEAN VALUE THEOREMS \[ \begin{aligned} \textbf{Rolle:}\quad & f \text{ continuous on } [a,b], \text{ differentiable on } (a,b), f(a) = f(b) \Rightarrow \exists c : f'(c) = 0 \\ \textbf{Lagrange:}\quad & f(b) - f(a) = (b-a) f'(c) \text{ for some } c \in (a,b) \\ \textbf{Cauchy:}\quad & \big[f(b) - f(a)\big]g'(c) = \big[g(b) - g(a)\big]f'(c) \\ \textbf{Taylor:}\quad & f(b) = \sum_{k=0}^{n-1}\frac{f^{(k)}(a)}{k!}(b-a)^{k} + \frac{f^{(n)}(c)}{n!}(b-a)^{n} \end{aligned} \]

Each is proved from the one above it by applying it to a cleverly chosen auxiliary function; all of them ultimately rest on the attainment of extrema on a compact interval, which is consequence (b) of section 5.

Where statistics uses Taylor's theorem. The delta method is Taylor to first order with the remainder controlled: if \(\sqrt{n}(T_n - \theta) \xrightarrow{d} N(0, \sigma^{2})\) and \(g\) is differentiable at \(\theta\) with \(g'(\theta) \ne 0\), then

\[ \sqrt{n}\big(g(T_n) - g(\theta)\big) \xrightarrow{d} N\big(0, \sigma^{2}[g'(\theta)]^{2}\big). \]

The proof writes \(g(T_n) - g(\theta) = g'(\xi_n)(T_n - \theta)\) by the mean value theorem, with \(\xi_n\) between \(T_n\) and \(\theta\), notes that \(\xi_n \xrightarrow{P} \theta\) and hence \(g'(\xi_n) \xrightarrow{P} g'(\theta)\), and finishes with Slutzky's theorem from Probability Theory, Unit 3. Every step is from this unit or that one.

Key Take-aways