Let \(\alpha\) be monotonically increasing on \([a,b]\). For a partition \(P = \{a = x_0 < x_1 < \cdots < x_n = b\}\) write
\[ \Delta\alpha_i = \alpha(x_i) - \alpha(x_{i-1}) \ (\ge 0), \qquad M_i = \sup_{[x_{i-1},x_i]} f, \qquad m_i = \inf_{[x_{i-1},x_i]} f, \]and form the upper and lower sums
\[ U(P, f, \alpha) = \sum_{i=1}^{n} M_i \Delta\alpha_i, \qquad L(P, f, \alpha) = \sum_{i=1}^{n} m_i \Delta\alpha_i. \]Then \(\overline{\int} f\,d\alpha = \inf_P U\) and \(\underline{\int} f\,d\alpha = \sup_P L\), and \(f\) is R–S integrable with respect to \(\alpha\), written \(f \in \mathcal{R}(\alpha)\), when the two agree. Their common value is \(\int_a^b f \, d\alpha\).
Taking \(\alpha(x) = x\) gives \(\Delta\alpha_i = \Delta x_i\) and recovers the Riemann integral exactly. Nothing is lost; the increments are simply allowed to be uneven, and \(\alpha\) decides how much weight each part of the interval carries.
Statement. \(f \in \mathcal{R}(\alpha)\) on \([a,b]\) if and only if for every \(\varepsilon > 0\) there is a partition \(P\) with
\[ U(P, f, \alpha) - L(P, f, \alpha) = \sum_{i=1}^{n}\left(M_i - m_i\right)\Delta\alpha_i \;<\; \varepsilon. \]Why it is the working criterion. It never mentions the value of the integral, so integrability can be settled before the integral is known. It also shows at once where trouble can come from: a large oscillation \(M_i - m_i\) is harmless if \(\Delta\alpha_i\) is small there, and a large \(\Delta\alpha_i\) is harmless if \(f\) barely moves. The integral fails to exist only when \(f\) and \(\alpha\) are discontinuous at the same point, from the same side.
Three sufficient conditions follow: \(f\) continuous and \(\alpha\) increasing; \(f\) monotonic and \(\alpha\) continuous and increasing; \(f\) bounded with finitely many discontinuities, none shared with \(\alpha\).
The second line — linearity in the integrator — has no counterpart in the Riemann integral, and it is what makes a mixed distribution decompose: write \(F = pF_d + (1-p)F_c\) and the expectation splits into a sum plus an ordinary integral.
Statement. Suppose \(\alpha\) is constant except for jumps of size \(c_1, \ldots, c_m\) at interior points \(s_1, \ldots, s_m\), and \(f\) is continuous at each \(s_k\). Then
\[ \int_a^b f \, d\alpha = \sum_{k=1}^{m} c_k \, f(s_k). \]Why. Take a partition whose subintervals isolate the jumps. On a subinterval containing no \(s_k\), \(\Delta\alpha_i = 0\) and the term vanishes whatever \(f\) does. On one containing \(s_k\), \(\Delta\alpha_i = c_k\), and shrinking that subinterval forces \(M_i\) and \(m_i\) both to \(f(s_k)\) by continuity. \(\blacksquare\)
The consequence. If \(X\) is discrete with \(P(X = x_k) = p_k\), then \(F\) is a step function with jumps \(p_k\) at \(x_k\), and
\[ E\big(g(X)\big) = \int_{-\infty}^{\infty} g \, dF = \sum_k p_k \, g(x_k), \]which is the discrete expectation formula — derived, not assumed.
Given. \(f(x) = x^{2}\) on \([0,5]\), and \(\alpha\) constant except for jumps of \(1\) at \(x = 1\), \(2\) at \(x = 2\) and \(3\) at \(x = 4\).
Asked. Evaluate \(\int_0^5 x^{2}\,d\alpha\).
Step 1 — check the hypotheses. \(f(x) = x^{2}\) is continuous everywhere, in particular at \(1, 2, 4\), so the theorem applies.
Step 2 — evaluate \(f\) at each jump point.
\[ f(1) = 1^{2} = 1, \qquad f(2) = 2^{2} = 4, \qquad f(4) = 4^{2} = 16. \]Step 3 — weight by the jumps and add.
\[ \int_0^5 x^{2}\,d\alpha = 1(1) + 2(4) + 3(16) = 1 + 8 + 48 = 57. \]Step 4 — a sanity check on the size. The bound from the linear properties gives
\[ \left|\int_0^5 x^{2}\,d\alpha\right| \le \left(\sup_{[0,5]} x^{2}\right) \left[\alpha(5) - \alpha(0)\right] = 25 \times (1 + 2 + 3) = 25 \times 6 = 150, \]and \(57 \le 150\). \(\checkmark\)
Interpretation. Read probabilistically with \(\alpha\) scaled to total \(1\): a variable taking \(1, 2, 4\) with probabilities \(\tfrac16, \tfrac26, \tfrac36\) has
\[ E\left(X^{2}\right) = \frac{57}{6} = 9.5. \]The whole of discrete expectation is the R–S integral against a step function, with no separate theory required.
Statement. If \(f \in \mathcal{R}(\alpha)\) on \([a,b]\) then \(\alpha \in \mathcal{R}(f)\) and
\[ \int_a^b f \, d\alpha + \int_a^b \alpha \, df = f(b)\alpha(b) - f(a)\alpha(a). \]The point. The relation is perfectly symmetric in \(f\) and \(\alpha\): either can play integrand or integrator, and each integral exists exactly when the other does. Nothing in the Riemann theory resembles this, because there the integrator is always \(x\).
Given. \(f(x) = x\) and \(\alpha(x) = x^{2}\) on \([0,2]\).
Step 1 — directly. \(\alpha\) is differentiable with \(\alpha'(x) = 2x\), so \(d\alpha = 2x\,dx\) and
\[ \int_0^2 x \, d\left(x^{2}\right) = \int_0^2 x (2x)\,dx = \int_0^2 2x^{2}\,dx = \left[\frac{2x^{3}}{3}\right]_0^2 = \frac{16}{3} = 5.333333. \]Step 2 — the boundary term.
\[ f(2)\alpha(2) - f(0)\alpha(0) = 2 \times 4 - 0 \times 0 = 8. \]Step 3 — the other integral. \(df = dx\), so
\[ \int_0^2 \alpha \, df = \int_0^2 x^{2}\,dx = \left[\frac{x^{3}}{3}\right]_0^2 = \frac{8}{3} = 2.666667. \]Step 4 — check the identity.
\[ \frac{16}{3} + \frac{8}{3} = \frac{24}{3} = 8 = f(2)\alpha(2) - f(0)\alpha(0). \checkmark \]Interpretation. Read with \(\alpha = F\), this is the identity behind the survival-function formula for a non-negative random variable:
\[ E(X) = \int_0^\infty x \, dF(x) = \int_0^\infty \big[1 - F(x)\big]\,dx, \]obtained by parts with the boundary terms vanishing when \(E(X) < \infty\). It is the reason a mean can be read off a survival curve as an area — the standard device in reliability and in actuarial work.
Statement. If \(f\) has a continuous derivative on \([a,b]\), then
\[ \sum_{a < n \le b} f(n) = \int_a^b f(x)\,dx + \int_a^b \left(x - \lfloor x \rfloor - \tfrac12\right) f'(x)\,dx + \tfrac12\left[f(a) - f(b)\right] + \text{(boundary adjustments)}. \]It is the R–S integral against the step function \(\lfloor x \rfloor\), whose jumps are \(1\) at each integer, compared with the ordinary integral. The difference between a sum and an integral is thereby made explicit rather than estimated.
Given. \(f(x) = 1/x\), so \(\sum_{k \le n} 1/k\) is the partial harmonic sum.
Step 1 — what Euler's formula delivers. Applied to \(1/x\) it gives
\[ \sum_{k=1}^{n} \frac{1}{k} = \ln n + \gamma + \frac{1}{2n} - \frac{1}{12n^{2}} + O\!\left(n^{-4}\right), \]with \(\gamma = 0.5772156649\) the Euler–Mascheroni constant.
Step 2 — compare, keeping only \(\ln n + \gamma + 1/(2n)\).
| \(n\) | exact sum | \(\ln n + \gamma + \frac{1}{2n}\) | difference |
|---|---|---|---|
| 10 | 2.92896825 | 2.92980076 | −8.33 × 10⁻⁴ |
| 100 | 5.18737752 | 5.18738585 | −8.33 × 10⁻⁶ |
| 1000 | 7.48547086 | 7.48547094 | −8.33 × 10⁻⁸ |
Step 3 — read the error. A tenfold rise in \(n\) divides the error by \(100\), so it is of order \(n^{-2}\); and the leading coefficient is visibly \(-1/12 = -0.0833\ldots\), which is exactly the next term of the expansion. The formula is not merely a good approximation — its error term is itself predicted, and confirmed here to three significant figures at every \(n\).
Interpretation. This is the technique behind Stirling's formula, and hence behind every normal approximation to a factorial — including the derivation of the De Moivre–Laplace theorem in Probability Theory, Unit 4. Note also the divergence it quantifies: \(\sum 1/k \approx \ln n\) grows without bound, which is why \(\sum 1/n\) diverges and, through Borel–Cantelli, why events of probability \(1/n\) occur infinitely often.
| type of \(X\) | what \(F\) does | what the integral becomes |
|---|---|---|
| discrete | steps, jumps \(p_k\) | \(\sum_k p_k\,g(x_k)\) |
| continuous | differentiable, \(F' = f\) | \(\int g(x) f(x)\,dx\) |
| mixed | jumps plus a continuous part | a sum plus an integral |
The mixed variable of Probability Theory, Unit 1, Example 1.5 — a quarter of its mass at \(0\), a quarter at \(1\), half spread uniformly — has
\[ E(X) = \underbrace{\tfrac14(0) + \tfrac14(1)}_{\text{the two jumps}} + \underbrace{\int_0^1 x \left(\tfrac12\right)dx}_{\text{the uniform part}} = 0 + \tfrac14 + \tfrac12 \times \tfrac12 = \tfrac14 + \tfrac14 = 0.5, \]a single calculation requiring no case analysis. That is the entire return on the Riemann–Stieltjes integral for a statistician, and it is why this unit sits in the programme.