For \(f\) on \([a,b]\) and a partition \(P = \{a = x_0 < \cdots < x_n = b\}\), put
\[ V(P, f) = \sum_{i=1}^{n}\left|f(x_i) - f(x_{i-1})\right|. \]The total variation is \(V_a^b(f) = \sup_P V(P,f)\), and \(f\) is of bounded variation, written \(f \in BV[a,b]\), when this supremum is finite.
Read it as total distance travelled, up and down, rather than net displacement. A monotonic \(f\) has \(V_a^b(f) = |f(b) - f(a)|\), because it never doubles back.
Statement. \(f \in BV[a,b]\) if and only if \(f\) is the difference of two increasing functions:
\[ f = f_1 - f_2, \qquad f_1, f_2 \text{ increasing}. \]One explicit choice is \(f_1(x) = V_a^x(f)\) and \(f_2(x) = V_a^x(f) - f(x)\), both of which can be checked to be increasing.
What it delivers. Unit 2 built \(\int f \, d\alpha\) for increasing \(\alpha\) only. By Jordan and linearity in the integrator,
\[ \int_a^b f \, d\alpha = \int_a^b f \, d\alpha_1 - \int_a^b f \, d\alpha_2, \]so the whole theory extends to every integrator of bounded variation at no cost. And since a function of bounded variation inherits from monotonic functions the property proved in Unit 1, it too has only jump discontinuities, at most countably many.
Given. \(f(x) = x^{2}\) on \([-1, 2]\).
Step 1 — split at the turning point. \(f\) decreases on \([-1,0]\) and increases on \([0,2]\), so the variation is additive across \(x = 0\) and each piece is monotonic.
Step 2 — the two pieces.
\[ V_{-1}^{0}(f) = |f(0) - f(-1)| = |0 - 1| = 1, \qquad V_{0}^{2}(f) = |f(2) - f(0)| = |4 - 0| = 4. \]Step 3 — add. \(V_{-1}^{2}(f) = 1 + 4 = 5\).
Step 4 — compare with the net change.
\[ f(2) - f(-1) = 4 - 1 = 3 \;<\; 5 = V_{-1}^{2}(f). \]Interpretation. The graph descends \(1\) and then climbs \(4\), a journey of \(5\), while arriving only \(3\) above where it began. The gap is exactly twice the descent, and it is zero precisely when \(f\) is monotonic.
Given.
\[ f(x) = \begin{cases} x \sin(1/x), & 0 < x \le 1, \\ 0, & x = 0. \end{cases} \]Step 1 — \(f\) is continuous. On \((0,1]\) it is a product of continuous functions. At \(0\), \(|f(x)| = |x \sin(1/x)| \le |x| \to 0 = f(0)\), so it is continuous there too. So continuity alone does not give bounded variation.
Step 2 — choose partition points at the peaks. Take \(x_k = \dfrac{2}{(2k+1)\pi}\), at which \(\sin(1/x_k) = \pm 1\), so
\[ |f(x_k)| = x_k = \frac{2}{(2k+1)\pi}. \]Consecutive peaks alternate in sign, so each step of the partition contributes at least \(x_k\) to \(V(P, f)\).
Step 3 — sum the peaks.
| peaks used, \(K\) | \(\sum_{k=1}^{K} \frac{2}{(2k+1)\pi}\) |
|---|---|
| 10 | 0.751768 |
| 100 | 1.457425 |
| 1000 | 2.187510 |
| 2000 | 2.407986 |
Step 4 — identify the growth. The terms behave like \(\dfrac{1}{\pi k}\), so the sum grows like \(\dfrac{\ln K}{\pi}\) — slowly, but without bound. Each tenfold rise in \(K\) adds about \(\dfrac{\ln 10}{\pi} = 0.7329\), and the table shows increments of \(0.7057\) and \(0.7301\), approaching it. So \(V_0^1(f) = \infty\) and \(f \notin BV[0,1]\).
Interpretation. Continuity controls how far \(f\) moves for a small change in \(x\); it says nothing about how often it changes direction. Here the oscillations shrink in height like \(1/k\) but there are infinitely many of them, and \(\sum 1/k\) diverges — the same harmonic divergence measured in Unit 2, Example 2.3. Replacing \(x\) by \(x^{2}\) in front, so that the peaks shrink like \(1/k^{2}\), makes the sum converge and the function is then of bounded variation. The borderline is exactly \(\sum 1/k^{p}\).
Let \(\alpha \in BV[a,b]\).
\[ \begin{aligned} \textbf{(S1)}\quad & f \text{ continuous on } [a,b] \Rightarrow f \in \mathcal{R}(\alpha) \\ \textbf{(S2)}\quad & f \in BV[a,b] \text{ and } \alpha \text{ continuous} \Rightarrow f \in \mathcal{R}(\alpha) \\ \textbf{(S3)}\quad & f \text{ and } \alpha \text{ share no common discontinuity} \Rightarrow f \in \mathcal{R}(\alpha) \\ \textbf{(S4)}\quad & \text{necessary and sufficient (Riemann's condition): } \forall \varepsilon > 0 \ \exists P : \sum_i \operatorname{osc}_i(f)\,\Delta\alpha_i < \varepsilon \\ \textbf{(S5)}\quad & \alpha' \text{ continuous} \Rightarrow \int_a^b f \, d\alpha = \int_a^b f(x)\,\alpha'(x)\,dx \end{aligned} \]where \(\operatorname{osc}_i(f) = M_i - m_i\).
(S3) is the one to remember, and the symmetry of (S1) and (S2) is worth noticing: whichever of \(f\) and \(\alpha\) is the rougher, the other must be smooth at the same places. (S5) is the bridge back to ordinary calculus, and with \(\alpha = F\) and \(F' = f\) it is exactly \(E(g(X)) = \int g(x)f(x)\,dx\).
Statement. Let \(I(t) = \displaystyle\int_a^b f(x,t)\,dx\). If \(f\) and \(\partial f/\partial t\) are continuous on \([a,b] \times [c,d]\), then \(I\) is differentiable and
\[ \frac{dI}{dt} = \int_a^b \frac{\partial f}{\partial t}(x,t)\,dx. \]With variable limits \(u(t)\) and \(v(t)\), also differentiable,
\[ \frac{d}{dt}\int_{u(t)}^{v(t)} f(x,t)\,dx = \int_{u(t)}^{v(t)} \frac{\partial f}{\partial t}\,dx + f\big(v(t), t\big)v'(t) - f\big(u(t), t\big)u'(t). \]The hypothesis is not decoration. Continuity of the partial derivative is what makes the difference quotient converge uniformly, and uniformity is what permits the interchange — the theme of Unit 4. Without it the rule fails, just as the interchange of limit and integral fails in Probability Theory, Unit 1, Example 1.4.
Given.
\[ I(a) = \int_0^1 \frac{x^{a} - 1}{\ln x}\,dx, \qquad a > -1. \]The integrand has no elementary antiderivative, so a direct attack fails.
Step 1 — differentiate with respect to the parameter. Since \(\dfrac{\partial}{\partial a} x^{a} = x^{a}\ln x\),
\[ \frac{\partial}{\partial a}\left(\frac{x^{a} - 1}{\ln x}\right) = \frac{x^{a}\ln x}{\ln x} = x^{a}. \]The logarithm cancels exactly — which is the whole trick.
Step 2 — integrate the derivative. By Leibniz's rule,
\[ I'(a) = \int_0^1 x^{a}\,dx = \left[\frac{x^{a+1}}{a+1}\right]_0^1 = \frac{1}{a+1}. \]Step 3 — integrate back in \(a\).
\[ I(a) = \int \frac{da}{a+1} = \ln(a+1) + C. \]Step 4 — fix the constant. At \(a = 0\) the integrand is \((x^{0} - 1)/\ln x = 0\), so \(I(0) = 0\). Then \(0 = \ln 1 + C = C\), giving
\[ I(a) = \ln(1 + a). \]Step 5 — check numerically. Evaluating the original integral directly:
| \(a\) | numerical value of the integral | \(\ln(1+a)\) |
|---|---|---|
| 1 | 0.693147 | 0.693147 |
| 2 | 1.098612 | 1.098612 |
| 3 | 1.386294 | 1.386294 |
Agreement to six decimals at all three values. \(\checkmark\)
Interpretation. The same manoeuvre is how moments are extracted from a moment generating function — \(E(X^{k}) = M^{(k)}(0)\) is differentiation under \(\int e^{tx} f(x)\,dx\) — and how the score function \(\partial \ln L / \partial \theta\) acquires mean zero in Estimation Theory (STS-201). That last result, \(E(\partial \ln L/\partial\theta) = 0\), is precisely the statement that \(\partial/\partial\theta\) and \(\int\) may be exchanged, and the regularity conditions quoted there are the hypotheses of this theorem. The uniform \(U(0,\theta)\) violates them because its support depends on \(\theta\), which is why the Cramér–Rao bound does not apply to it.
Fubini's theorem. If \(f\) is integrable over \(A \times B\), that is \(\displaystyle\iint |f| < \infty\), then
\[ \int_A\left(\int_B f(x,y)\,dy\right)dx = \int_B\left(\int_A f(x,y)\,dx\right)dy. \]Tonelli's theorem. If \(f \ge 0\) and measurable, the same equality holds with no integrability assumption — both sides may be \(+\infty\) together.
How they are used in practice. Tonelli on \(|f|\) first, to decide whether the double integral is finite; then Fubini on \(f\) to do the swap. Absolute integrability is the condition, and without it the two orders can genuinely differ.
Given. \(X \ge 0\) with density \(f\) and \(E(X) < \infty\).
Asked. Prove \(E(X) = \displaystyle\int_0^\infty P(X > t)\,dt\).
Step 1 — write \(x\) as an integral. For \(x \ge 0\), \(x = \int_0^{x} dt\), so
\[ E(X) = \int_0^\infty x f(x)\,dx = \int_0^\infty \left(\int_0^{x} dt\right) f(x)\,dx. \]Step 2 — make the region explicit. The inner limits say \(0 \le t \le x\), so the double integral runs over \(\{(x,t) : 0 \le t \le x < \infty\}\):
\[ = \iint_{0 \le t \le x} f(x)\,dt\,dx. \]Step 3 — check the swap is legitimate. The integrand \(f(x) \ge 0\), so Tonelli applies with no further condition.
Step 4 — swap, and re-read the region. Fixing \(t\) first, the condition \(t \le x\) means \(x\) runs from \(t\) to \(\infty\):
\[ = \int_0^\infty \left(\int_t^\infty f(x)\,dx\right)dt. \]Step 5 — recognise the inner integral. \(\int_t^\infty f(x)\,dx = P(X > t)\), so
\[ E(X) = \int_0^\infty P(X > t)\,dt. \quad \blacksquare \]Step 6 — verify on a case where both sides are known. For \(X \sim\) Exponential\((\theta)\), \(P(X > t) = e^{-\theta t}\) and
\[ \int_0^\infty e^{-\theta t}\,dt = \left[\frac{-e^{-\theta t}}{\theta}\right]_0^\infty = \frac{1}{\theta} = E(X). \checkmark \]Interpretation. Compare Unit 2, Example 2.2, where the same identity came from integration by parts. Two routes, one result — and this one extends without change to \(E(X^{k}) = \int_0^\infty k t^{k-1} P(X > t)\,dt\), which is the standard way to obtain higher moments from a survival function.