Skip to the content

Topics Covered

Bounded Variation Jordan Decomposition Existence Conditions Leibniz's Rule Fubini Tonelli
On this page
  1. 1. Functions of Bounded Variation
  2. 2. Existence Conditions for the R–S Integral
  3. 3. Differentiation Under the Integral Sign
  4. 4. Interchanging the Order of Integration
  5. Key Take-aways
Where this unit starts. It continues directly from Unit 2. That unit required the integrator \(\alpha\) to be increasing. This one removes that restriction — the right class is functions of bounded variation — and then takes up the two interchange questions that every derivation in statistics runs into: differentiating inside an integral, and swapping the order of a double integral.

1. Functions of Bounded Variation

DEFINITION

For \(f\) on \([a,b]\) and a partition \(P = \{a = x_0 < \cdots < x_n = b\}\), put

\[ V(P, f) = \sum_{i=1}^{n}\left|f(x_i) - f(x_{i-1})\right|. \]

The total variation is \(V_a^b(f) = \sup_P V(P,f)\), and \(f\) is of bounded variation, written \(f \in BV[a,b]\), when this supremum is finite.

Read it as total distance travelled, up and down, rather than net displacement. A monotonic \(f\) has \(V_a^b(f) = |f(b) - f(a)|\), because it never doubles back.

PROPERTIES \[ \begin{aligned} &f \in BV \Rightarrow f \text{ is bounded} \\ &f, g \in BV \Rightarrow f \pm g,\ fg \in BV \\ &V_a^b(f) = V_a^c(f) + V_c^b(f) \quad (a < c < b) \\ &f \text{ monotonic} \Rightarrow f \in BV \text{ with } V_a^b(f) = |f(b) - f(a)| \\ &f' \text{ exists and is bounded on } [a,b] \Rightarrow f \in BV \end{aligned} \]
JORDAN'S DECOMPOSITION, AND WHY IT IS THE POINT

Statement. \(f \in BV[a,b]\) if and only if \(f\) is the difference of two increasing functions:

\[ f = f_1 - f_2, \qquad f_1, f_2 \text{ increasing}. \]

One explicit choice is \(f_1(x) = V_a^x(f)\) and \(f_2(x) = V_a^x(f) - f(x)\), both of which can be checked to be increasing.

What it delivers. Unit 2 built \(\int f \, d\alpha\) for increasing \(\alpha\) only. By Jordan and linearity in the integrator,

\[ \int_a^b f \, d\alpha = \int_a^b f \, d\alpha_1 - \int_a^b f \, d\alpha_2, \]

so the whole theory extends to every integrator of bounded variation at no cost. And since a function of bounded variation inherits from monotonic functions the property proved in Unit 1, it too has only jump discontinuities, at most countably many.

EXAMPLE 3.1 — VARIATION AGAINST NET CHANGE

Given. \(f(x) = x^{2}\) on \([-1, 2]\).

Step 1 — split at the turning point. \(f\) decreases on \([-1,0]\) and increases on \([0,2]\), so the variation is additive across \(x = 0\) and each piece is monotonic.

Step 2 — the two pieces.

\[ V_{-1}^{0}(f) = |f(0) - f(-1)| = |0 - 1| = 1, \qquad V_{0}^{2}(f) = |f(2) - f(0)| = |4 - 0| = 4. \]

Step 3 — add. \(V_{-1}^{2}(f) = 1 + 4 = 5\).

Step 4 — compare with the net change.

\[ f(2) - f(-1) = 4 - 1 = 3 \;<\; 5 = V_{-1}^{2}(f). \]

Interpretation. The graph descends \(1\) and then climbs \(4\), a journey of \(5\), while arriving only \(3\) above where it began. The gap is exactly twice the descent, and it is zero precisely when \(f\) is monotonic.

EXAMPLE 3.2 — A CONTINUOUS FUNCTION THAT IS NOT OF BOUNDED VARIATION

Given.

\[ f(x) = \begin{cases} x \sin(1/x), & 0 < x \le 1, \\ 0, & x = 0. \end{cases} \]

Step 1 — \(f\) is continuous. On \((0,1]\) it is a product of continuous functions. At \(0\), \(|f(x)| = |x \sin(1/x)| \le |x| \to 0 = f(0)\), so it is continuous there too. So continuity alone does not give bounded variation.

Step 2 — choose partition points at the peaks. Take \(x_k = \dfrac{2}{(2k+1)\pi}\), at which \(\sin(1/x_k) = \pm 1\), so

\[ |f(x_k)| = x_k = \frac{2}{(2k+1)\pi}. \]

Consecutive peaks alternate in sign, so each step of the partition contributes at least \(x_k\) to \(V(P, f)\).

Step 3 — sum the peaks.

peaks used, \(K\)\(\sum_{k=1}^{K} \frac{2}{(2k+1)\pi}\)
100.751768
1001.457425
10002.187510
20002.407986

Step 4 — identify the growth. The terms behave like \(\dfrac{1}{\pi k}\), so the sum grows like \(\dfrac{\ln K}{\pi}\) — slowly, but without bound. Each tenfold rise in \(K\) adds about \(\dfrac{\ln 10}{\pi} = 0.7329\), and the table shows increments of \(0.7057\) and \(0.7301\), approaching it. So \(V_0^1(f) = \infty\) and \(f \notin BV[0,1]\).

Interpretation. Continuity controls how far \(f\) moves for a small change in \(x\); it says nothing about how often it changes direction. Here the oscillations shrink in height like \(1/k\) but there are infinitely many of them, and \(\sum 1/k\) diverges — the same harmonic divergence measured in Unit 2, Example 2.3. Replacing \(x\) by \(x^{2}\) in front, so that the peaks shrink like \(1/k^{2}\), makes the sum converge and the function is then of bounded variation. The borderline is exactly \(\sum 1/k^{p}\).

2. Existence Conditions for the R–S Integral

THE STATEMENTS, AS THE SYLLABUS PRESCRIBES

Let \(\alpha \in BV[a,b]\).

\[ \begin{aligned} \textbf{(S1)}\quad & f \text{ continuous on } [a,b] \Rightarrow f \in \mathcal{R}(\alpha) \\ \textbf{(S2)}\quad & f \in BV[a,b] \text{ and } \alpha \text{ continuous} \Rightarrow f \in \mathcal{R}(\alpha) \\ \textbf{(S3)}\quad & f \text{ and } \alpha \text{ share no common discontinuity} \Rightarrow f \in \mathcal{R}(\alpha) \\ \textbf{(S4)}\quad & \text{necessary and sufficient (Riemann's condition): } \forall \varepsilon > 0 \ \exists P : \sum_i \operatorname{osc}_i(f)\,\Delta\alpha_i < \varepsilon \\ \textbf{(S5)}\quad & \alpha' \text{ continuous} \Rightarrow \int_a^b f \, d\alpha = \int_a^b f(x)\,\alpha'(x)\,dx \end{aligned} \]

where \(\operatorname{osc}_i(f) = M_i - m_i\).

(S3) is the one to remember, and the symmetry of (S1) and (S2) is worth noticing: whichever of \(f\) and \(\alpha\) is the rougher, the other must be smooth at the same places. (S5) is the bridge back to ordinary calculus, and with \(\alpha = F\) and \(F' = f\) it is exactly \(E(g(X)) = \int g(x)f(x)\,dx\).

3. Differentiation Under the Integral Sign

LEIBNIZ'S RULE

Statement. Let \(I(t) = \displaystyle\int_a^b f(x,t)\,dx\). If \(f\) and \(\partial f/\partial t\) are continuous on \([a,b] \times [c,d]\), then \(I\) is differentiable and

\[ \frac{dI}{dt} = \int_a^b \frac{\partial f}{\partial t}(x,t)\,dx. \]

With variable limits \(u(t)\) and \(v(t)\), also differentiable,

\[ \frac{d}{dt}\int_{u(t)}^{v(t)} f(x,t)\,dx = \int_{u(t)}^{v(t)} \frac{\partial f}{\partial t}\,dx + f\big(v(t), t\big)v'(t) - f\big(u(t), t\big)u'(t). \]

The hypothesis is not decoration. Continuity of the partial derivative is what makes the difference quotient converge uniformly, and uniformity is what permits the interchange — the theme of Unit 4. Without it the rule fails, just as the interchange of limit and integral fails in Probability Theory, Unit 1, Example 1.4.

EXAMPLE 3.3 — AN INTEGRAL DONE BY DIFFERENTIATING IT

Given.

\[ I(a) = \int_0^1 \frac{x^{a} - 1}{\ln x}\,dx, \qquad a > -1. \]

The integrand has no elementary antiderivative, so a direct attack fails.

Step 1 — differentiate with respect to the parameter. Since \(\dfrac{\partial}{\partial a} x^{a} = x^{a}\ln x\),

\[ \frac{\partial}{\partial a}\left(\frac{x^{a} - 1}{\ln x}\right) = \frac{x^{a}\ln x}{\ln x} = x^{a}. \]

The logarithm cancels exactly — which is the whole trick.

Step 2 — integrate the derivative. By Leibniz's rule,

\[ I'(a) = \int_0^1 x^{a}\,dx = \left[\frac{x^{a+1}}{a+1}\right]_0^1 = \frac{1}{a+1}. \]

Step 3 — integrate back in \(a\).

\[ I(a) = \int \frac{da}{a+1} = \ln(a+1) + C. \]

Step 4 — fix the constant. At \(a = 0\) the integrand is \((x^{0} - 1)/\ln x = 0\), so \(I(0) = 0\). Then \(0 = \ln 1 + C = C\), giving

\[ I(a) = \ln(1 + a). \]

Step 5 — check numerically. Evaluating the original integral directly:

\(a\)numerical value of the integral\(\ln(1+a)\)
10.6931470.693147
21.0986121.098612
31.3862941.386294

Agreement to six decimals at all three values. \(\checkmark\)

Interpretation. The same manoeuvre is how moments are extracted from a moment generating function — \(E(X^{k}) = M^{(k)}(0)\) is differentiation under \(\int e^{tx} f(x)\,dx\) — and how the score function \(\partial \ln L / \partial \theta\) acquires mean zero in Estimation Theory (STS-201). That last result, \(E(\partial \ln L/\partial\theta) = 0\), is precisely the statement that \(\partial/\partial\theta\) and \(\int\) may be exchanged, and the regularity conditions quoted there are the hypotheses of this theorem. The uniform \(U(0,\theta)\) violates them because its support depends on \(\theta\), which is why the Cramér–Rao bound does not apply to it.

4. Interchanging the Order of Integration

FUBINI AND TONELLI

Fubini's theorem. If \(f\) is integrable over \(A \times B\), that is \(\displaystyle\iint |f| < \infty\), then

\[ \int_A\left(\int_B f(x,y)\,dy\right)dx = \int_B\left(\int_A f(x,y)\,dx\right)dy. \]

Tonelli's theorem. If \(f \ge 0\) and measurable, the same equality holds with no integrability assumption — both sides may be \(+\infty\) together.

How they are used in practice. Tonelli on \(|f|\) first, to decide whether the double integral is finite; then Fubini on \(f\) to do the swap. Absolute integrability is the condition, and without it the two orders can genuinely differ.

EXAMPLE 3.4 — THE SURVIVAL FORMULA, BY SWAPPING

Given. \(X \ge 0\) with density \(f\) and \(E(X) < \infty\).

Asked. Prove \(E(X) = \displaystyle\int_0^\infty P(X > t)\,dt\).

Step 1 — write \(x\) as an integral. For \(x \ge 0\), \(x = \int_0^{x} dt\), so

\[ E(X) = \int_0^\infty x f(x)\,dx = \int_0^\infty \left(\int_0^{x} dt\right) f(x)\,dx. \]

Step 2 — make the region explicit. The inner limits say \(0 \le t \le x\), so the double integral runs over \(\{(x,t) : 0 \le t \le x < \infty\}\):

\[ = \iint_{0 \le t \le x} f(x)\,dt\,dx. \]

Step 3 — check the swap is legitimate. The integrand \(f(x) \ge 0\), so Tonelli applies with no further condition.

Step 4 — swap, and re-read the region. Fixing \(t\) first, the condition \(t \le x\) means \(x\) runs from \(t\) to \(\infty\):

\[ = \int_0^\infty \left(\int_t^\infty f(x)\,dx\right)dt. \]

Step 5 — recognise the inner integral. \(\int_t^\infty f(x)\,dx = P(X > t)\), so

\[ E(X) = \int_0^\infty P(X > t)\,dt. \quad \blacksquare \]

Step 6 — verify on a case where both sides are known. For \(X \sim\) Exponential\((\theta)\), \(P(X > t) = e^{-\theta t}\) and

\[ \int_0^\infty e^{-\theta t}\,dt = \left[\frac{-e^{-\theta t}}{\theta}\right]_0^\infty = \frac{1}{\theta} = E(X). \checkmark \]

Interpretation. Compare Unit 2, Example 2.2, where the same identity came from integration by parts. Two routes, one result — and this one extends without change to \(E(X^{k}) = \int_0^\infty k t^{k-1} P(X > t)\,dt\), which is the standard way to obtain higher moments from a survival function.

Key Take-aways