At Foundation level expectation is defined twice — a sum for discrete variables, an integral for continuous ones — and a variable that is neither gets no definition at all. Measure theory gives one definition for all three:
\[ E(X) = \int_{\Omega} X \, dP, \]the integral of the measurable function \(X\) with respect to the probability measure \(P\), provided \(E|X| < \infty\). The two familiar formulae are special cases, not rivals:
\[ E(X) = \sum_{i} x_i \, p(x_i) \quad \text{(discrete)}, \qquad E(X) = \int_{-\infty}^{\infty} x f(x) \, dx \quad \text{(continuous)}. \]The mixed variable of Unit 1, Example 1.5 — two atoms plus a uniform part — has an expectation under the general definition and under neither of the special ones.
For a measurable \(g : \mathbb{R} \to \mathbb{R}\) with \(E|g(X)| < \infty\),
\[ E\big(g(X)\big) = \int_{\Omega} g(X) \, dP = \int_{-\infty}^{\infty} g(x) \, dF(x), \]the last written as a Riemann–Stieltjes integral against the distribution function. The content of the equality — sometimes called the law of the unconscious statistician — is that the distribution of \(X\) alone is enough; the underlying \(\Omega\) never has to be examined.
For jointly distributed \(X\) and \(Y\), the conditional expectation of \(X\) given \(Y = y\) is
\[ E(X \mid Y = y) = \sum_{x} x \, P(X = x \mid Y = y) \quad \text{or} \quad \int_{-\infty}^{\infty} x \, f_{X \mid Y}(x \mid y) \, dx. \]This is a number for each \(y\). Writing \(h(y) = E(X \mid Y = y)\) and substituting the random \(Y\) gives \(E(X \mid Y) = h(Y)\), which is itself a random variable — a function of \(Y\), so measurable with respect to the sigma-field generated by \(Y\). That shift, from a number to a random variable, is what makes the next two identities possible.
The conditional variance is defined the same way:
\[ \operatorname{Var}(X \mid Y = y) = E\big(X^{2} \mid Y = y\big) - \big[E(X \mid Y = y)\big]^{2}. \]The second is read as: total variability = variability within the groups defined by \(Y\), plus variability between those groups. It is the identity behind the ANOVA decomposition, behind the Rao–Blackwell theorem of Estimation Theory (STS-201), and behind the variance formula for two-stage sampling in Sampling Theory (STS-204).
Statement. If \(E|X| < \infty\) then \(E[E(X \mid Y)] = E(X)\).
Step 1 — write the outer expectation as a sum over \(y\). The random variable \(E(X \mid Y)\) takes the value \(E(X \mid Y = y)\) when \(Y = y\), so by the definition of expectation
\[ E\big[E(X \mid Y)\big] = \sum_{y} E(X \mid Y = y) \, P(Y = y). \]Step 2 — expand the inner conditional expectation.
\[ = \sum_{y} \left[ \sum_{x} x \, P(X = x \mid Y = y) \right] P(Y = y). \]Step 3 — use the multiplication rule. By the definition of conditional probability, \(P(X = x \mid Y = y) \, P(Y = y) = P(X = x, Y = y)\). Taking \(P(Y = y)\) inside the inner sum, which is legitimate because it does not depend on \(x\),
\[ = \sum_{y} \sum_{x} x \, P(X = x, Y = y). \]Step 4 — interchange the order of summation. Absolute convergence is guaranteed by \(E|X| < \infty\), so Fubini's theorem permits the exchange:
\[ = \sum_{x} x \sum_{y} P(X = x, Y = y). \]Step 5 — the inner sum is a marginal. Summing a joint probability over all values of \(Y\) gives the marginal of \(X\): \(\sum_{y} P(X = x, Y = y) = P(X = x)\). Hence
\[ E\big[E(X \mid Y)\big] = \sum_{x} x \, P(X = x) = E(X). \quad \blacksquare \]Given. \(X\) and \(Y\) each take the values \(0, 1, 2\) with joint probabilities
| \(p(x,y)\) | \(x = 0\) | \(x = 1\) | \(x = 2\) | \(P(Y = y)\) |
|---|---|---|---|---|
| \(y = 0\) | 0.10 | 0.10 | 0.05 | 0.25 |
| \(y = 1\) | 0.15 | 0.20 | 0.10 | 0.45 |
| \(y = 2\) | 0.05 | 0.15 | 0.10 | 0.30 |
| \(P(X = x)\) | 0.30 | 0.45 | 0.25 | 1.00 |
Asked. Verify the tower property and the variance decomposition.
Step 1 — the marginals, and the check that they sum to 1. Adding each row: \(0.10 + 0.10 + 0.05 = 0.25\); \(0.15 + 0.20 + 0.10 = 0.45\); \(0.05 + 0.15 + 0.10 = 0.30\). These are \(P(Y = 0), P(Y = 1), P(Y = 2)\) and total \(0.25 + 0.45 + 0.30 = 1.00\). Adding each column gives \(P(X = 0) = 0.30\), \(P(X = 1) = 0.45\), \(P(X = 2) = 0.25\), again totalling \(1.00\).
Step 2 — the unconditional mean and variance of \(X\).
\[ E(X) = 0(0.30) + 1(0.45) + 2(0.25) = 0 + 0.45 + 0.50 = 0.95 \] \[ E(X^{2}) = 0^{2}(0.30) + 1^{2}(0.45) + 2^{2}(0.25) = 0 + 0.45 + 1.00 = 1.45 \] \[ \operatorname{Var}(X) = 1.45 - (0.95)^{2} = 1.45 - 0.9025 = 0.5475 \]Step 3 — the conditional distribution given \(Y = 0\). Divide the \(y = 0\) row by \(P(Y = 0) = 0.25\):
\[ P(X = 0 \mid Y = 0) = \frac{0.10}{0.25} = 0.4, \quad P(X = 1 \mid Y = 0) = \frac{0.10}{0.25} = 0.4, \quad P(X = 2 \mid Y = 0) = \frac{0.05}{0.25} = 0.2 \](these sum to \(1.0\), as a conditional distribution must). Hence
\[ E(X \mid Y = 0) = 0(0.4) + 1(0.4) + 2(0.2) = 0.8, \] \[ E(X^{2} \mid Y = 0) = 0(0.4) + 1(0.4) + 4(0.2) = 1.2, \qquad \operatorname{Var}(X \mid Y = 0) = 1.2 - 0.8^{2} = 1.2 - 0.64 = 0.56. \]Step 4 — the same for \(Y = 1\). Divide the \(y = 1\) row by \(0.45\): the conditional probabilities are \(\tfrac{0.15}{0.45} = \tfrac13\), \(\tfrac{0.20}{0.45} = \tfrac49\), \(\tfrac{0.10}{0.45} = \tfrac29\), summing to 1.
\[ E(X \mid Y = 1) = 0\left(\tfrac13\right) + 1\left(\tfrac49\right) + 2\left(\tfrac29\right) = \tfrac{4}{9} + \tfrac{4}{9} = \tfrac{8}{9} \approx 0.888889 \] \[ E(X^{2} \mid Y = 1) = \tfrac49 + 4\left(\tfrac29\right) = \tfrac{4}{9} + \tfrac{8}{9} = \tfrac{12}{9} = \tfrac43 \] \[ \operatorname{Var}(X \mid Y = 1) = \tfrac43 - \left(\tfrac89\right)^{2} = \tfrac{108}{81} - \tfrac{64}{81} = \tfrac{44}{81} \approx 0.543210 \]Step 5 — the same for \(Y = 2\). Divide the \(y = 2\) row by \(0.30\): \(\tfrac{0.05}{0.30} = \tfrac16\), \(\tfrac{0.15}{0.30} = \tfrac12\), \(\tfrac{0.10}{0.30} = \tfrac13\), summing to 1.
\[ E(X \mid Y = 2) = 0\left(\tfrac16\right) + 1\left(\tfrac12\right) + 2\left(\tfrac13\right) = \tfrac12 + \tfrac23 = \tfrac{7}{6} \approx 1.166667 \] \[ E(X^{2} \mid Y = 2) = \tfrac12 + 4\left(\tfrac13\right) = \tfrac12 + \tfrac43 = \tfrac{11}{6} \] \[ \operatorname{Var}(X \mid Y = 2) = \tfrac{11}{6} - \left(\tfrac76\right)^{2} = \tfrac{66}{36} - \tfrac{49}{36} = \tfrac{17}{36} \approx 0.472222 \]Step 6 — the tower property. Average the three conditional means against the marginal of \(Y\):
\[ E\big[E(X \mid Y)\big] = 0.8(0.25) + \tfrac89(0.45) + \tfrac76(0.30) = 0.20 + 0.40 + 0.35 = 0.95. \]This equals \(E(X) = 0.95\) from Step 2. \(\checkmark\)
Step 7 — the within-group term.
\[ E\big[\operatorname{Var}(X \mid Y)\big] = 0.56(0.25) + \tfrac{44}{81}(0.45) + \tfrac{17}{36}(0.30) \] \[ = 0.14 + 0.244444\ldots + 0.141666\ldots = \tfrac{947}{1800} = 0.526111\ldots \]Step 8 — the between-group term. The conditional means are \(0.8\), \(\tfrac89\), \(\tfrac76\), with mean \(0.95\) by Step 6, so
\[ \operatorname{Var}\big[E(X \mid Y)\big] = (0.8 - 0.95)^{2}(0.25) + \left(\tfrac89 - 0.95\right)^{2}(0.45) + \left(\tfrac76 - 0.95\right)^{2}(0.30) \] \[ = (-0.15)^{2}(0.25) + (-0.061111)^{2}(0.45) + (0.216667)^{2}(0.30) \] \[ = 0.005625 + 0.001681 + 0.014083 = \tfrac{77}{3600} = 0.021389\ldots \]Step 9 — add them.
\[ 0.526111\ldots + 0.021389\ldots = \frac{947}{1800} + \frac{77}{3600} = \frac{1894 + 77}{3600} = \frac{1971}{3600} = \frac{219}{400} = 0.5475. \]This equals \(\operatorname{Var}(X) = 0.5475\) from Step 2. \(\checkmark\)
Interpretation. Knowing \(Y\) removes only \(0.021389 / 0.5475 = 3.91\%\) of the variability of \(X\); the other \(96.09\%\) survives inside the groups. In regression language that ratio is \(R^{2}\), and a value this small says \(Y\) is a weak predictor of \(X\) here even though the two are certainly not independent — if they were, every row of conditional probabilities would be identical, and they are not.
The characteristic function of a random variable \(X\) is
\[ \phi_X(t) = E\left(e^{itX}\right) = E\big(\cos tX\big) + i\,E\big(\sin tX\big), \qquad t \in \mathbb{R}. \]The moment generating function \(M_X(t) = E(e^{tX})\), met in Theory of Probability Unit 5, can fail to exist — the Cauchy distribution has no MGF for any \(t \ne 0\). The characteristic function always exists, because \(|e^{itX}| = 1\) so the expectation is of a bounded function. That is the entire reason the limit theorems of Unit 4 are proved with \(\phi\) and not with \(M\).
Property (1) follows from \(\phi_X(0) = E(e^{0}) = E(1) = 1\) and \(|\phi_X(t)| = |E(e^{itX})| \le E|e^{itX}| = E(1) = 1\), using \(|E(Z)| \le E|Z|\). Property (5) is the workhorse: the characteristic function turns the convolution of two distributions into an ordinary product.
Given. \(X \sim\) Poisson\((\lambda)\), so \(P(X = k) = e^{-\lambda} \lambda^{k} / k!\) for \(k = 0, 1, 2, \ldots\)
Asked. Derive \(\phi_X(t)\) and use it to obtain \(E(X)\) and \(\operatorname{Var}(X)\).
Step 1 — write the expectation as a series.
\[ \phi_X(t) = E\left(e^{itX}\right) = \sum_{k=0}^{\infty} e^{itk} \, \frac{e^{-\lambda}\lambda^{k}}{k!}. \]Step 2 — pull out what does not depend on \(k\) and gather the powers.
\[ = e^{-\lambda} \sum_{k=0}^{\infty} \frac{\left(\lambda e^{it}\right)^{k}}{k!}. \]Step 3 — recognise the exponential series. For any complex \(z\), \(\sum_{k \ge 0} z^{k}/k! = e^{z}\). With \(z = \lambda e^{it}\),
\[ \phi_X(t) = e^{-\lambda} \, e^{\lambda e^{it}} = \exp\left\{\lambda\left(e^{it} - 1\right)\right\}. \]Step 4 — first derivative at 0. Differentiating by the chain rule, with \(\tfrac{d}{dt} e^{it} = i e^{it}\),
\[ \phi_X'(t) = \exp\left\{\lambda\left(e^{it} - 1\right)\right\} \cdot \lambda \, i e^{it}. \]At \(t = 0\): \(e^{i \cdot 0} = 1\), so the exponential factor is \(e^{0} = 1\) and
\[ \phi_X'(0) = 1 \cdot \lambda i = i\lambda. \]By property (6) with \(k = 1\), \(\phi_X'(0) = i E(X)\), so \(E(X) = \lambda\).
Step 5 — second derivative at 0. Differentiate the product in Step 4:
\[ \phi_X''(t) = \exp\left\{\lambda\left(e^{it} - 1\right)\right\} \left(\lambda i e^{it}\right)^{2} + \exp\left\{\lambda\left(e^{it} - 1\right)\right\} \cdot \lambda \, i^{2} e^{it}. \]At \(t = 0\), using \(i^{2} = -1\),
\[ \phi_X''(0) = (\lambda i)^{2} + \lambda i^{2} = -\lambda^{2} - \lambda. \]By property (6) with \(k = 2\), \(\phi_X''(0) = i^{2} E(X^{2}) = -E(X^{2})\), so \(E(X^{2}) = \lambda^{2} + \lambda\).
Step 6 — the variance.
\[ \operatorname{Var}(X) = E(X^{2}) - [E(X)]^{2} = \left(\lambda^{2} + \lambda\right) - \lambda^{2} = \lambda. \]Interpretation. Mean and variance are both \(\lambda\), the result quoted without proof in Discrete Distributions, Unit 2. Property (5) now gives the additivity of the Poisson in one line: if \(X \sim\) Poisson\((\lambda_1)\) and \(Y \sim\) Poisson\((\lambda_2)\) are independent, then \(\phi_{X+Y}(t) = e^{\lambda_1(e^{it}-1)} e^{\lambda_2(e^{it}-1)} = e^{(\lambda_1+\lambda_2)(e^{it}-1)}\), which is the characteristic function of Poisson\((\lambda_1 + \lambda_2)\). By the uniqueness theorem below, that is the distribution of the sum.
Uniqueness. Two random variables have the same distribution function if and only if they have the same characteristic function. So \(\phi\) determines \(F\) completely, and the step taken at the end of Example 2.2 — recognising a characteristic function and naming the distribution — is valid.
Inversion (continuous case). If \(\phi_X\) is absolutely integrable, \(X\) has a bounded continuous density given by
\[ f(x) = \frac{1}{2\pi} \int_{-\infty}^{\infty} e^{-itx} \phi_X(t) \, dt. \]Inversion (lattice case). If \(X\) takes only integer values,
\[ P(X = k) = \frac{1}{2\pi} \int_{-\pi}^{\pi} e^{-itk} \phi_X(t) \, dt. \]Continuity theorem (Lévy). Let \(F_n\) have characteristic functions \(\phi_n\). Then \(F_n \to F\) at every continuity point of \(F\) if and only if \(\phi_n(t) \to \phi(t)\) for every \(t\), where \(\phi\) is continuous at \(t = 0\); and then \(\phi\) is the characteristic function of \(F\).
The continuity theorem is the engine of Unit 4. Every central limit theorem there is proved by showing that a sequence of characteristic functions converges to \(e^{-t^{2}/2}\), the characteristic function of the standard normal, and then invoking this theorem to conclude that the distributions converge.
Given. Four candidate functions of a real variable \(t\):
\[ \text{(a) } \cos t, \qquad \text{(b) } \frac{1}{1 + t^{2}}, \qquad \text{(c) } t, \qquad \text{(d) } 2 e^{-t^{2}}. \]Asked. Decide which are characteristic functions, giving the distribution where one exists and the failing property where one does not.
(a) \(\cos t\) — yes. Let \(X\) take the values \(+1\) and \(-1\), each with probability \(\tfrac12\). Then
\[ \phi_X(t) = \tfrac12 e^{it(1)} + \tfrac12 e^{it(-1)} = \tfrac12\left(e^{it} + e^{-it}\right) = \cos t, \]the last step by Euler's identity \(e^{i\theta} + e^{-i\theta} = 2\cos\theta\). It is real for every \(t\), consistent with property (7) and with the fact that this \(X\) is symmetric about \(0\).
(b) \(1/(1 + t^{2})\) — yes. It is the characteristic function of the standard Laplace (double exponential) distribution with density \(f(x) = \tfrac12 e^{-|x|}\) on \(\mathbb{R}\). Two quick consistency checks: at \(t = 0\) it equals \(1\), as property (1) demands; and it is real and even, matching the symmetry of the Laplace density about \(0\).
(c) \(t\) — no. Property (1) requires \(\phi(0) = 1\), but here \(\phi(0) = 0\). A second, independent failure: \(|\phi(t)| \le 1\) is required, yet \(|t| \to \infty\). Either failure alone settles it.
(d) \(2e^{-t^{2}}\) — no. At \(t = 0\) it equals \(2 \ne 1\), so property (1) fails. Note that the unscaled \(e^{-t^{2}}\) is a characteristic function — that of \(N(0, 2)\), since the \(N(0,\sigma^{2})\) characteristic function is \(e^{-\sigma^{2}t^{2}/2}\) and \(\sigma^{2} = 2\) gives \(e^{-t^{2}}\). Multiplying by a constant destroys the normalisation and nothing else.
Interpretation. \(\phi(0) = 1\) and \(|\phi| \le 1\) are necessary but not sufficient. The full characterisation is Bochner's theorem: a continuous \(\phi\) with \(\phi(0) = 1\) is a characteristic function if and only if it is non-negative definite. In an examination, the quick necessary conditions dispose of most candidates, and a known distribution identifies the rest.
Statement. For \(a > 0\) and \(r > 0\), \(P(|X| \ge a) \le E|X|^{r} / a^{r}\).
Step 1 — name the event and its indicator. Let \(A = \{\omega : |X(\omega)| \ge a\}\) and let \(\mathbf{1}_A\) be its indicator, which is \(1\) on \(A\) and \(0\) off it.
Step 2 — a pointwise inequality. At every \(\omega\),
\[ |X(\omega)|^{r} \;\ge\; a^{r} \, \mathbf{1}_A(\omega). \]Justification, by cases. On \(A\) we have \(|X| \ge a \ge 0\), and \(u \mapsto u^{r}\) is increasing on \([0,\infty)\), so \(|X|^{r} \ge a^{r}\), which is the right-hand side there. Off \(A\) the right-hand side is \(0\) and the left-hand side is non-negative. So the inequality holds everywhere.
Step 3 — take expectations. Expectation is monotone: if \(U \ge V\) pointwise then \(E(U) \ge E(V)\). Applying this,
\[ E|X|^{r} \;\ge\; a^{r} \, E\big(\mathbf{1}_A\big) = a^{r} \, P(A), \]the last equality because the expectation of an indicator is the probability of its event.
Step 4 — divide. Since \(a^{r} > 0\), dividing preserves the direction:
\[ P\big(|X| \ge a\big) \le \frac{E|X|^{r}}{a^{r}}. \quad \blacksquare \]Chebyshev. Apply Markov to the variable \(X - \mu\) with \(r = 2\) and \(a = k\):
\[ P\big(|X - \mu| \ge k\big) \le \frac{E|X - \mu|^{2}}{k^{2}} = \frac{\sigma^{2}}{k^{2}}. \]So Chebyshev is not an independent result; it is Markov at \(r = 2\), centred at the mean.
Given. \(X \sim\) Poisson\((\lambda = 4)\), for which \(\mu = 4\) and \(\sigma^{2} = 4\) by Example 2.2.
Asked. Bound \(P(|X - 4| \ge 4)\) by Chebyshev, compute the exact value, and compare.
Step 1 — the bound. With \(k = 4\),
\[ P\big(|X - 4| \ge 4\big) \le \frac{\sigma^{2}}{k^{2}} = \frac{4}{16} = 0.25. \]Step 2 — identify the event exactly. \(|X - 4| \ge 4\) means \(X \le 0\) or \(X \ge 8\). Since \(X\) takes only non-negative integers, \(X \le 0\) is the single outcome \(X = 0\). So
\[ P\big(|X - 4| \ge 4\big) = P(X = 0) + P(X \ge 8). \]Step 3 — the first term.
\[ P(X = 0) = \frac{e^{-4} 4^{0}}{0!} = e^{-4} = 0.0183156. \]Step 4 — the second term, as one minus a partial sum.
\[ P(X \ge 8) = 1 - \sum_{k=0}^{7} \frac{e^{-4} 4^{k}}{k!} = 0.0511336. \]Step 5 — add.
\[ P\big(|X - 4| \ge 4\big) = 0.0183156 + 0.0511336 = 0.0694493. \]Step 6 — compare. The bound is \(0.25\), the truth is \(0.0694\). The bound is larger by a factor of
\[ \frac{0.25}{0.0694493} = 3.600. \]Interpretation. Chebyshev is correct but weak — here it overstates the tail probability roughly three-and-a-half-fold. That is the price of using only the variance and nothing else about the shape. Its value is precisely that generality: it needs no distributional assumption at all, which is why it, and not any sharper bound, is what proves the weak law of large numbers in Unit 4.
Given. \(X\) takes the values \(1, 2, 6\) with probabilities \(\tfrac12, \tfrac13, \tfrac16\). (Check: \(\tfrac12 + \tfrac13 + \tfrac16 = \tfrac{3+2+1}{6} = 1\).)
Asked. (i) Verify Liapunov's inequality for \(r = 1, 2, 3, 4\); (ii) verify Jensen's inequality for the convex function \(g(u) = 1/u\).
Step 1 — the four absolute moments.
\[ E|X|^{1} = 1\left(\tfrac12\right) + 2\left(\tfrac13\right) + 6\left(\tfrac16\right) = \tfrac12 + \tfrac23 + 1 = \tfrac{3 + 4 + 6}{6} = \tfrac{13}{6} \] \[ E|X|^{2} = 1\left(\tfrac12\right) + 4\left(\tfrac13\right) + 36\left(\tfrac16\right) = \tfrac12 + \tfrac43 + 6 = \tfrac{3 + 8 + 36}{6} = \tfrac{47}{6} \] \[ E|X|^{3} = 1\left(\tfrac12\right) + 8\left(\tfrac13\right) + 216\left(\tfrac16\right) = \tfrac12 + \tfrac83 + 36 = \tfrac{3 + 16 + 216}{6} = \tfrac{235}{6} \] \[ E|X|^{4} = 1\left(\tfrac12\right) + 16\left(\tfrac13\right) + 1296\left(\tfrac16\right) = \tfrac12 + \tfrac{16}{3} + 216 = \tfrac{3 + 32 + 1296}{6} = \tfrac{1331}{6} \]Step 2 — take the \(r\)-th roots.
| \(r\) | \(E|X|^{r}\) | decimal | \(\left(E|X|^{r}\right)^{1/r}\) |
|---|---|---|---|
| 1 | \(13/6\) | 2.166667 | 2.166667 |
| 2 | \(47/6\) | 7.833333 | 2.798809 |
| 3 | \(235/6\) | 39.166667 | 3.396035 |
| 4 | \(1331/6\) | 221.833333 | 3.859284 |
Step 3 — read off Liapunov. The last column increases at every step:
\[ 2.166667 < 2.798809 < 3.396035 < 3.859284, \]which is exactly \(\left(E|X|^{r}\right)^{1/r} \le \left(E|X|^{s}\right)^{1/s}\) for \(r < s\). \(\checkmark\)
Step 4 — Jensen with \(g(u) = 1/u\). This \(g\) is convex on \((0, \infty)\), since \(g''(u) = 2/u^{3} > 0\) there. Compute both sides:
\[ E\left(\frac{1}{X}\right) = \frac{1}{1}\left(\tfrac12\right) + \frac{1}{2}\left(\tfrac13\right) + \frac{1}{6}\left(\tfrac16\right) = \tfrac12 + \tfrac16 + \tfrac1{36} = \tfrac{18 + 6 + 1}{36} = \tfrac{25}{36} = 0.694444, \] \[ \frac{1}{E(X)} = \frac{1}{13/6} = \frac{6}{13} = 0.461538. \]Step 5 — compare. \(0.694444 \ge 0.461538\), so \(E(1/X) \ge 1/E(X)\) as Jensen requires. \(\checkmark\)
Interpretation. The gap in Step 5 is not small: the expected reciprocal is about \(1.50\) times the reciprocal of the expectation. This is the reason a ratio estimator is biased, the point taken up in Sampling Theory (STS-204), Unit 2: averaging and inverting do not commute, and Jensen tells you in advance which way the bias runs.