Skip to the content

Topics Covered

Expectation Moments via E Covariance Uncorrelated vs Independent Addition Theorem Multiplication Theorem Properties of E, Var, Cov Chebyshev's Inequality β₂ ≥ β₁ + 1 Cauchy-Schwarz Inequality
On this page
  1. 1. Mathematical Expectation
  2. 2. Moments via Expectation
  3. 3. Covariance via Expectation
  4. 4. Addition Theorem of Expectation
  5. 5. Multiplication Theorem of Expectation
  6. 6. Properties
  7. 7. Chebyshev's Inequality
  8. 8. Cauchy–Schwarz Inequality
  9. Worked Problems on Mathematical Expectation
  10. Key Take-aways from Unit 4

1. Mathematical Expectation

DEFINITION

The mathematical expectation (or expected value, mean) of a random variable \(X\) is the long-run average value of \(X\) over many repetitions of the experiment.

FORMULA \[ E(X) \;=\; \begin{cases} \displaystyle \sum_x x\, p(x), & \text{discrete}\\[6pt] \displaystyle \int_{-\infty}^{\infty} x\, f(x)\,dx, & \text{continuous} \end{cases} \]

Provided the sum / integral converges absolutely.

Expectation of a Function

\[ E[g(X)] \;=\; \begin{cases} \displaystyle \sum_x g(x)\, p(x), & \text{discrete}\\[6pt] \displaystyle \int_{-\infty}^{\infty} g(x)\, f(x)\,dx, & \text{continuous} \end{cases} \]
EXAMPLE 1 (Discrete)

\(X\) has PMF \(p(0)=0.2,\; p(1)=0.5,\; p(2)=0.3\).
\(E(X) = 0(0.2) + 1(0.5) + 2(0.3) = 1.1\).
\(E(X^2) = 0 + 1(0.5) + 4(0.3) = 1.7\).

EXAMPLE 2 (Continuous)

For \(f(x) = 2x,\; 0 \le x \le 1\):
\(E(X) = \int_0^1 x \cdot 2x \,dx = 2/3\).
\(E(X^2) = \int_0^1 x^2 \cdot 2x\,dx = 1/2\).
Variance = \(1/2 - (2/3)^2 = 1/2 - 4/9 = 1/18\).

Conditional Expectation and Conditional Variance

The conditional expectation of \(Y\) given \(X = x\) is the mean of the conditional distribution:

\[ E(Y \mid X = x) = \sum_y y\,P(Y = y \mid X = x) \quad\text{or}\quad \int y\, f_{Y\mid X}(y\mid x)\,dy. \]

Regarded as a function of the random variable \(X\), \(E(Y\mid X)\) is itself a random variable. Two fundamental identities follow:

\[ \textbf{Tower / total-expectation: } \; E(Y) = E\big[E(Y \mid X)\big]; \] \[ \textbf{Variance decomposition: } \; \text{Var}(Y) = E\big[\text{Var}(Y\mid X)\big] + \text{Var}\big[E(Y\mid X)\big]. \]

The second (the "law of total variance") splits total variability into the average within-group variance plus the variance of the group means — the same idea that underlies ANOVA.

EXAMPLE

Joint pmf of \((X,Y)\): \(P(0,0)=0.10\), \(P(0,1)=0.20\), \(P(0,2)=0.10\), \(P(1,0)=0.20\), \(P(1,1)=0.30\), \(P(1,2)=0.10\). Then \(P(X=0)=0.4,\ P(X=1)=0.6\), and

\(E(Y\mid X=0) = (0\cdot0.10 + 1\cdot0.20 + 2\cdot0.10)/0.4 = 1.0\); \(E(Y\mid X=1) = (0\cdot0.20 + 1\cdot0.30 + 2\cdot0.10)/0.6 = 0.833\).

Check the tower property: \(E[E(Y\mid X)] = 0.4(1.0) + 0.6(0.833) = 0.90 = E(Y)\). ✓

When the Expectation Exists

ABSOLUTE CONVERGENCE

The sum \(\sum x\,p(x)\) or the integral \(\int x f(x)\,dx\) defines \(E(X)\) only when it converges absolutely, that is, when

\[ \sum_x |x|\,p(x) < \infty \qquad\text{or}\qquad \int_{-\infty}^{\infty} |x|\,f(x)\,dx < \infty . \]

The reason is that an expectation must not depend on the order in which the values are listed. A series that converges but not absolutely can be rearranged to give any total at all (Riemann's rearrangement theorem), so it cannot define a mean.

An example with no mean. \(f(x) = 1/x^2\) for \(x \ge 1\) is a density, since \(\int_1^\infty x^{-2}\,dx = \big[-x^{-1}\big]_1^\infty = 1\). But

\[ \int_1^\infty x \cdot \frac{1}{x^2}\,dx = \int_1^\infty \frac{dx}{x} = \big[\ln x\big]_1^\infty = \infty , \]

so \(E(X)\) does not exist. A distribution can be perfectly proper and still have no mean.

THE STATISTICAL AVERAGES ARE ALL EXPECTATIONS

Choosing \(g\) in \(E[g(X)]\) gives every average used in this course:

The expectation of a constant. If \(X = C\) with probability 1, then \(E(C) = C \cdot 1 = C\); in the continuous form, \(E(C) = \int C f(x)\,dx = C\int f(x)\,dx = C\).

2. Moments via Expectation

RAW MOMENT \[ \mu'_r \;=\; E(X^r) \]
CENTRAL MOMENT \[ \mu_r \;=\; E[(X - \mu)^r] \]

The variance, skewness and kurtosis are special moments:

Central Moments from Raw Moments

DERIVATION OF \(\mu_2 = \mu_2' - \mu_1'^{\,2}\)

Write \(\mu = \mu_1' = E(X)\). Expand the square, then use the linearity of expectation (\(\mu\) is a constant):

\[ \begin{aligned} \mu_2 = E(X - \mu)^2 &= E\big(X^2 - 2\mu X + \mu^2\big) \\ &= E(X^2) - 2\mu\,E(X) + \mu^2 \\ &= \mu_2' - 2\mu_1'\mu_1' + \mu_1'^{\,2} = \mu_2' - \mu_1'^{\,2} . \end{aligned} \]

The same steps with \((X - \mu)^3\) and \((X - \mu)^4\), expanded by the binomial theorem, give

\[ \mu_3 = \mu_3' - 3\mu_2'\mu_1' + 2\mu_1'^{\,3}, \] \[ \mu_4 = \mu_4' - 4\mu_3'\mu_1' + 6\mu_2'\mu_1'^{\,2} - 3\mu_1'^{\,4} . \]

For \(\mu_3\): \(E(X - \mu)^3 = \mu_3' - 3\mu\,\mu_2' + 3\mu^2\mu_1' - \mu^3\), and with \(\mu = \mu_1'\) the last two terms combine to \(3\mu_1'^{\,3} - \mu_1'^{\,3} = 2\mu_1'^{\,3}\). These formulae are used again in Unit 5, where the cumulants are expressed through the moments.

3. Covariance via Expectation

DEFINITION

The covariance of two random variables \(X\) and \(Y\) measures the joint variability:

\[ \text{Cov}(X,Y) \;=\; E[(X - E(X))(Y - E(Y))] \;=\; E(XY) - E(X) E(Y) \]

The correlation coefficient is the standardized covariance:

\[ \rho_{XY} \;=\; \dfrac{\text{Cov}(X,Y)}{\sigma_X\, \sigma_Y}, \qquad -1 \le \rho \le 1. \]
EXAMPLE 1

From Unit 3 Sec 2 Ex 1: \(E(X) = 0(0.4)+1(0.6)=0.6\); \(E(Y) = 1(0.3)+2(0.5)+3(0.2)=1.9\).

The three cells with \(X = 0\) contribute 0, so

\[ E(XY) = 1\cdot 1\cdot 0.20 + 1\cdot 2\cdot 0.30 + 1\cdot 3\cdot 0.10 = 0.20 + 0.60 + 0.30 = 1.10 . \]

Cov = \(1.10 - 0.6 \times 1.9 = 1.10 - 1.14 = -0.04\) (slight negative association).

EXAMPLE 2 (Independent ⇒ Cov = 0)

If \(X,Y\) independent, then \(E(XY) = E(X)E(Y)\), hence Cov\(=0\). The converse is not always true.

Uncorrelated Is Not the Same as Independent

THE SHORTCUT FORMULA, STEP BY STEP

Write \(\mu_X = E(X)\) and \(\mu_Y = E(Y)\), both constants. Multiply out and take expectations term by term:

\[ \begin{aligned} \operatorname{Cov}(X,Y) &= E\big[(X - \mu_X)(Y - \mu_Y)\big] \\ &= E\big(XY - \mu_Y X - \mu_X Y + \mu_X\mu_Y\big) \\ &= E(XY) - \mu_Y\mu_X - \mu_X\mu_Y + \mu_X\mu_Y = E(XY) - E(X)E(Y) . \end{aligned} \]
CORRECTION — WHAT \(E(XY) = E(X)E(Y)\) MEANS

The textbook says that \(X\) and \(Y\) “are said to be independent if \(E(XY) = E(X)E(Y)\), i.e. \(\operatorname{Cov}(X,Y) = 0\)”. That condition defines uncorrelated variables, not independent ones. Independence means that the joint distribution factorises, \(P(X = x, Y = y) = P(X = x)\,P(Y = y)\) for every pair. Independence implies zero covariance (§5); the converse is false.

Counterexample. Let \(X\) take the values \(-1, 0, 1\) with probability \(\tfrac13\) each, and let \(Y = X^2\). Then

\[ E(X) = 0, \qquad E(Y) = \tfrac23, \qquad E(XY) = E(X^3) = \tfrac13(-1 + 0 + 1) = 0, \] \[ \operatorname{Cov}(X,Y) = 0 - 0 \cdot \tfrac23 = 0 . \]

Yet \(Y\) is a function of \(X\), as dependent as two variables can be: \(P(X = 0, Y = 0) = \tfrac13\), while \(P(X = 0)\,P(Y = 0) = \tfrac13 \cdot \tfrac13 = \tfrac19\). Covariance measures only linear association, and the relation here is a parabola.

4. Addition Theorem of Expectation

STATEMENT

For any random variables \(X\) and \(Y\) (independent or not),

\[ E(X + Y) \;=\; E(X) + E(Y) \]

More generally, for constants \(a_1, \ldots, a_n\) and r.v.s \(X_1, \ldots, X_n\):

\[ E(a_1 X_1 + a_2 X_2 + \cdots + a_n X_n) \;=\; a_1 E(X_1) + \cdots + a_n E(X_n) \]
PROOF (continuous case)

Let \((X,Y)\) have joint density \(f(x,y)\) with marginals \(f_X\) and \(f_Y\). By the expectation of a function of \((X,Y)\),

\[ \begin{aligned} E(X + Y) &= \int_{-\infty}^{\infty}\!\!\int_{-\infty}^{\infty} (x + y)\, f(x,y)\, dx\, dy \\ &= \iint x\, f(x,y)\, dx\, dy \;+\; \iint y\, f(x,y)\, dx\, dy \\ &= \int_{-\infty}^{\infty} x \Big[\int_{-\infty}^{\infty} f(x,y)\, dy\Big] dx \;+\; \int_{-\infty}^{\infty} y \Big[\int_{-\infty}^{\infty} f(x,y)\, dx\Big] dy \\ &= \int_{-\infty}^{\infty} x\, f_X(x)\, dx \;+\; \int_{-\infty}^{\infty} y\, f_Y(y)\, dy \;=\; E(X) + E(Y). \end{aligned} \]

No independence is assumed — linearity of expectation holds for any \(X, Y\). The discrete case is identical with sums replacing integrals.

EXAMPLE 1

\(X\) = score on first die, \(Y\) = score on second die. \(E(X) = E(Y) = 3.5\). So \(E(X+Y) = 7\).

EXAMPLE 2

\(X\) ~ Binomial(\(n,p\)) is the sum of \(n\) independent Bernoulli(\(p\)) trials. Each Bernoulli has mean \(p\); by linearity \(E(X) = np\).

The Addition Theorem for \(n\) Variables

PROOF BY INDUCTION

Claim: \(E(X_1 + X_2 + \cdots + X_n) = E(X_1) + E(X_2) + \cdots + E(X_n)\) for every \(n\).

  1. Base. For \(n = 2\) this is the theorem just proved.
  2. Step. Suppose it holds for \(r\) variables. Put \(S_r = X_1 + \cdots + X_r\), a single random variable. By the two-variable theorem and then the hypothesis, \[ E(S_r + X_{r+1}) = E(S_r) + E(X_{r+1}) = E(X_1) + \cdots + E(X_r) + E(X_{r+1}) . \]
  3. So it holds for \(r + 1\), and by induction for every \(n\).

Nothing about independence is used at any step.

5. Multiplication Theorem of Expectation

STATEMENT (for independent r.v.)

If \(X\) and \(Y\) are independent random variables, then

\[ E(XY) \;=\; E(X) \cdot E(Y). \]

For \(n\) mutually independent r.v.: \(E(X_1 X_2 \cdots X_n) = \prod E(X_i)\).

EXAMPLE 1

Toss two independent dice. \(E(XY) = E(X) E(Y) = 3.5 \times 3.5 = 12.25\).

EXAMPLE 2

If two independent stocks have expected returns 8 % and 10 %, the expected return of their product (rare in practice) factorizes: \(E(R_1 R_2) = 0.08 \times 0.10 = 0.008\).

Proof of the Multiplication Theorem

PROOF (discrete case)

Let \(X\) take the values \(x_i\) and \(Y\) the values \(y_j\). Independence means \(P(X = x_i, Y = y_j) = p_i\,q_j\) with \(p_i = P(X = x_i)\), \(q_j = P(Y = y_j)\). Then

\[ \begin{aligned} E(XY) &= \sum_i\sum_j x_i y_j\,P(X = x_i, Y = y_j) = \sum_i\sum_j x_i y_j\,p_i q_j \\ &= \Big(\sum_i x_i p_i\Big)\Big(\sum_j y_j q_j\Big) = E(X)\,E(Y) . \end{aligned} \]

The double sum splits into a product only because the joint probability is a product. In the continuous case \(f(x,y) = f_X(x) f_Y(y)\) splits the double integral the same way.

FOR \(n\) VARIABLES: MUTUAL INDEPENDENCE IS NEEDED

The induction runs as for the addition theorem: write \(P_r = X_1 X_2 \cdots X_r\) and use \(E(P_r X_{r+1}) = E(P_r)E(X_{r+1})\). That step needs \(P_r\) to be independent of \(X_{r+1}\), which is what mutual independence guarantees. Pairwise independence is not enough.

Example. Let \(X\) and \(Y\) be independent, each \(\pm1\) with probability \(\tfrac12\), and \(Z = XY\). Each pair among \(X, Y, Z\) is independent (every pair of signs has probability \(\tfrac14\)), and \(E(X) = E(Y) = E(Z) = 0\). But

\[ E(XYZ) = E(X^2Y^2) = 1 \ne 0 = E(X)E(Y)E(Z) . \]

6. Properties

6.1 Properties of Expectation

  1. \(E(c) = c\) for any constant \(c\).
  2. \(E(cX) = c\, E(X)\).
  3. \(E(aX + b) = a\, E(X) + b\).
  4. \(E(X + Y) = E(X) + E(Y)\) (always; linearity).
  5. If \(X \ge 0\) then \(E(X) \ge 0\).
  6. If \(X \le Y\) then \(E(X) \le E(Y)\).
  7. \(|E(X)| \le E(|X|)\).

6.2 Properties of Variance

  1. \(\text{Var}(c) = 0\).
  2. \(\text{Var}(aX + b) = a^2\, \text{Var}(X)\).
  3. \(\text{Var}(X) = E(X^2) - [E(X)]^2 \ge 0\).
  4. \(\text{Var}(X \pm Y) = \text{Var}(X) + \text{Var}(Y) \pm 2\,\text{Cov}(X,Y)\).
  5. If \(X,Y\) independent: \(\text{Var}(X + Y) = \text{Var}(X) + \text{Var}(Y)\).

6.3 Properties of Covariance

  1. Cov\((X, X) = \text{Var}(X)\).
  2. Cov\((X, Y) = \) Cov\((Y, X)\) (symmetric).
  3. Cov\((aX + b,\; cY + d) = ac\,\)Cov\((X, Y)\).
  4. Cov\((X + Y, Z) = \) Cov\((X, Z) + \) Cov\((Y, Z)\) (bilinear).
  5. If \(X, Y\) independent ⇒ Cov\(=0\); converse not generally true.
EXAMPLE 1

If Var\((X) = 4\), Var\((Y)=9\), Cov\((X,Y) = 2\). Then Var\((X + Y) = 4 + 9 + 2(2) = 17\).

EXAMPLE 2

For \(Z = 3X - 2Y\) with the above: Var\((Z) = 9(4) + 4(9) - 2(3)(2)(2) = 36 + 36 - 24 = 48\).

Why the Properties Hold

EXPECTATION

\(E(aX + b) = aE(X) + b\). In the continuous case,

\[ E(aX + b) = \int (ax + b) f(x)\,dx = a\int x f(x)\,dx + b\int f(x)\,dx = aE(X) + b . \]

\(|E(X)| \le E|X|\). Since \(-|x| \le x \le |x|\) for every \(x\), taking expectations (which keeps inequalities, property 6) gives \(-E|X| \le E(X) \le E|X|\).

VARIANCE

\(\operatorname{Var}(aX \pm bY) = a^2\operatorname{Var}(X) + b^2\operatorname{Var}(Y) \pm 2ab\operatorname{Cov}(X,Y)\). Let \(U = aX \pm bY\), so \(U - E(U) = a(X - \mu_X) \pm b(Y - \mu_Y)\). Square and take expectations:

\[ \begin{aligned} \operatorname{Var}(U) &= E\big[a^2(X - \mu_X)^2 + b^2(Y - \mu_Y)^2 \pm 2ab(X - \mu_X)(Y - \mu_Y)\big] \\ &= a^2\operatorname{Var}(X) + b^2\operatorname{Var}(Y) \pm 2ab\operatorname{Cov}(X,Y) . \end{aligned} \]

With \(a = b = 1\) this is property 4, and for independent \(X, Y\) the covariance term vanishes.

Change of origin and scale. If \(U = (X - a)/h\), then \(U - E(U) = (X - \mu_X)/h\), so

\[ \operatorname{Var}(U) = \frac{\operatorname{Var}(X)}{h^2} . \]

Variance is unaffected by the change of origin \(a\) but is divided by the square of the scale \(h\).

COVARIANCE

A change of origin does not alter covariance. \((X + a) - E(X + a) = X - \mu_X\), and likewise for \(Y + b\), so \(\operatorname{Cov}(X + a, Y + b) = \operatorname{Cov}(X,Y)\).

Bilinearity. Expanding the product of deviations as for the variance,

\[ \operatorname{Cov}(aX - bY,\; cX - dY) = ac\operatorname{Var}(X) + bd\operatorname{Var}(Y) - (ad + bc)\operatorname{Cov}(X,Y), \]

because the deviations are \(a(X - \mu_X) - b(Y - \mu_Y)\) and \(c(X - \mu_X) - d(Y - \mu_Y)\), and their product has expectation \(ac\operatorname{Var}(X) - ad\operatorname{Cov} - bc\operatorname{Cov} + bd\operatorname{Var}(Y)\).

7. Chebyshev's Inequality

STATEMENT

If \(X\) is any r.v. with mean \(\mu\) and finite variance \(\sigma^2\), then for any \(k > 0\):

\[ P(|X - \mu| \ge k\sigma) \;\le\; \dfrac{1}{k^2} \]

Equivalently: \(P(|X - \mu| < k\sigma) \ge 1 - \dfrac{1}{k^2}\).

This is a distribution-free bound — it holds for any distribution with finite variance.

μ − kσμμ + kσ P(|X−μ| < kσ) ≥ 1 − 1/k² tails tails both tails combined ≤ 1/k² (e.g. k = 2 ⇒ ≤ 25 %, so centre ≥ 75 %)
Fig 4.1 — Chebyshev caps the probability in the two tails beyond \(k\) standard deviations at \(1/k^2\), whatever the shape of the distribution — so at least \(1 - 1/k^2\) of the mass must sit in the central band. The bound is loose (a normal curve keeps ~95 % within \(\pm 2\sigma\), not just 75 %) but it needs no distributional assumption.
EXAMPLE 1

For \(k = 2\): \(P(|X - \mu| \ge 2\sigma) \le 1/4\). At least 75 % of any distribution lies within 2 SDs of mean.

EXAMPLE 2

Let \(X\) have \(\mu = 50,\; \sigma = 4\). Find a lower bound for \(P(40 < X < 60)\).
\(|X - 50| < 10\) means \(k\sigma = 10 \Rightarrow k = 2.5\).
\(P(|X - 50| < 10) \ge 1 - 1/6.25 = 0.84\). At least 84 % of values lie in (40, 60).

Applications of Chebyshev

Markov's Inequality

For a non-negative random variable \(X\) with finite mean and any \(a > 0\),

\[ P(X \ge a) \le \dfrac{E(X)}{a}. \]

Markov's inequality is the one-sided ancestor of Chebyshev's inequality: applying it to \((X-\mu)^2\) with \(a = k^2\sigma^2\) immediately yields Chebyshev's \(P(|X-\mu|\ge k\sigma) \le 1/k^2\).

Proof of Chebyshev's Inequality

PROOF (continuous case)

Let \(X\) have density \(f\), mean \(\mu\) and variance \(\sigma^2\), and let \(k > 0\). Split the variance integral into the central band and the two tails:

\[ \sigma^2 = \int_{-\infty}^{\infty} (x - \mu)^2 f(x)\,dx = \int_{-\infty}^{\mu - k\sigma} + \int_{\mu - k\sigma}^{\mu + k\sigma} + \int_{\mu + k\sigma}^{\infty} . \]

Every piece is non-negative, so dropping the middle one can only decrease the total:

\[ \sigma^2 \ge \int_{-\infty}^{\mu - k\sigma} (x - \mu)^2 f(x)\,dx + \int_{\mu + k\sigma}^{\infty} (x - \mu)^2 f(x)\,dx . \]

In both remaining ranges \(|x - \mu| \ge k\sigma\), so \((x - \mu)^2 \ge k^2\sigma^2\):

\[ \sigma^2 \ge k^2\sigma^2\Big[P(X \le \mu - k\sigma) + P(X \ge \mu + k\sigma)\Big] = k^2\sigma^2\,P\big(|X - \mu| \ge k\sigma\big) . \]

Divide by \(k^2\sigma^2\) to get \(P(|X - \mu| \ge k\sigma) \le 1/k^2\). The discrete proof is the same with sums.

Printing note. The textbook writes the two tails as \(P(X < \mu - k\sigma)\) and \(P(X > \mu + k\sigma)\). To reach \(P(|X - \mu| \ge k\sigma)\) the tails must include their end-points, \(\le\) and \(\ge\), leaving the open interval in the middle. For a continuous variable the end-points carry no probability and the two readings agree; for a discrete one they can, and only the \(\le\)/\(\ge\) form gives the stated result.

THE FORM USED IN PRACTICE

Putting \(c = k\sigma\),

\[ P\big(|X - \mu| \ge c\big) \le \frac{\sigma^2}{c^2}, \qquad P\big(|X - \mu| < c\big) \ge 1 - \frac{\sigma^2}{c^2} . \]

The bound tells us something only if \(\sigma^2/c^2 < 1\), that is \(c > \sigma\) (\(k > 1\)). For a band narrower than one standard deviation, any probability at all is possible in the tails, and the inequality is true but empty. Worked Problem 11 below meets exactly this case.

A Moment Inequality: \(\beta_2 \ge \beta_1 + 1\)

PROOF

Standardise: \(Z = (X - \mu)/\sigma\) has \(E(Z) = 0\), \(E(Z^2) = 1\), \(E(Z^3) = \mu_3/\sigma^3 = \gamma_1\) with \(\gamma_1^2 = \beta_1 = \mu_3^2/\mu_2^3\), and \(E(Z^4) = \mu_4/\mu_2^2 = \beta_2\). For any constants \(t\) and \(K\), a square has non-negative expectation:

\[ \begin{aligned} 0 \le E\big(Z^2 + tZ + K\big)^2 &= E(Z^4) + t^2E(Z^2) + K^2 + 2tE(Z^3) + 2KE(Z^2) + 2tKE(Z) \\ &= \beta_2 + t^2 + 2t\gamma_1 + K^2 + 2K . \end{aligned} \]

The right side is smallest at \(t = -\gamma_1\), where \(t^2 + 2t\gamma_1 = -\gamma_1^2 = -\beta_1\). So \(\beta_2 \ge \beta_1 - (K^2 + 2K)\) for every \(K\), and the strongest choice minimises \(K^2 + 2K\), at \(K = -1\), where it equals \(-1\):

\[ \beta_2 \ge \beta_1 + 1, \qquad\text{and in particular}\qquad \beta_2 \ge 1 . \]

Equality holds only when \(Z^2 - \gamma_1 Z - 1 = 0\) with probability 1, that is, for a distribution on two points. (A Bernoulli variable with \(p = \tfrac14\) has \(\beta_1 = \tfrac43\), \(\beta_2 = \tfrac73\): equality.)

8. Cauchy–Schwarz Inequality

STATEMENT

For any two r.v. \(X, Y\) with finite second moments,

\[ \big[E(XY)\big]^2 \;\le\; E(X^2)\, E(Y^2) \]

Equivalently using deviations:

\[ \big[\text{Cov}(X,Y)\big]^2 \;\le\; \text{Var}(X)\, \text{Var}(Y) \]

This implies the correlation coefficient \(|\rho_{XY}| \le 1\).

EXAMPLE 1

If \(\text{Var}(X) = 16,\; \text{Var}(Y) = 25\), then by Cauchy-Schwarz, \(\text{Cov}(X,Y)^2 \le 400\), i.e., \(|\text{Cov}(X,Y)| \le 20\).

EXAMPLE 2

Equality holds when \(Y = aX + b\) (linear relationship), giving \(\rho = \pm 1\).

Proof of the Cauchy–Schwarz Inequality

PROOF

For every real \(t\), \((X + tY)^2 \ge 0\), so its expectation is non-negative:

\[ Z(t) = E(X + tY)^2 = E(X^2) + 2t\,E(XY) + t^2\,E(Y^2) \ge 0 . \]

Suppose \(E(Y^2) > 0\). Then \(Z(t)\) is a quadratic in \(t\) with positive leading coefficient that is never negative, so it has at most one real root and its discriminant is not positive:

\[ \big[2E(XY)\big]^2 - 4E(X^2)E(Y^2) \le 0 \quad\Longrightarrow\quad \big[E(XY)\big]^2 \le E(X^2)\,E(Y^2) . \]

If \(E(Y^2) = 0\), then \(Y = 0\) with probability 1, \(E(XY) = 0\), and both sides are 0. (The textbook omits this case; the discriminant argument needs \(E(Y^2) > 0\), because otherwise \(Z(t)\) is not a quadratic.)

Equality holds exactly when the discriminant is zero: then \(Z(t_0) = 0\) for some \(t_0\), so \(X + t_0Y = 0\) with probability 1, and \(X\) is a constant multiple of \(Y\). Applied to \(X - \mu_X\) and \(Y - \mu_Y\), the same argument gives the covariance form, with equality when \(Y\) is a linear function \(aX + b\) of \(X\).

Worked Problems on Mathematical Expectation

Fourteen problems in the textbook's order: seven that find an expectation from a stated distribution, one on a linear change of variable, four applications of Chebyshev's inequality and two checks of the Cauchy–Schwarz inequality, followed by the exercises with their answers checked. Every one uses the same two moves: write down the distribution, then weight each value by its probability (a sum) or by its density (an integral).

Source note. Every answer below was recomputed exactly, as a fraction or a closed-form integral, and the Chebyshev bounds were set beside the exact probabilities they bound. Twelve of the fourteen worked answers agree. Worked Problem 5 differs in the third decimal place, a rounding artefact, and is left standing with a note. Worked Problem 11 reaches a “bound” larger than 1, and is corrected. All eight exercise answers are right.

A. Expectation from a Distribution

WORKED PROBLEM 1 — heads in three tosses

Three coins are tossed. Find the expected number of heads.

Let \(X\) = number of heads. Of the 8 equally likely outcomes, 1 has no head, 3 have one (\(HTT, THT, TTH\)), 3 have two and 1 has three:

\(x\)0123
\(p(x)\)\(\tfrac18\)\(\tfrac38\)\(\tfrac38\)\(\tfrac18\)
\[ E(X) = \sum x\,p(x) = 0 \cdot \tfrac18 + 1 \cdot \tfrac38 + 2 \cdot \tfrac38 + 3 \cdot \tfrac18 = \frac{0 + 3 + 6 + 3}{8} = \frac{12}{8} = \frac32 = 1.5 . \]

On average one and a half heads, though no single toss of three coins can show 1.5. An expectation is a long-run average, not a value that must occur.

WORKED PROBLEM 2 — \(E(X)\), \(E(X^2)\) and \(E(2X + 1)^2\)

\(X\) takes the values \(-3, 6, 9\) with probabilities \(\tfrac16, \tfrac12, \tfrac13\). Find \(E(X)\), \(E(X^2)\) and \(E(2X + 1)^2\).

\[ E(X) = -3 \cdot \tfrac16 + 6 \cdot \tfrac12 + 9 \cdot \tfrac13 = -\tfrac12 + 3 + 3 = \frac{11}{2} = 5.5, \] \[ E(X^2) = 9 \cdot \tfrac16 + 36 \cdot \tfrac12 + 81 \cdot \tfrac13 = 1.5 + 18 + 27 = \frac{93}{2} = 46.5 . \]

For the third, expand the square and use linearity rather than squaring each value:

\[ E(2X + 1)^2 = E(4X^2 + 4X + 1) = 4(46.5) + 4(5.5) + 1 = 186 + 22 + 1 = 209 . \]

Check directly: \(2X + 1\) takes \(-5, 13, 19\), and \(25 \cdot \tfrac16 + 169 \cdot \tfrac12 + 361 \cdot \tfrac13 = \tfrac{25 + 507 + 722}{6} = \tfrac{1254}{6} = 209\).

WORKED PROBLEM 3 — the sum of two dice

Two dice are thrown. Find the expected value of the sum of the numbers shown.

Counting the 36 ordered pairs for each sum:

\(x\)234567
\(36\,p(x)\)123456
\(x\)89101112
\(36\,p(x)\)54321
\[ \sum x \cdot 36\,p(x) = 2 + 6 + 12 + 20 + 30 + 42 + 40 + 36 + 30 + 22 + 12 = 252, \] \[ E(X) = \frac{252}{36} = 7 . \]

The addition theorem (§4) gets there in one line: each die has mean \(\tfrac{1 + 2 + \cdots + 6}{6} = 3.5\), so \(E(X_1 + X_2) = 3.5 + 3.5 = 7\).

WORKED PROBLEM 4 — four coins: mean and variance

Four coins are tossed. Find the mean and variance of the number of heads.

\(P(X = x) = \binom4x/16\), giving \(\tfrac1{16}, \tfrac4{16}, \tfrac6{16}, \tfrac4{16}, \tfrac1{16}\) for \(x = 0, 1, 2, 3, 4\).

\[ E(X) = \frac{0 + 4 + 12 + 12 + 4}{16} = \frac{32}{16} = 2, \] \[ E(X^2) = \frac{0 + 4 + 24 + 36 + 16}{16} = \frac{80}{16} = 5, \] \[ \operatorname{Var}(X) = E(X^2) - [E(X)]^2 = 5 - 4 = 1 . \]

These agree with the binomial formulae \(np = 4 \cdot \tfrac12 = 2\) and \(npq = 4 \cdot \tfrac12 \cdot \tfrac12 = 1\).

WORKED PROBLEM 5 — \(E(X)\), \(E(X^2)\) and \(E(X - 1)^2\)

\(X\) takes the values \(0, 1, 2, 3\) with probabilities \(\tfrac13, \tfrac12, \tfrac1{24}, \tfrac18\). Find \(E(X)\), \(E(X^2)\) and \(E(X - 1)^2\).

First check that the probabilities add to 1: \(\tfrac{8 + 12 + 1 + 3}{24} = 1\). In twenty-fourths,

\[ E(X) = \frac{0 + 12 + 2 + 9}{24} = \frac{23}{24}, \qquad E(X^2) = \frac{0 + 12 + 4 + 27}{24} = \frac{43}{24}, \] \[ E(X - 1)^2 = E(X^2) - 2E(X) + 1 = \frac{43 - 46 + 24}{24} = \frac{21}{24} = \frac78 = 0.875 . \]

Directly: \((X - 1)^2\) is \(1, 0, 1, 4\), and \(\tfrac13 + 0 + \tfrac1{24} + \tfrac48 = \tfrac{8 + 1 + 12}{24} = \tfrac{21}{24}\).

Rounding note. The textbook prints \(0.876\). That comes from rounding \(E(X^2) = 1.7917\) and \(2E(X) = 1.9167\) to three places before subtracting (\(1.792 - 1.916 + 1 = 0.876\)). The exact value is \(\tfrac78 = 0.875\).

WORKED PROBLEM 6 — tossing until the first head

A coin is tossed until a head appears. Find the expected number of tosses.

The first head comes on toss \(x\) when the first \(x - 1\) tosses are tails and toss \(x\) is a head: \(P(X = x) = (\tfrac12)^{x-1}\tfrac12 = \tfrac{1}{2^x}\), \(x = 1, 2, 3, \ldots\). So

\[ S = E(X) = \frac12 + \frac{2}{4} + \frac{3}{8} + \frac{4}{16} + \cdots . \]

Halve it and shift one place:

\[ \frac{S}{2} = \frac14 + \frac28 + \frac{3}{16} + \cdots . \]

Subtract term by term: each numerator drops by exactly 1,

\[ S - \frac S2 = \frac12 + \frac14 + \frac18 + \cdots = 1 \quad\Longrightarrow\quad S = 2 . \]

The geometric series converges, and every term is positive, so the rearrangement is legitimate (§1: the series converges absolutely).

WORKED PROBLEM 7 — a continuous expectation

\(X\) has density \(f(x) = 3x^2\), \(0 \le x \le 1\). Find \(E(X)\).

First, \(\int_0^1 3x^2\,dx = [x^3]_0^1 = 1\), so \(f\) is a density. Then

\[ E(X) = \int_0^1 x \cdot 3x^2\,dx = \int_0^1 3x^3\,dx = \Big[\frac{3x^4}{4}\Big]_0^1 = \frac34 . \]

The density rises towards 1, so the mean sits to the right of the midpoint \(\tfrac12\).

B. A Linear Change of Variable

WORKED PROBLEM 8 — standardising \(X\)

\(E(X) = 10\) and \(\operatorname{Var}(X) = 25\). Find positive constants \(a, b\) such that \(Y = aX - b\) has mean 0 and variance 1.

By the properties of §6,

\[ E(Y) = aE(X) - b = 10a - b = 0, \qquad \operatorname{Var}(Y) = a^2\operatorname{Var}(X) = 25a^2 = 1 . \]

The second gives \(a = \tfrac15\) (taking the positive root), and then the first gives \(b = 10a = 2\). So \(Y = \tfrac15X - 2 = \dfrac{X - 10}{5}\): subtract the mean, divide by the standard deviation. That is exactly how any variable is standardised.

C. Chebyshev's Inequality in Use

WORKED PROBLEM 9 — sixes in 720 throws

A fair die is thrown 720 times. Use Chebyshev's inequality to find a lower bound for the probability of getting between 100 and 140 sixes.

\(X\) = number of sixes \(\sim B(n, p)\) with \(n = 720\), \(p = \tfrac16\):

\[ \mu = np = 120, \qquad \sigma^2 = npq = 720 \cdot \tfrac16 \cdot \tfrac56 = 100, \qquad \sigma = 10 . \]

The range 100 to 140 is \(120 \pm 20\), and \(20 = k\sigma\) gives \(k = 2\):

\[ P(100 \le X \le 140) \ge P\big(|X - 120| < 20\big) \ge 1 - \frac{1}{2^2} = \frac34 = 0.75 . \]

The exact binomial probability is \(0.9598\): the bound is true and far from sharp.

Printing note. The textbook writes “\(X \sim B(np, npq)\)”. The binomial is written \(B(n, p)\), with mean \(np\) and variance \(npq\).

WORKED PROBLEM 10 — heads in 2000 tosses

A coin is tossed 2000 times. Show, using Chebyshev's inequality, that the probability of between 900 and 1100 heads is at least \(\tfrac{19}{20}\).

\(X \sim B(2000, \tfrac12)\): \(\mu = 1000\), \(\sigma^2 = 2000 \cdot \tfrac12 \cdot \tfrac12 = 500\), \(\sigma = 22.36\). The range is \(1000 \pm 100\), so \(k = 100/22.36 = 4.47\); exactly, \(k^2 = 100^2/500 = 20\):

\[ P(900 \le X \le 1100) \ge P\big(|X - 1000| < 100\big) \ge 1 - \frac{\sigma^2}{100^2} = 1 - \frac{500}{10000} = \frac{19}{20} . \]

(The exact probability is \(0.99999\).)

Printing note. The textbook's closing sentence says the coin is thrown “200 times”; the problem, and the working, use 2000.

WORKED PROBLEM 11 — when Chebyshev says nothing

\(X\) takes the values \(-1, 1, 3, 5\) with probabilities \(\tfrac16, \tfrac16, \tfrac16, \tfrac12\). Use Chebyshev's inequality to find an upper bound for \(P(|X - 3| \ge 1)\).

\[ \mu = \frac{-1 + 1 + 3}{6} + \frac52 = \frac12 + \frac52 = 3, \qquad E(X^2) = \frac{1 + 1 + 9}{6} + \frac{25}{2} = \frac{43}{3}, \] \[ \sigma^2 = \frac{43}{3} - 9 = \frac{16}{3}, \qquad \sigma = \frac{4}{\sqrt3} = 2.3094 . \]

With \(c = 1\),

\[ P\big(|X - 3| \ge 1\big) \le \frac{\sigma^2}{c^2} = \frac{16}{3} = 5.33 . \]

Correction. The textbook stops here, with 5.33 as the “upper bound”. A probability cannot exceed 1, so a bound of 5.33 carries no information at all. It happens because \(c = 1\) is smaller than \(\sigma = 2.31\) (§7: the inequality is informative only for \(c > \sigma\), i.e. \(k > 1\); here \(k = c/\sigma = \sqrt3/4 = 0.43\)). The honest conclusion is that Chebyshev's inequality gives no useful bound for this band. The exact answer is easy, since only \(X = 3\) lies within 1 of the mean:

\[ P\big(|X - 3| \ge 1\big) = 1 - P(X = 3) = 1 - \frac16 = \frac56 . \]
|X − 3| < 1 1/6 1/6 1/6 1/2 -2 -1 0 1 2 3 4 5 6 μ = 3 x P(|X − 3| ≥ 1) = 1/6 + 1/6 + 1/2 = 5/6
Fig 4.2 — Worked Problem 11. Only the value 3 lies within 1 of the mean, so the exact probability of a deviation of at least 1 is \(5/6\). Chebyshev can say nothing useful here: with \(\sigma = 4/\sqrt3 \approx 2.31\) the band is less than one standard deviation wide, and the bound \(\sigma^2/c^2 = 16/3\) exceeds 1.
WORKED PROBLEM 12 — a bound for an exponential variable

\(X\) has density \(f(x) = e^{-x}\), \(x \ge 0\). Use Chebyshev's inequality to bound \(P(|X - 1| < 2)\), and compare with the exact value.

Integrating by parts, \(\int_0^\infty x e^{-x}\,dx = \big[-xe^{-x}\big]_0^\infty + \int_0^\infty e^{-x}\,dx = 0 + 1\), so \(\mu = 1\). Once more, \(E(X^2) = \int_0^\infty x^2 e^{-x}\,dx = 2\int_0^\infty x e^{-x}\,dx = 2\). Hence

\[ \sigma^2 = 2 - 1^2 = 1, \qquad \sigma = 1 . \]

With \(c = 2\) (\(k = 2\)),

\[ P\big(|X - 1| < 2\big) = P(-1 < X < 3) \ge 1 - \frac{1}{2^2} = \frac34 = 0.75 . \]

Since \(X \ge 0\), the event is \(0 \le X < 3\), with exact probability \(\int_0^3 e^{-x}\,dx = 1 - e^{-3} = 0.9502\). The tail Chebyshev allows is \(0.25\); the real tail is \(e^{-3} = 0.0498\).

c ≤ σ: bound ≥ 1, no information 0 0.25 0.5 0.75 1 0 1 2 3 4 c (deviation from the mean μ = 1; here σ = 1) 0.25 0.0498 at c = 2 (dots): bound 1/4, exact e⁻³ = 0.0498 Chebyshev bound σ²/c² exact P(|X − 1| ≥ c)
Fig 4.3 — Worked Problem 12. For \(f(x) = e^{-x}\) (mean 1, standard deviation 1) the exact chance of a deviation of at least \(c\) sits far below Chebyshev's \(\sigma^2/c^2\). The bound is honest for every distribution with this variance, which is why it is loose for any one of them; and for \(c \le \sigma\) it says nothing at all.

D. The Cauchy–Schwarz Inequality Checked

WORKED PROBLEM 13 — a discrete joint distribution

Verify the Cauchy–Schwarz inequality for the joint distribution

\(Y = 1\)\(Y = 2\)\(p_X\)
\(X = 1\)\(\tfrac18\)\(\tfrac38\)\(\tfrac12\)
\(X = 2\)\(\tfrac14\)\(\tfrac14\)\(\tfrac12\)
\(p_Y\)\(\tfrac38\)\(\tfrac58\)1

From the margins,

\[ E(X^2) = 1 \cdot \tfrac12 + 4 \cdot \tfrac12 = \frac52, \qquad E(Y^2) = 1 \cdot \tfrac38 + 4 \cdot \tfrac58 = \frac{23}{8}, \]

and from the four cells,

\[ E(XY) = 1 \cdot \tfrac18 + 2 \cdot \tfrac38 + 2 \cdot \tfrac14 + 4 \cdot \tfrac14 = \frac{1 + 6 + 4 + 8}{8} = \frac{19}{8} . \] \[ \big[E(XY)\big]^2 = \frac{361}{64} = 5.64 \;\le\; E(X^2)\,E(Y^2) = \frac{115}{16} = 7.19 . \]

The inequality holds, and is strict: \(X\) is not a multiple of \(Y\).

WORKED PROBLEM 14 — a continuous joint density

Verify the Cauchy–Schwarz inequality for \(f(x, y) = e^{-(x + y)}\), \(x, y \ge 0\).

The density factorises, \(e^{-(x+y)} = e^{-x} \cdot e^{-y}\), so the marginals are \(f_X(x) = e^{-x}\) and \(f_Y(y) = e^{-y}\), and \(X, Y\) are independent. From Worked Problem 12, \(E(X^2) = E(Y^2) = 2\). The double integral splits:

\[ E(XY) = \int_0^\infty x e^{-x}\,dx \int_0^\infty y e^{-y}\,dy = \Big[-xe^{-x} - e^{-x}\Big]_0^\infty \Big[-ye^{-y} - e^{-y}\Big]_0^\infty = 1 \cdot 1 = 1 . \] \[ \big[E(XY)\big]^2 = 1 \;<\; E(X^2)\,E(Y^2) = 2 \cdot 2 = 4 . \]

Exercises, with Answers Checked

PRACTICE
  1. Show that \(E(aX - bY) = aE(X) - bE(Y)\) for constants \(a, b\). By the addition theorem and \(E(cX) = cE(X)\).
  2. Show that \(\operatorname{Cov}(aX - bY, cX - dY) = ac\operatorname{Var}(X) + bd\operatorname{Var}(Y) - (ad + bc)\operatorname{Cov}(X,Y)\). Proved in §6, “Why the Properties Hold”.
  3. Show that \(\operatorname{Var}(a + bX) = b^2\operatorname{Var}(X)\). The deviation of \(a + bX\) from its mean is \(b(X - \mu)\).
  4. \(X\) takes \(8, 12, 16, 20, 24\) with probabilities \(\tfrac18, \tfrac16, \tfrac38, \tfrac14, \tfrac1{12}\). Find the mean and variance. Ans. mean 16, variance 20: \(E(X) = 1 + 2 + 6 + 5 + 2 = 16\), \(E(X^2) = 8 + 24 + 96 + 100 + 48 = 276\), and \(276 - 256 = 20\).
  5. \(X\) takes \(-2, 3, 1\) with probabilities \(\tfrac13, \tfrac12, \tfrac16\). Find (i) \(E(2X + 5)\), (ii) \(E(X^2)\). Ans. (i) 7, (ii) 6: \(E(X) = -\tfrac23 + \tfrac32 + \tfrac16 = 1\), so \(E(2X + 5) = 7\); \(E(X^2) = \tfrac43 + \tfrac92 + \tfrac16 = 6\).
  6. A coin is tossed until a tail appears. Find the expected number of tosses. Ans. 2, as in Worked Problem 6 with the faces exchanged.
  7. Five coins are tossed. Find the expected number of heads. Ans. \(\tfrac{80}{32} = \tfrac52\): \(\sum x\binom5x/32 = (5 + 20 + 30 + 20 + 5)/32\).
  8. \(E(X) = 3\) and \(E(X^2) = 13\). Use Chebyshev's inequality to find a lower bound for \(P(-2 < X < 8)\). Ans. \(\tfrac{21}{25}\): \(\sigma^2 = 13 - 9 = 4\), the interval is \(|X - 3| < 5\), and \(1 - 4/25 = 21/25\).

Key Take-aways from Unit 4