Skip to the content

Topics Covered

Axiomatic Definition (Kolmogorov, 1933) Important Consequences Multiplication Rule Total Probability Theorem Bayes' Theorem Independent Events Properties of CDF Discrete vs Continuous Moments Variance Properties Key Properties Memoryless Property of Geometric

Topic Overview — What & Why

Unit I lays the mathematical foundation for the rest of statistics. Every estimator, every hypothesis test, and every model relies on the laws of probability and on the behaviour of random variables.

  • Probability axioms & conditional probability: rules that quantify uncertainty. Bayes' theorem inverts the direction of conditioning — the engine behind modern Bayesian inference and many diagnostic tests.
  • Random variables & distribution functions: abstract numerical encodings of experimental outcomes. CDFs unify discrete and continuous cases.
  • Expectation, moments & MGF: summarise location, spread, skewness, kurtosis. The MGF, when it exists, uniquely determines a distribution and turns convolution of independent r.v.s into multiplication.
  • Standard distributions: Bernoulli through Cauchy — the alphabet of probabilistic modelling. Recognising which distribution applies in a problem is half the work.
  • Joint & conditional distributions: framework for studying dependence; covariance, correlation, conditional expectation are tools for prediction.
  • Chebyshev & Markov inequalities: distribution-free bounds. They guarantee something about deviations even when we don't know the full distribution.
  • Sampling distributions: the distributions of statistics computed from samples — backbone of inference ($t$, $\chi^2$, $F$).
  • Transformations & characteristic function: how distributions change under functional changes. The characteristic function exists for every distribution and is the key to limit theorems.
  • Convergence, LLN & CLT: capstone results — sample averages stabilise (LLN) and become approximately normal (CLT). Without these, statistics would not work asymptotically.

1. Basic Concepts of Probability

Why this section? Before we can talk about random variables, we need a precise language for events and their likelihoods. Kolmogorov's axioms (1933) put probability on rigorous measure-theoretic footing — ensuring that everything we derive later (means, tests, estimators) is internally consistent.

Random Experiment: An experiment whose outcome cannot be predicted with certainty although all possible outcomes are known. The set of all possible outcomes is the sample space $\Omega$. A subset of $\Omega$ is called an event.

Axiomatic Definition (Kolmogorov, 1933)

Let $\Omega$ be the sample space and $\mathcal{F}$ a $\sigma$-field of events. A probability measure $P:\mathcal{F}\to[0,1]$ satisfies:

  1. Non-negativity: $P(A) \geq 0$ for every $A \in \mathcal{F}$.
  2. Normalization: $P(\Omega) = 1$.
  3. Countable Additivity: If $A_1, A_2,\ldots$ are pairwise disjoint, then $$P\left(\bigcup_{i=1}^{\infty} A_i\right) = \sum_{i=1}^{\infty} P(A_i).$$

Important Consequences

EXAMPLE 1 Two-die problem. Two fair dice are rolled. Find the probability that the sum is 7 or one die shows 5.
$|\Omega| = 36$. Let $A$ = "sum is 7" $= \{(1,6),(2,5),(3,4),(4,3),(5,2),(6,1)\}$, $|A|=6$. Let $B$ = "one die shows 5", $|B|=11$. $A \cap B = \{(2,5),(5,2)\}$, $|A\cap B|=2$. $$P(A\cup B)=\tfrac{6}{36}+\tfrac{11}{36}-\tfrac{2}{36}=\tfrac{15}{36}=\tfrac{5}{12}.$$
EXAMPLE 2 Birthday matching. In a class of $n=23$ students, find the probability that at least two share a birthday (ignoring leap years). $$P(\text{no match}) = \frac{365 \cdot 364 \cdots (365-n+1)}{365^n}.$$ For $n=23$: $P(\text{match}) = 1 - 0.4927 \approx 0.5073$ — surprisingly high!

🌍 Where it's used in real life

  1. Weather forecasting — the chance of rain tomorrow.
  2. Insurance — pricing policies from accident probabilities.
  3. Casinos and lotteries — working out the odds of a game.
  4. Medical screening — a patient's risk before any test.
  5. Spam filters — the probability an email is junk.

2. Conditional Probability and Bayes' Theorem

Conditional Probability: For events $A,B$ with $P(B)>0$, $$P(A\mid B)=\frac{P(A\cap B)}{P(B)}.$$

Multiplication Rule

$$P(A\cap B)=P(A\mid B)P(B)=P(B\mid A)P(A).$$

Total Probability Theorem

If $\{B_1,B_2,\ldots,B_n\}$ is a partition of $\Omega$ with $P(B_i)>0$, then for any event $A$: $$P(A)=\sum_{i=1}^{n}P(A\mid B_i)P(B_i).$$

Bayes' Theorem

Bayes' Theorem. For a partition $\{B_i\}$ and event $A$ with $P(A)>0$: $$P(B_k\mid A)=\frac{P(A\mid B_k)P(B_k)}{\sum_{i=1}^{n}P(A\mid B_i)P(B_i)}.$$ $P(B_k)$ is called the prior, $P(A\mid B_k)$ the likelihood, and $P(B_k\mid A)$ the posterior.

Independent Events

$A$ and $B$ are independent iff $P(A\cap B)=P(A)P(B)$. Equivalently $P(A\mid B)=P(A)$.

EXAMPLE 1 Disease testing. A disease has prevalence $1\%$. A test has sensitivity $99\%$ and specificity $95\%$. If a person tests positive, what is the probability they have the disease?
Let $D$ = disease, $T^+$ = positive test. $$P(D\mid T^+)=\frac{P(T^+\mid D)P(D)}{P(T^+\mid D)P(D)+P(T^+\mid D^c)P(D^c)}=\frac{0.99 \times 0.01}{0.99 \times 0.01 + 0.05 \times 0.99}=\frac{0.0099}{0.0594}\approx 0.167.$$ Only 16.7% — a useful low base-rate illustration.
EXAMPLE 2 Three urns. Urn 1: 2W,3B; Urn 2: 4W,1B; Urn 3: 3W,2B. An urn is chosen at random and a white ball is drawn. Find $P(\text{Urn 2}\mid W)$.
$P(W\mid U_1)=2/5,\ P(W\mid U_2)=4/5,\ P(W\mid U_3)=3/5$, each $P(U_i)=1/3$. $$P(U_2\mid W)=\frac{(4/5)(1/3)}{(2/5+4/5+3/5)(1/3)}=\frac{4/15}{9/15}=\tfrac{4}{9}.$$

🌍 Where it's used in real life

  1. Medical diagnosis — updating disease risk after a positive test.
  2. Spam detection — P(spam | email says "free offer").
  3. Courtroom evidence — updating guilt given a DNA match.
  4. Search and rescue — narrowing a lost ship's location as areas are cleared.
  5. Recommendations — P(you'll like a film | you liked similar ones).

3. Random Variables and Distribution Functions

A random variable (r.v.) $X$ is a measurable function $X:\Omega\to\mathbb{R}$. The cumulative distribution function (CDF) is $$F_X(x)=P(X\le x),\quad x\in\mathbb{R}.$$

Properties of CDF

Discrete vs Continuous

Discrete: Range is countable; probability mass function (pmf) $p(x)=P(X=x)$ with $\sum p(x)=1$.

Continuous: $F$ is absolutely continuous; probability density function (pdf) $f(x)=F'(x)$ with $\int_{-\infty}^{\infty} f(x)\,dx=1$.

EXAMPLE 1 Let $X$ be the number of heads in 3 tosses of a fair coin. Then $X\sim\text{Bin}(3,1/2)$ with pmf $p(0)=1/8,\,p(1)=3/8,\,p(2)=3/8,\,p(3)=1/8$. CDF: $F(0)=1/8,\,F(1)=4/8,\,F(2)=7/8,\,F(3)=1$.
EXAMPLE 2 $X$ has pdf $f(x)=2x$ for $0\le x\le 1$, zero otherwise. Then $F(x)=x^2$ for $0\le x\le 1$. Hence $P(0.3<X\le 0.7)=F(0.7)-F(0.3)=0.49-0.09=0.40.$

🌍 Where it's used in real life

  1. Number of customers arriving at a shop in an hour.
  2. Daily rainfall recorded in a city.
  3. Number of defective items in a production batch.
  4. Waiting time at a bus stop.
  5. Exam scores across a class.

4. Expectation and Moments

Expectation: $$E(X)=\sum_x x\,p(x)\quad(\text{discrete}),\qquad E(X)=\int_{-\infty}^{\infty} x\,f(x)\,dx\quad(\text{continuous}).$$ For a function $g$, $E[g(X)]=\sum g(x)p(x)$ or $\int g(x)f(x)dx$.

Properties

Moments

The $r$-th raw moment about origin: $\mu'_r=E(X^r)$. The $r$-th central moment: $\mu_r=E[(X-\mu)^r]$ where $\mu=E(X)$.

Variance Properties

$$\text{Var}(X)=E(X^2)-[E(X)]^2,\quad \text{Var}(aX+b)=a^2\text{Var}(X).$$
EXAMPLE 1 For a fair die, $E(X)=\frac{1+2+\cdots+6}{6}=3.5$, $E(X^2)=\frac{91}{6}$, so $\text{Var}(X)=\frac{91}{6}-(3.5)^2=\frac{35}{12}\approx 2.917.$
EXAMPLE 2 For $X$ with pdf $f(x)=2x,\,0\le x\le 1$: $E(X)=\int_0^1 2x^2 dx=\tfrac{2}{3}$, $E(X^2)=\int_0^1 2x^3 dx=\tfrac{1}{2}$, so $\text{Var}(X)=\tfrac{1}{2}-\tfrac{4}{9}=\tfrac{1}{18}.$

🌍 Where it's used in real life

  1. Expected profit or loss of a business decision.
  2. Average claim size an insurer must budget for.
  3. Expected winnings in a game or bet.
  4. Mean delivery time a courier promises.
  5. Comparing funds by average return and volatility.

5. Moment Generating Function (MGF)

$$M_X(t)=E(e^{tX}),\quad \text{whenever the expectation exists in a neighbourhood of }0.$$ The $r$-th moment is obtained by $E(X^r)=M_X^{(r)}(0).$

Key Properties

EXAMPLE 1 $X\sim\text{Exp}(\lambda)$ has $M_X(t)=\frac{\lambda}{\lambda-t},\,t<\lambda.$ Then $E(X)=M'(0)=1/\lambda$, $E(X^2)=2/\lambda^2$, $\text{Var}(X)=1/\lambda^2.$
EXAMPLE 2 If $X_i\overset{\text{iid}}{\sim}\text{Bin}(n_i,p)$, then $M_{X_i}(t)=(q+pe^t)^{n_i}$. By independence, $\sum X_i$ has MGF $(q+pe^t)^{\sum n_i}\Rightarrow \sum X_i\sim\text{Bin}(\sum n_i,p).$

🌍 Where it's used in real life

  1. Quickly deriving a model's mean and variance in actuarial work.
  2. Showing that sums of independent risks stay in the same family.
  3. Physics — generating moments of energy distributions.
  4. Combining independent call-arrival counts in telecom.
  5. Deriving properties of standard distributions.

6. Standard Discrete Distributions

A discrete distribution is the probability model of a discrete random variable — it is fixed by a probability mass function $p(x)=P(X=x)$ over a countable set of values, with $p(x)\ge 0$ and $\sum_x p(x)=1$. The families below are the standard ones that recur throughout the syllabus.
DistributionpmfMeanVarianceMGF
Bernoulli($p$)$p^x(1-p)^{1-x},\,x=0,1$$p$$pq$$q+pe^t$
Binomial($n,p$)$\binom{n}{x}p^x q^{n-x}$$np$$npq$$(q+pe^t)^n$
Poisson($\lambda$)$\frac{e^{-\lambda}\lambda^x}{x!}$$\lambda$$\lambda$$e^{\lambda(e^t-1)}$
Geometric($p$)$pq^{x-1},\,x\ge 1$$1/p$$q/p^2$$\frac{pe^t}{1-qe^t}$
Neg. Binomial($r,p$)$\binom{x-1}{r-1}p^r q^{x-r}$$r/p$$rq/p^2$$\left(\frac{pe^t}{1-qe^t}\right)^r$
Hypergeometric$(N,K,n)$$\frac{\binom{K}{x}\binom{N-K}{n-x}}{\binom{N}{n}}$$nK/N$$\frac{nK(N-K)(N-n)}{N^2(N-1)}$—
Discrete Uniform$\{1,\ldots,N\}$$1/N$$\frac{N+1}{2}$$\frac{N^2-1}{12}$$\frac{1}{N}\sum_{k=1}^N e^{kt}$

Memoryless Property of Geometric

$$P(X>m+n\mid X>m)=P(X>n).$$ The geometric is the only discrete distribution with this property.

Poisson as Limit of Binomial

If $n\to\infty$, $p\to 0$, $np\to\lambda$, then $\text{Bin}(n,p)\to\text{Poisson}(\lambda).$
EXAMPLE 1 A coin with $p=0.3$ is tossed 10 times. $P(X=4)=\binom{10}{4}(0.3)^4(0.7)^6\approx 0.2001$. Mean $=3$, Var $=2.1$.
EXAMPLE 2 Calls arrive at a switchboard at rate $\lambda=4$/min. Probability of exactly 6 calls in a minute: $P(X=6)=\frac{e^{-4}4^6}{6!}\approx 0.1042.$

🌍 Where it's used in real life

  1. Binomial — pass/fail counts, yes/no poll answers.
  2. Poisson — calls to a call centre or goals in a match.
  3. Geometric — sales calls made until the first success.
  4. Hypergeometric — quality checks drawing items without replacement.
  5. Negative binomial — trials until a target number of wins.

7. Standard Continuous Distributions

A continuous distribution is the probability model of a continuous random variable — it is fixed by a probability density function $f(x)\ge 0$ with $\int_{-\infty}^{\infty} f(x)\,dx=1$, so that $P(a\le X\le b)=\int_a^b f(x)\,dx$. The standard families are listed below.
DistributionpdfMeanVarianceMGF
Uniform$(a,b)$$\frac{1}{b-a}$$(a+b)/2$$(b-a)^2/12$$\frac{e^{tb}-e^{ta}}{t(b-a)}$
Normal$(\mu,\sigma^2)$$\frac{1}{\sigma\sqrt{2\pi}}e^{-(x-\mu)^2/(2\sigma^2)}$$\mu$$\sigma^2$$e^{\mu t+\sigma^2 t^2/2}$
Exponential$(\lambda)$$\lambda e^{-\lambda x},\,x\ge 0$$1/\lambda$$1/\lambda^2$$\lambda/(\lambda-t)$
Gamma$(\alpha,\beta)$$\frac{\beta^\alpha}{\Gamma(\alpha)}x^{\alpha-1}e^{-\beta x}$$\alpha/\beta$$\alpha/\beta^2$$\left(\frac{\beta}{\beta-t}\right)^\alpha$
Beta$(\alpha,\beta)$$\frac{x^{\alpha-1}(1-x)^{\beta-1}}{B(\alpha,\beta)}$$\frac{\alpha}{\alpha+\beta}$$\frac{\alpha\beta}{(\alpha+\beta)^2(\alpha+\beta+1)}$—
Chi-square$(\nu)$$\frac{1}{2^{\nu/2}\Gamma(\nu/2)}x^{\nu/2-1}e^{-x/2}$$\nu$$2\nu$$(1-2t)^{-\nu/2}$
Cauchy$(\mu,\sigma)$$\frac{1}{\pi\sigma[1+((x-\mu)/\sigma)^2]}$———

Memoryless Property of Exponential

$$P(X>s+t\mid X>s)=P(X>t).$$ The only continuous distribution with this property.

Standard Normal

$Z=(X-\mu)/\sigma\sim N(0,1)$ has pdf $\phi(z)=\frac{1}{\sqrt{2\pi}}e^{-z^2/2}$ and CDF $\Phi(z).$
2.1%13.6%34.1%34.1%13.6%2.1%μ−3σμ−2σμ−σμμ+σμ+2σμ+3σ Normal distribution — the 68–95–99.7 rule
Empirical rule. About 68% of a normal population lies within ±1σ of the mean, 95% within ±2σ, and 99.7% within ±3σ.
EXAMPLE 1 $X\sim N(50,100)$. Find $P(40<X<65)$.
Standardize: $Z=(X-50)/10$. $P(-1<Z<1.5)=\Phi(1.5)-\Phi(-1)=0.9332-0.1587=0.7745.$
EXAMPLE 2 Lifetime of a bulb is exponential with mean 500 hours. Probability it lasts at least 800 hours given it has lasted 300 hours: $$P(X>800\mid X>300)=P(X>500)=e^{-500/500}=e^{-1}\approx 0.368.$$

🌍 Where it's used in real life

  1. Normal — heights, IQ scores, measurement errors.
  2. Exponential — time between machine breakdowns.
  3. Uniform — random-number generation in simulations.
  4. Gamma — insurance claim sizes and rainfall totals.
  5. Beta — modelling proportions like conversion rates.

8. Jointly Distributed Random Variables

Joint pdf/pmf: $f_{X,Y}(x,y)$ such that $\iint f=1$ (or $\sum\sum=1$).
Marginal: $f_X(x)=\int f(x,y)\,dy$.
Conditional: $f_{Y\mid X}(y\mid x)=f(x,y)/f_X(x)$ when $f_X(x)>0.$

Independence

$X,Y$ are independent iff $f_{X,Y}(x,y)=f_X(x)f_Y(y)$ for all $(x,y).$

Covariance and Correlation

$$\text{Cov}(X,Y)=E(XY)-E(X)E(Y),\qquad \rho=\frac{\text{Cov}(X,Y)}{\sigma_X\sigma_Y},\quad -1\le\rho\le 1.$$

$\text{Var}(aX+bY)=a^2\text{Var}(X)+b^2\text{Var}(Y)+2ab\,\text{Cov}(X,Y).$

Conditional Expectation

$$E(Y\mid X=x)=\int y\,f_{Y\mid X}(y\mid x)\,dy.$$ $$E(Y)=E[E(Y\mid X)],\quad \text{Var}(Y)=E[\text{Var}(Y\mid X)]+\text{Var}[E(Y\mid X)].$$
EXAMPLE 1 Let $f(x,y)=2$ for $0<x<y<1$. Then $f_X(x)=\int_x^1 2\,dy=2(1-x)$, $f_{Y\mid X}(y\mid x)=\frac{2}{2(1-x)}=\frac{1}{1-x}$ on $(x,1)$. So $E(Y\mid X=x)=\frac{1+x}{2}.$
EXAMPLE 2 For two fair dice $X,Y$, $\text{Cov}(X,Y)=0$ (independence), $\text{Var}(X+Y)=2(35/12)=35/6.$

🌍 Where it's used in real life

  1. Height and weight of people considered together.
  2. Returns of two stocks for portfolio risk.
  3. Temperature and ice-cream sales.
  4. Advertising spend and revenue.
  5. Rainfall and crop yield together.

9. Chebyshev's Inequality

Chebyshev's Inequality. For r.v. $X$ with mean $\mu$ and variance $\sigma^2<\infty$, for any $k>0$: $$P(|X-\mu|\ge k\sigma)\le\frac{1}{k^2}.$$ Equivalently $P(|X-\mu|\ge\varepsilon)\le\sigma^2/\varepsilon^2.$

Markov's Inequality

For non-negative $X$ and $a>0$: $P(X\ge a)\le E(X)/a.$ Chebyshev's follows by applying Markov's to $(X-\mu)^2.$
EXAMPLE 1 $\mu=50,\sigma=5$. Bound on $P(40<X<60)$: $P(|X-50|\ge 10)\le 25/100=0.25\Rightarrow P(40<X<60)\ge 0.75.$
EXAMPLE 2 For $X\sim\text{Poisson}(9)$, $P(|X-9|\ge 6)\le 9/36=1/4.$ Compared to exact $\approx 0.043$ — Chebyshev is loose but distribution-free.

🌍 Where it's used in real life

  1. Guaranteeing quality limits without knowing the exact distribution.
  2. Setting conservative risk bounds in finance.
  3. Choosing a safe sample size when little is known.
  4. Flagging outliers many standard deviations from the mean.
  5. Reliability tolerances in engineering.

10. Sampling Distributions

A statistic is any function of the sample (such as $\bar X$ or $S^2$) that involves no unknown parameters. Its probability distribution, arising from the randomness of the sample, is called a sampling distribution — the foundation of all inference, since it tells us how an estimator or test statistic behaves from sample to sample.

Let $X_1,\ldots,X_n$ be iid $N(\mu,\sigma^2)$. Define $\bar X=\frac{1}{n}\sum X_i$, $S^2=\frac{1}{n-1}\sum(X_i-\bar X)^2.$

Key Distributions

EXAMPLE 1 A sample of $n=25$ from $N(\mu,16)$. Then $\bar X\sim N(\mu,16/25)$, so SD$(\bar X)=0.8$. $P(|\bar X-\mu|<1)=P(|Z|<1.25)\approx 0.789.$
EXAMPLE 2 For $n=10$ from a normal population with unknown $\sigma$, the pivot $\frac{\bar X-\mu}{S/\sqrt{10}}\sim t_9.$ For 95% CI, use $t_{9,0.025}=2.262.$

🌍 Where it's used in real life

  1. Estimating average income from a survey sample.
  2. Quality control on the mean of sampled items.
  3. Poll margins of error before an election.
  4. Comparing two teaching methods' mean scores (t).
  5. Testing a process variance (chi-square / F).

11. Transformation of Random Variables

A transformation of a random variable derives the distribution of a new variable $Y=g(X)$ (or $\mathbf Y=g(\mathbf X)$) from the known distribution of $X$. The methods below — the CDF method, the Jacobian method, and the probability integral transform — obtain the pmf/pdf of $Y$.

One-to-one Transformation (Continuous)

If $Y=g(X)$ is monotonic with inverse $X=g^{-1}(Y)$: $$f_Y(y)=f_X(g^{-1}(y))\left|\frac{dx}{dy}\right|.$$

Multivariate (Jacobian)

If $(Y_1,Y_2)=g(X_1,X_2)$ with Jacobian $J$: $$f_{Y_1,Y_2}(y_1,y_2)=f_{X_1,X_2}(x_1,x_2)\cdot|J|.$$

Probability Integral Transform

If $X$ has continuous CDF $F$, then $U=F(X)\sim\text{Uniform}(0,1).$ Useful for simulation.
EXAMPLE 1 $X\sim\text{Exp}(1)$, $Y=X^2$. Then $x=\sqrt y$, $|dx/dy|=1/(2\sqrt y)$. So $f_Y(y)=e^{-\sqrt y}/(2\sqrt y),\,y>0.$
EXAMPLE 2 $X,Y\overset{\text{iid}}{\sim}N(0,1)$. Polar transform $R=\sqrt{X^2+Y^2}$, $\Theta=\arctan(Y/X)$ gives $R^2\sim\chi^2_2=\text{Exp}(1/2)$ and $\Theta\sim\text{Uniform}(0,2\pi)$ — basis of Box-Muller.

🌍 Where it's used in real life

  1. Converting measurement units inside a model.
  2. Simulating random values (inverse-transform) for games.
  3. Log-transforming skewed income or price data.
  4. Finding the distribution of area from a random radius.
  5. Box–Muller generating normal noise from uniform numbers.

12. Characteristic Function

$$\varphi_X(t)=E(e^{itX}),\quad t\in\mathbb{R}.$$ Always exists (since $|e^{itX}|=1$). Uniquely determines the distribution.

Properties

EXAMPLE 1 $X\sim N(0,1)$: $\varphi(t)=e^{-t^2/2}.$ Hence $X+Y$ for independent standard normals has $\varphi=e^{-t^2}\Rightarrow X+Y\sim N(0,2).$
EXAMPLE 2 $X\sim\text{Cauchy}(0,1)$: $\varphi(t)=e^{-|t|}.$ For iid $X_1,\ldots,X_n$: $\bar X$ has $\varphi(t)=e^{-|t|/n\cdot n}=e^{-|t|}$, i.e. $\bar X\sim\text{Cauchy}(0,1)$ — CLT fails (no finite mean).

🌍 Where it's used in real life

  1. Proving the Central Limit Theorem.
  2. Adding independent insurance risks (convolution).
  3. Signal processing in the frequency domain.
  4. Handling heavy-tailed data where the MGF fails.
  5. Deriving limit results for large samples.

13. Modes of Convergence of Random Variables

ModeDefinitionNotation
Almost sure$P(\lim_{n\to\infty}X_n=X)=1$$X_n\xrightarrow{a.s.}X$
In probability$\forall\varepsilon>0,\,P(|X_n-X|>\varepsilon)\to 0$$X_n\xrightarrow{P}X$
In $r$-th mean$E|X_n-X|^r\to 0$$X_n\xrightarrow{L^r}X$
In distribution$F_n(x)\to F(x)$ at continuity points of $F$$X_n\xrightarrow{d}X$

Implications

$$X_n\xrightarrow{a.s.}X\Rightarrow X_n\xrightarrow{P}X\Rightarrow X_n\xrightarrow{d}X.$$ $$X_n\xrightarrow{L^r}X\Rightarrow X_n\xrightarrow{P}X.$$ Convergence in distribution to a constant implies convergence in probability.
almost sure in probability in distribution in r-th mean (Lʳ) Implications among modes of convergence
One-way implications. Almost-sure and $L^r$ convergence each imply convergence in probability, which implies convergence in distribution. None of the arrows reverses in general (the lone exception: a limit that is a constant).

Slutsky's Theorem

If $X_n\xrightarrow{d}X$ and $Y_n\xrightarrow{P}c$, then $X_n+Y_n\xrightarrow{d}X+c$, $X_nY_n\xrightarrow{d}cX.$
EXAMPLE 1 $X_n=1$ on $[0,1/n]$ and $0$ elsewhere, on Uniform$(0,1)$. Since $P(X_n=1)=1/n\to 0$, we have $X_n\xrightarrow{P}0$. Moreover, for every fixed $\omega>0$ we have $X_n(\omega)=0$ as soon as $n>1/\omega$, so $X_n(\omega)\to 0$ for all $\omega\in(0,1]$; hence $X_n\xrightarrow{a.s.}0$ as well.
EXAMPLE 2 Let $X_n\sim N(0,1)$ for all $n$. Then $X_n\xrightarrow{d}X\sim N(0,1)$ but does not converge in probability (each $X_n$ is independent).

🌍 Where it's used in real life

  1. Justifying that estimates improve as data grows.
  2. Accuracy guarantees for Monte-Carlo simulation.
  3. Reliability of long-run averages in insurance.
  4. Convergence of machine-learning training.
  5. Approximating distributions in large samples.

14. Weak and Strong Laws of Large Numbers

Weak Law of Large Numbers (WLLN). If $X_1,\ldots,X_n$ are iid with finite mean $\mu$, then $$\bar X_n=\frac{1}{n}\sum X_i\xrightarrow{P}\mu.$$
Strong Law of Large Numbers (Kolmogorov, SLLN). If $X_i$ iid with $E|X|<\infty$, $$\bar X_n\xrightarrow{a.s.}\mu.$$

Khintchin's WLLN

For iid $X_i$ with finite mean, $\bar X_n\xrightarrow{P}\mu$ (no second moment needed).

Chebyshev's WLLN

For uncorrelated $X_i$ with bounded variances, $\bar X_n - E(\bar X_n)\xrightarrow{P}0.$
EXAMPLE 1 Toss a fair coin $n$ times; $X_i$ = 1 if head. SLLN: relative frequency of heads $\to 1/2$ a.s.
EXAMPLE 2 For Cauchy distribution, mean does not exist; $\bar X_n$ does NOT converge — both LLNs fail.

🌍 Where it's used in real life

  1. Why casinos always profit in the long run.
  2. Insurance — average claims stabilise over many policies.
  3. Opinion polls sharpen with larger samples.
  4. A/B testing conversion rates over many visitors.
  5. Estimating π or integrals by simulation.

15. Central Limit Theorem (CLT)

Lindeberg–Lévy CLT. Let $X_1,X_2,\ldots$ be iid with mean $\mu$ and finite variance $\sigma^2>0$. Then $$\frac{\bar X_n - \mu}{\sigma/\sqrt n}\xrightarrow{d}N(0,1)\quad\text{as } n\to\infty.$$ Equivalently, $\sum_{i=1}^n X_i\approx N(n\mu,n\sigma^2)$ for large $n$.

Other CLT variants

Continuity correction

For integer-valued r.v.: $P(X\le k)\approx \Phi\left(\frac{k+0.5-np}{\sqrt{npq}}\right).$
EXAMPLE 1 $X_i\sim\text{Bernoulli}(0.4)$, $n=100$. $P(\bar X>0.5)=P\left(Z>\frac{0.5-0.4}{\sqrt{0.24/100}}\right)=P(Z>2.04)\approx 0.0207.$
EXAMPLE 2 Suppose $X_i\sim\text{Exp}(1)$, $n=50$. $E(\bar X)=1$, $\text{Var}(\bar X)=1/50.$ Then $P(\bar X<0.8)\approx P(Z<(0.8-1)/\sqrt{0.02})=P(Z<-1.414)\approx 0.0786.$

🌍 Where it's used in real life

  1. Why averages and errors look bell-shaped.
  2. Poll margins of error and confidence intervals.
  3. Control charts assuming normal sample means.
  4. Pooling many small risks in finance.
  5. Approximating binomial counts with a normal curve.