Topics Covered
Contents
- 1. Basic Probability
- 2. Conditional Probability & Bayes' Theorem
- 3. Random Variables & Distribution Functions
- 4. Expectation and Moments
- 5. Moment Generating Function
- 6. Standard Discrete Distributions
- 7. Standard Continuous Distributions
- 8. Jointly Distributed Random Variables
- 9. Chebyshev's Inequality
- 10. Sampling Distributions
- 11. Transformation of Random Variables
- 12. Characteristic Function
- 13. Modes of Convergence
- 14. Laws of Large Numbers
- 15. Central Limit Theorem
Topic Overview — What & Why
Unit I lays the mathematical foundation for the rest of statistics. Every estimator, every hypothesis test, and every model relies on the laws of probability and on the behaviour of random variables.
- Probability axioms & conditional probability: rules that quantify uncertainty. Bayes' theorem inverts the direction of conditioning — the engine behind modern Bayesian inference and many diagnostic tests.
- Random variables & distribution functions: abstract numerical encodings of experimental outcomes. CDFs unify discrete and continuous cases.
- Expectation, moments & MGF: summarise location, spread, skewness, kurtosis. The MGF, when it exists, uniquely determines a distribution and turns convolution of independent r.v.s into multiplication.
- Standard distributions: Bernoulli through Cauchy — the alphabet of probabilistic modelling. Recognising which distribution applies in a problem is half the work.
- Joint & conditional distributions: framework for studying dependence; covariance, correlation, conditional expectation are tools for prediction.
- Chebyshev & Markov inequalities: distribution-free bounds. They guarantee something about deviations even when we don't know the full distribution.
- Sampling distributions: the distributions of statistics computed from samples — backbone of inference ($t$, $\chi^2$, $F$).
- Transformations & characteristic function: how distributions change under functional changes. The characteristic function exists for every distribution and is the key to limit theorems.
- Convergence, LLN & CLT: capstone results — sample averages stabilise (LLN) and become approximately normal (CLT). Without these, statistics would not work asymptotically.
1. Basic Concepts of Probability
Why this section? Before we can talk about random variables, we need a precise language for events and their likelihoods. Kolmogorov's axioms (1933) put probability on rigorous measure-theoretic footing — ensuring that everything we derive later (means, tests, estimators) is internally consistent.
Axiomatic Definition (Kolmogorov, 1933)
Let $\Omega$ be the sample space and $\mathcal{F}$ a $\sigma$-field of events. A probability measure $P:\mathcal{F}\to[0,1]$ satisfies:
- Non-negativity: $P(A) \geq 0$ for every $A \in \mathcal{F}$.
- Normalization: $P(\Omega) = 1$.
- Countable Additivity: If $A_1, A_2,\ldots$ are pairwise disjoint, then $$P\left(\bigcup_{i=1}^{\infty} A_i\right) = \sum_{i=1}^{\infty} P(A_i).$$
Important Consequences
- $P(A^c) = 1 - P(A)$
- $P(A \cup B) = P(A) + P(B) - P(A \cap B)$ (Addition Theorem)
- Inclusion–Exclusion: $P\!\left(\bigcup_{i=1}^n A_i\right) = \sum_{i} P(A_i) - \sum_{i<j} P(A_i \cap A_j) + \sum_{i<j<k} P(A_i \cap A_j \cap A_k) - \cdots + (-1)^{n+1} P(A_1 \cap \cdots \cap A_n)$
- Boole's Inequality: $P(\bigcup A_i) \leq \sum P(A_i)$
$|\Omega| = 36$. Let $A$ = "sum is 7" $= \{(1,6),(2,5),(3,4),(4,3),(5,2),(6,1)\}$, $|A|=6$. Let $B$ = "one die shows 5", $|B|=11$. $A \cap B = \{(2,5),(5,2)\}$, $|A\cap B|=2$. $$P(A\cup B)=\tfrac{6}{36}+\tfrac{11}{36}-\tfrac{2}{36}=\tfrac{15}{36}=\tfrac{5}{12}.$$
🌍 Where it's used in real life
- Weather forecasting — the chance of rain tomorrow.
- Insurance — pricing policies from accident probabilities.
- Casinos and lotteries — working out the odds of a game.
- Medical screening — a patient's risk before any test.
- Spam filters — the probability an email is junk.
2. Conditional Probability and Bayes' Theorem
Multiplication Rule
$$P(A\cap B)=P(A\mid B)P(B)=P(B\mid A)P(A).$$Total Probability Theorem
If $\{B_1,B_2,\ldots,B_n\}$ is a partition of $\Omega$ with $P(B_i)>0$, then for any event $A$: $$P(A)=\sum_{i=1}^{n}P(A\mid B_i)P(B_i).$$
Bayes' Theorem
Independent Events
$A$ and $B$ are independent iff $P(A\cap B)=P(A)P(B)$. Equivalently $P(A\mid B)=P(A)$.
Let $D$ = disease, $T^+$ = positive test. $$P(D\mid T^+)=\frac{P(T^+\mid D)P(D)}{P(T^+\mid D)P(D)+P(T^+\mid D^c)P(D^c)}=\frac{0.99 \times 0.01}{0.99 \times 0.01 + 0.05 \times 0.99}=\frac{0.0099}{0.0594}\approx 0.167.$$ Only 16.7% — a useful low base-rate illustration.
$P(W\mid U_1)=2/5,\ P(W\mid U_2)=4/5,\ P(W\mid U_3)=3/5$, each $P(U_i)=1/3$. $$P(U_2\mid W)=\frac{(4/5)(1/3)}{(2/5+4/5+3/5)(1/3)}=\frac{4/15}{9/15}=\tfrac{4}{9}.$$
🌍 Where it's used in real life
- Medical diagnosis — updating disease risk after a positive test.
- Spam detection — P(spam | email says "free offer").
- Courtroom evidence — updating guilt given a DNA match.
- Search and rescue — narrowing a lost ship's location as areas are cleared.
- Recommendations — P(you'll like a film | you liked similar ones).
3. Random Variables and Distribution Functions
Properties of CDF
- $F$ is non-decreasing.
- $\lim_{x\to-\infty}F(x)=0,\ \lim_{x\to\infty}F(x)=1.$
- $F$ is right-continuous: $\lim_{h\to 0^+}F(x+h)=F(x).$
- $P(a<X\le b)=F(b)-F(a).$
Discrete vs Continuous
Discrete: Range is countable; probability mass function (pmf) $p(x)=P(X=x)$ with $\sum p(x)=1$.
Continuous: $F$ is absolutely continuous; probability density function (pdf) $f(x)=F'(x)$ with $\int_{-\infty}^{\infty} f(x)\,dx=1$.
🌍 Where it's used in real life
- Number of customers arriving at a shop in an hour.
- Daily rainfall recorded in a city.
- Number of defective items in a production batch.
- Waiting time at a bus stop.
- Exam scores across a class.
4. Expectation and Moments
Properties
- Linearity: $E(aX+bY+c)=aE(X)+bE(Y)+c.$
- If $X\ge 0$ then $E(X)\ge 0$.
- If $X,Y$ independent: $E(XY)=E(X)E(Y).$
Moments
The $r$-th raw moment about origin: $\mu'_r=E(X^r)$. The $r$-th central moment: $\mu_r=E[(X-\mu)^r]$ where $\mu=E(X)$.
- $\mu_1=0$, $\mu_2=\text{Var}(X)=\sigma^2$.
- Skewness: $\beta_1=\mu_3^2/\mu_2^3$, $\gamma_1=\mu_3/\sigma^3$.
- Kurtosis: $\beta_2=\mu_4/\mu_2^2$, $\gamma_2=\beta_2-3$ (excess).
Variance Properties
$$\text{Var}(X)=E(X^2)-[E(X)]^2,\quad \text{Var}(aX+b)=a^2\text{Var}(X).$$🌍 Where it's used in real life
- Expected profit or loss of a business decision.
- Average claim size an insurer must budget for.
- Expected winnings in a game or bet.
- Mean delivery time a courier promises.
- Comparing funds by average return and volatility.
5. Moment Generating Function (MGF)
Key Properties
- $M_{aX+b}(t)=e^{bt}M_X(at).$
- If $X,Y$ independent, $M_{X+Y}(t)=M_X(t)M_Y(t).$
- Uniqueness Theorem: If two MGFs are equal in a neighbourhood of 0, the distributions are equal.
🌍 Where it's used in real life
- Quickly deriving a model's mean and variance in actuarial work.
- Showing that sums of independent risks stay in the same family.
- Physics — generating moments of energy distributions.
- Combining independent call-arrival counts in telecom.
- Deriving properties of standard distributions.
6. Standard Discrete Distributions
| Distribution | pmf | Mean | Variance | MGF |
|---|---|---|---|---|
| Bernoulli($p$) | $p^x(1-p)^{1-x},\,x=0,1$ | $p$ | $pq$ | $q+pe^t$ |
| Binomial($n,p$) | $\binom{n}{x}p^x q^{n-x}$ | $np$ | $npq$ | $(q+pe^t)^n$ |
| Poisson($\lambda$) | $\frac{e^{-\lambda}\lambda^x}{x!}$ | $\lambda$ | $\lambda$ | $e^{\lambda(e^t-1)}$ |
| Geometric($p$) | $pq^{x-1},\,x\ge 1$ | $1/p$ | $q/p^2$ | $\frac{pe^t}{1-qe^t}$ |
| Neg. Binomial($r,p$) | $\binom{x-1}{r-1}p^r q^{x-r}$ | $r/p$ | $rq/p^2$ | $\left(\frac{pe^t}{1-qe^t}\right)^r$ |
| Hypergeometric$(N,K,n)$ | $\frac{\binom{K}{x}\binom{N-K}{n-x}}{\binom{N}{n}}$ | $nK/N$ | $\frac{nK(N-K)(N-n)}{N^2(N-1)}$ | — |
| Discrete Uniform$\{1,\ldots,N\}$ | $1/N$ | $\frac{N+1}{2}$ | $\frac{N^2-1}{12}$ | $\frac{1}{N}\sum_{k=1}^N e^{kt}$ |
Memoryless Property of Geometric
$$P(X>m+n\mid X>m)=P(X>n).$$ The geometric is the only discrete distribution with this property.Poisson as Limit of Binomial
If $n\to\infty$, $p\to 0$, $np\to\lambda$, then $\text{Bin}(n,p)\to\text{Poisson}(\lambda).$🌍 Where it's used in real life
- Binomial — pass/fail counts, yes/no poll answers.
- Poisson — calls to a call centre or goals in a match.
- Geometric — sales calls made until the first success.
- Hypergeometric — quality checks drawing items without replacement.
- Negative binomial — trials until a target number of wins.
7. Standard Continuous Distributions
| Distribution | Mean | Variance | MGF | |
|---|---|---|---|---|
| Uniform$(a,b)$ | $\frac{1}{b-a}$ | $(a+b)/2$ | $(b-a)^2/12$ | $\frac{e^{tb}-e^{ta}}{t(b-a)}$ |
| Normal$(\mu,\sigma^2)$ | $\frac{1}{\sigma\sqrt{2\pi}}e^{-(x-\mu)^2/(2\sigma^2)}$ | $\mu$ | $\sigma^2$ | $e^{\mu t+\sigma^2 t^2/2}$ |
| Exponential$(\lambda)$ | $\lambda e^{-\lambda x},\,x\ge 0$ | $1/\lambda$ | $1/\lambda^2$ | $\lambda/(\lambda-t)$ |
| Gamma$(\alpha,\beta)$ | $\frac{\beta^\alpha}{\Gamma(\alpha)}x^{\alpha-1}e^{-\beta x}$ | $\alpha/\beta$ | $\alpha/\beta^2$ | $\left(\frac{\beta}{\beta-t}\right)^\alpha$ |
| Beta$(\alpha,\beta)$ | $\frac{x^{\alpha-1}(1-x)^{\beta-1}}{B(\alpha,\beta)}$ | $\frac{\alpha}{\alpha+\beta}$ | $\frac{\alpha\beta}{(\alpha+\beta)^2(\alpha+\beta+1)}$ | — |
| Chi-square$(\nu)$ | $\frac{1}{2^{\nu/2}\Gamma(\nu/2)}x^{\nu/2-1}e^{-x/2}$ | $\nu$ | $2\nu$ | $(1-2t)^{-\nu/2}$ |
| Cauchy$(\mu,\sigma)$ | $\frac{1}{\pi\sigma[1+((x-\mu)/\sigma)^2]}$ | — | — | — |
Memoryless Property of Exponential
$$P(X>s+t\mid X>s)=P(X>t).$$ The only continuous distribution with this property.Standard Normal
$Z=(X-\mu)/\sigma\sim N(0,1)$ has pdf $\phi(z)=\frac{1}{\sqrt{2\pi}}e^{-z^2/2}$ and CDF $\Phi(z).$Standardize: $Z=(X-50)/10$. $P(-1<Z<1.5)=\Phi(1.5)-\Phi(-1)=0.9332-0.1587=0.7745.$
🌍 Where it's used in real life
- Normal — heights, IQ scores, measurement errors.
- Exponential — time between machine breakdowns.
- Uniform — random-number generation in simulations.
- Gamma — insurance claim sizes and rainfall totals.
- Beta — modelling proportions like conversion rates.
8. Jointly Distributed Random Variables
Marginal: $f_X(x)=\int f(x,y)\,dy$.
Conditional: $f_{Y\mid X}(y\mid x)=f(x,y)/f_X(x)$ when $f_X(x)>0.$
Independence
$X,Y$ are independent iff $f_{X,Y}(x,y)=f_X(x)f_Y(y)$ for all $(x,y).$Covariance and Correlation
$$\text{Cov}(X,Y)=E(XY)-E(X)E(Y),\qquad \rho=\frac{\text{Cov}(X,Y)}{\sigma_X\sigma_Y},\quad -1\le\rho\le 1.$$$\text{Var}(aX+bY)=a^2\text{Var}(X)+b^2\text{Var}(Y)+2ab\,\text{Cov}(X,Y).$
Conditional Expectation
$$E(Y\mid X=x)=\int y\,f_{Y\mid X}(y\mid x)\,dy.$$ $$E(Y)=E[E(Y\mid X)],\quad \text{Var}(Y)=E[\text{Var}(Y\mid X)]+\text{Var}[E(Y\mid X)].$$🌍 Where it's used in real life
- Height and weight of people considered together.
- Returns of two stocks for portfolio risk.
- Temperature and ice-cream sales.
- Advertising spend and revenue.
- Rainfall and crop yield together.
9. Chebyshev's Inequality
Markov's Inequality
For non-negative $X$ and $a>0$: $P(X\ge a)\le E(X)/a.$ Chebyshev's follows by applying Markov's to $(X-\mu)^2.$🌍 Where it's used in real life
- Guaranteeing quality limits without knowing the exact distribution.
- Setting conservative risk bounds in finance.
- Choosing a safe sample size when little is known.
- Flagging outliers many standard deviations from the mean.
- Reliability tolerances in engineering.
10. Sampling Distributions
Let $X_1,\ldots,X_n$ be iid $N(\mu,\sigma^2)$. Define $\bar X=\frac{1}{n}\sum X_i$, $S^2=\frac{1}{n-1}\sum(X_i-\bar X)^2.$
- $\bar X\sim N(\mu,\sigma^2/n).$
- $\frac{(n-1)S^2}{\sigma^2}\sim\chi^2_{n-1}$ and is independent of $\bar X.$
- $T=\frac{\bar X-\mu}{S/\sqrt n}\sim t_{n-1}.$
- If $X\sim\chi^2_m, Y\sim\chi^2_n$ independent, $\frac{X/m}{Y/n}\sim F_{m,n}.$
Key Distributions
- $\chi^2_n$: sum of $n$ iid squared $N(0,1)$. Mean $n$, variance $2n$.
- $t_n$: $Z/\sqrt{V/n}$ where $Z\sim N(0,1)$, $V\sim\chi^2_n$ independent.
- $F_{m,n}$: ratio of independent $\chi^2$'s scaled by df.
🌍 Where it's used in real life
- Estimating average income from a survey sample.
- Quality control on the mean of sampled items.
- Poll margins of error before an election.
- Comparing two teaching methods' mean scores (t).
- Testing a process variance (chi-square / F).
11. Transformation of Random Variables
One-to-one Transformation (Continuous)
If $Y=g(X)$ is monotonic with inverse $X=g^{-1}(Y)$: $$f_Y(y)=f_X(g^{-1}(y))\left|\frac{dx}{dy}\right|.$$Multivariate (Jacobian)
If $(Y_1,Y_2)=g(X_1,X_2)$ with Jacobian $J$: $$f_{Y_1,Y_2}(y_1,y_2)=f_{X_1,X_2}(x_1,x_2)\cdot|J|.$$Probability Integral Transform
If $X$ has continuous CDF $F$, then $U=F(X)\sim\text{Uniform}(0,1).$ Useful for simulation.🌍 Where it's used in real life
- Converting measurement units inside a model.
- Simulating random values (inverse-transform) for games.
- Log-transforming skewed income or price data.
- Finding the distribution of area from a random radius.
- Box–Muller generating normal noise from uniform numbers.
12. Characteristic Function
Properties
- $\varphi_X(0)=1,\ |\varphi_X(t)|\le 1.$
- $\varphi_X(t)$ is uniformly continuous on $\mathbb{R}.$
- $\varphi_{aX+b}(t)=e^{ibt}\varphi_X(at).$
- If $X,Y$ independent: $\varphi_{X+Y}=\varphi_X\varphi_Y.$
- If $E|X|^k<\infty$: $\varphi^{(k)}(0)=i^k E(X^k).$
- Lévy continuity: $X_n\xrightarrow{d}X$ iff $\varphi_{X_n}(t)\to\varphi_X(t)$ for all $t.$
🌍 Where it's used in real life
- Proving the Central Limit Theorem.
- Adding independent insurance risks (convolution).
- Signal processing in the frequency domain.
- Handling heavy-tailed data where the MGF fails.
- Deriving limit results for large samples.
13. Modes of Convergence of Random Variables
| Mode | Definition | Notation |
|---|---|---|
| Almost sure | $P(\lim_{n\to\infty}X_n=X)=1$ | $X_n\xrightarrow{a.s.}X$ |
| In probability | $\forall\varepsilon>0,\,P(|X_n-X|>\varepsilon)\to 0$ | $X_n\xrightarrow{P}X$ |
| In $r$-th mean | $E|X_n-X|^r\to 0$ | $X_n\xrightarrow{L^r}X$ |
| In distribution | $F_n(x)\to F(x)$ at continuity points of $F$ | $X_n\xrightarrow{d}X$ |
Implications
$$X_n\xrightarrow{a.s.}X\Rightarrow X_n\xrightarrow{P}X\Rightarrow X_n\xrightarrow{d}X.$$ $$X_n\xrightarrow{L^r}X\Rightarrow X_n\xrightarrow{P}X.$$ Convergence in distribution to a constant implies convergence in probability.Slutsky's Theorem
If $X_n\xrightarrow{d}X$ and $Y_n\xrightarrow{P}c$, then $X_n+Y_n\xrightarrow{d}X+c$, $X_nY_n\xrightarrow{d}cX.$🌍 Where it's used in real life
- Justifying that estimates improve as data grows.
- Accuracy guarantees for Monte-Carlo simulation.
- Reliability of long-run averages in insurance.
- Convergence of machine-learning training.
- Approximating distributions in large samples.
14. Weak and Strong Laws of Large Numbers
Khintchin's WLLN
For iid $X_i$ with finite mean, $\bar X_n\xrightarrow{P}\mu$ (no second moment needed).Chebyshev's WLLN
For uncorrelated $X_i$ with bounded variances, $\bar X_n - E(\bar X_n)\xrightarrow{P}0.$🌍 Where it's used in real life
- Why casinos always profit in the long run.
- Insurance — average claims stabilise over many policies.
- Opinion polls sharpen with larger samples.
- A/B testing conversion rates over many visitors.
- Estimating π or integrals by simulation.
15. Central Limit Theorem (CLT)
Other CLT variants
- Lindeberg-Feller CLT: for independent but not identical r.v.'s satisfying the Lindeberg condition.
- Liapunov CLT: stronger moment condition.
- De Moivre-Laplace CLT: $\text{Bin}(n,p)\approx N(np,npq)$ for large $n.$
Continuity correction
For integer-valued r.v.: $P(X\le k)\approx \Phi\left(\frac{k+0.5-np}{\sqrt{npq}}\right).$🌍 Where it's used in real life
- Why averages and errors look bell-shaped.
- Poll margins of error and confidence intervals.
- Control charts assuming normal sample means.
- Pooling many small risks in finance.
- Approximating binomial counts with a normal curve.