Topics Covered
Contents
Topic Overview — What & Why
Unit V develops the formal framework for deciding between competing claims about a population using sample data. Estimation tells us what the parameter is; testing tells us whether a hypothesised value is plausible.
- Basic concepts: hypotheses, type I/II errors, size, power, $p$-value — the vocabulary every test-of-significance shares.
- Neyman-Pearson lemma: for simple-vs-simple testing, the likelihood-ratio test is most powerful. Foundation of optimal testing.
- Monotone likelihood ratio & Karlin-Rubin: when the likelihood ratio is monotone in a sufficient statistic, UMP tests for one-sided composite hypotheses exist.
- UMP, UMPU, UMPI tests: hierarchy of optimality concepts — UMP is rare, UMPU/UMPI broaden the class of "best" tests when UMP fails (e.g., two-sided tests).
- Likelihood ratio test (LRT): general-purpose construction; under regularity, $-2\log\Lambda$ is asymptotically $\chi^2$ (Wilks).
- Wald's SPRT: sequential testing — sample size random; on average uses fewer observations than fixed-sample tests for the same error rates.
- Chi-square tests: goodness-of-fit, independence, homogeneity. The workhorse of categorical data analysis.
- Sign & Wilcoxon signed-rank: distribution-free alternatives to the one-sample $t$-test for medians.
- Mann-Whitney U: nonparametric two-sample test; robust to non-normality and outliers.
- Linear rank tests: general framework $T=\sum c_i a(R_i)$ — Wilcoxon, normal scores, Ansari-Bradley, Mood are all special cases.
- Kruskal-Wallis: distribution-free analogue of one-way ANOVA for $k>2$ groups.
Where the derivations are
This unit is written to be read straight through as exam preparation: the statement, the condition, a worked example, and the MCQs at the end. The proofs behind it are written out elsewhere on this site, one step at a time, and they are worth reading once before you rely on the summaries here.
- Randomized tests and the complete Neyman–Pearson lemma — why a discrete distribution admits no non-randomized test of exact size, and the constant that repairs it, solved exactly.
- UMP tests, monotone likelihood ratio and similar regions — the Karlin–Rubin theorem proved, and a worked demonstration that no UMP test exists against a two-sided alternative.
- The likelihood ratio test, Wald and Rao score — the LRT reduced to the familiar tests, and one finite-sample case where all three give different answers.
- Sequential analysis and decision theory — the SPRT with its operating characteristic and average sample number computed and tabulated.
- Inferential Statistics, Unit 2 — the Foundation course: critical regions, both kinds of error as integrals, power, and the Neyman–Pearson lemma with its proof.
Where this unit is examined by other papers, see the ISS Paper II map and the CSIR NET map.
1. Basic Concepts
Why this section? Without precise definitions of error rates and power, all subsequent tests would be "shooting in the dark". Mastery of these terms makes every subsequent test transparent.
Type I error: reject $H_0$ when true. Probability $\alpha$ (size).
Type II error: accept $H_0$ when false. Probability $\beta$.
Power: $1-\beta=P(\text{reject }H_0\mid H_1).$
p-value: probability under $H_0$ of obtaining a test statistic as or more extreme than observed.
Common misconception. The $p$-value is not $P(H_0\text{ is true})$ and it is not the probability the result is due to chance. It is computed assuming $H_0$ holds: it answers “if $H_0$ were true, how surprising would data this extreme be?” A small $p$-value means the observed data sit far in the tail of the null distribution.
Test Function
A test is a function $\phi(x)\in[0,1]$: probability of rejection given data. $E_\theta[\phi(X)]=$ power function.Simple vs Composite
$H_0:\theta=\theta_0$ is simple; $H_0:\theta\le\theta_0$ is composite.🌍 Where it's used in real life
- Deciding whether a new drug really works.
- Checking if a coin or machine is fair.
- Did a website change lift sales?
- Is average wait time above target?
- Detecting bias in a process.
2. Neyman–Pearson Lemma
Intuition. To maximise power at a fixed size, spend your allowed Type-I “budget” on the sample outcomes where the data are relatively far more likely under $H_1$ than under $H_0$. Ranking outcomes by the likelihood ratio and rejecting the top ones until the size reaches $\alpha$ is provably the most powerful strategy — every unit of $\alpha$ is bought where it buys the most power.
🌍 Where it's used in real life
- Radar: is a signal present or absent?
- Setting a medical test's cut-off.
- Optimal spam vs not-spam rule.
- Accept/reject decisions in quality control.
- Fraud-detection thresholds.
3. Monotone Likelihood Ratio (MLR)
One-parameter exponential family has MLR in the natural sufficient statistic.
Karlin–Rubin Theorem
If MLR in $T$, the UMP test for $H_0:\theta\le\theta_0$ vs $H_1:\theta>\theta_0$ rejects when $T\ge c.$🌍 Where it's used in real life
- One-sided quality tests (is the defect rate too high?).
- Is mean strength above the specification?
- Detecting a rising failure rate.
- One-sided drug-improvement tests.
- Watching for increasing pollution levels.
4. UMP, UMPU, UMPI Tests
- UMP: uniformly most powerful — at every $\theta\in H_1$, beats any other test.
- UMPU: UMP among unbiased tests (power $\ge\alpha$ throughout $H_1$). Used when no UMP exists (e.g., two-sided tests).
- UMPI: UMP among invariant tests under a group of transformations.
Two-sided tests $H_0:\theta=\theta_0$ vs $H_1:\theta\ne\theta_0$ generally have no UMP; UMPU exists.
🌍 Where it's used in real life
- Picking the most powerful test available.
- Two-sided tests around a target value.
- Tests that don't depend on the units used.
- Best process-monitoring tests.
- Optimal clinical-trial test design.
5. Likelihood Ratio Test (LRT)
Wilks' Theorem
Under regularity, $-2\log\Lambda\xrightarrow{d}\chi^2_r$ under $H_0$, where $r=\dim\Theta-\dim\Theta_0.$Important LRTs
- One-sample $t$-test: $H_0:\mu=\mu_0,\sigma$ unknown.
- Two-sample $t$-test: $H_0:\mu_1=\mu_2.$
- $F$-test for variances: $H_0:\sigma_1^2=\sigma_2^2.$
- $\chi^2$ test for $\sigma^2$.
🌍 Where it's used in real life
- Comparing nested regression models.
- Testing genetic linkage models.
- Testing equality of several means.
- Model selection in machine learning.
- Testing whether data fit a distribution.
6. Wald's Sequential Probability Ratio Test (SPRT)
Test $H_0:\theta=\theta_0$ vs $H_1:\theta=\theta_1$ sequentially. After observing $X_1,\ldots,X_n$, compute likelihood ratio: $$\lambda_n=\frac{L_1}{L_0}=\prod_{i=1}^n\frac{f(x_i;\theta_1)}{f(x_i;\theta_0)}.$$ Define bounds $A=(1-\beta)/\alpha,\,B=\beta/(1-\alpha)$:
- If $\lambda_n\ge A$: reject $H_0.$
- If $\lambda_n\le B$: accept $H_0.$
- Otherwise: take another observation.
Operating Characteristic (OC)
$L(\theta)=P_\theta(\text{accept }H_0)$ — Wald's approximation: $$L(\theta)\approx\frac{A^{h(\theta)}-1}{A^{h(\theta)}-B^{h(\theta)}}.$$Average Sample Number (ASN)
$$E_\theta(N)\approx\frac{L(\theta)\log B+(1-L(\theta))\log A}{E_\theta\log[f(X;\theta_1)/f(X;\theta_0)]}.$$Wald's Identity
For SPRT: $E(\sum_{i=1}^N Z_i)=E(N)E(Z),$ where $Z_i=\log[f_1(X_i)/f_0(X_i)].$🌍 Where it's used in real life
- Acceptance sampling that can stop early.
- Online A/B tests that end as soon as clear.
- Clinical trials with early stopping for safety.
- Real-time fault detection.
- Inspection that minimises the number sampled.
7. Chi-square Tests
Goodness of Fit
Observed $O_i$ vs expected $E_i$ for $k$ categories: $$\chi^2=\sum_{i=1}^k\frac{(O_i-E_i)^2}{E_i}\sim\chi^2_{k-1-r},$$ where $r$ = number of parameters estimated from data.Independence in $r\times c$ Contingency Table
$E_{ij}=R_iC_j/N$: $$\chi^2=\sum\sum\frac{(O_{ij}-E_{ij})^2}{E_{ij}}\sim\chi^2_{(r-1)(c-1)}.$$Homogeneity
Same statistic; tests whether several populations have the same distribution across $c$ categories.Yates' Continuity Correction
For $2\times 2$ tables: $\chi^2=\sum\frac{(|O-E|-0.5)^2}{E}.$🌍 Where it's used in real life
- Is a die or roulette wheel fair? (goodness of fit).
- Gender vs product preference (independence).
- Survey response vs region.
- Genetics — observed vs expected ratios.
- Comparing defect patterns across machines.
8. Sign Test
For $H_0$: median = $m_0$. Let $S^+=\#(X_i>m_0),\,S^-=\#(X_i<m_0).$ Under $H_0$, $S^+\sim\text{Bin}(n,1/2)$ (after deleting ties).
Two-sample Paired Sign Test
Apply to differences $D_i=X_i-Y_i.$🌍 Where it's used in real life
- Did scores improve after training (before/after)?
- Taste test: do people prefer A or B?
- Is the median house price above a claim?
- A quick paired test needing no normality.
- Did a treatment change symptoms?
9. Wilcoxon Signed Rank Test
For symmetric distribution; tests $H_0:$ median = $m_0.$
- Compute $D_i=X_i-m_0.$
- Rank $|D_i|$, attach signs.
- $W^+=\sum$ ranks of positive $D_i.$
🌍 Where it's used in real life
- Before/after weight change on non-normal data.
- Pain scores before and after treatment.
- Matched-pair product comparisons.
- Median tests on small samples.
- Rating changes after a policy.
10. Mann–Whitney U Test (Wilcoxon Rank-Sum)
Two independent samples of sizes $m,n.$ Combine, rank from 1 to $m+n.$ Let $R_X$ = sum of ranks of $X$-sample. $$U_X=R_X-\tfrac{m(m+1)}{2},\,U_Y=mn-U_X.$$ Under $H_0$ (same distribution): $E(U)=mn/2,\,V(U)=mn(m+n+1)/12.$ Large samples: $z=(U-mn/2)/\sqrt{V}\xrightarrow{d}N(0,1).$
🌍 Where it's used in real life
- Comparing two groups' incomes (skewed data).
- Test scores from two teaching methods.
- Recovery times: drug vs placebo (non-normal).
- Satisfaction across two stores.
- Comparing rankings of two products.
11. Linear Rank Tests for Location and Scale
General form: $T=\sum_{i=1}^n c_i a(R_i),$ with constants $c_i$ and score function $a(\cdot).$
Common Scores
- Wilcoxon scores: $a(i)=i$ (location).
- Median scores: $a(i)=I(i>(N+1)/2)$ (location).
- van der Waerden / normal scores: $a(i)=\Phi^{-1}(i/(N+1))$.
- Ansari–Bradley: $a(i)=|i-(N+1)/2|$ (scale).
- Mood test: $a(i)=(i-(N+1)/2)^2$ (scale).
Asymptotic Distribution
$T$ is asymptotically normal under $H_0$, with $E$ and $V$ in closed form depending on scores.🌍 Where it's used in real life
- Detecting shifts in location without assuming normality.
- Comparing the spread of two processes (Ansari–Bradley).
- Robust two-sample comparisons.
- Comparing environmental measurements.
- Reliability data with outliers.
12. Kruskal–Wallis Test
Nonparametric analogue of one-way ANOVA for $k$ samples; tests $H_0$: all $k$ populations have the same continuous distribution.
Statistic
Combine all $N=\sum n_i$ observations and rank. $R_i$ = sum of ranks in group $i.$ $$H=\frac{12}{N(N+1)}\sum_{i=1}^k\frac{R_i^2}{n_i}-3(N+1).$$ Under $H_0$: $H\xrightarrow{d}\chi^2_{k-1}$ for $n_i\to\infty.$Tie Correction
$H_{\text{adj}}=H/[1-\frac{\sum(t_j^3-t_j)}{N^3-N}].$🌍 Where it's used in real life
- Comparing 3+ groups on satisfaction ratings.
- Yields across several varieties (non-normal).
- Salaries across departments.
- Pain relief across several drugs.
- Ratings across multiple websites.