Skip to the content

Topics Covered

Test Function Simple vs Composite Karlin–Rubin Theorem Wilks' Theorem Important LRTs Operating Characteristic (OC) Average Sample Number (ASN) Wald's Identity Goodness of Fit Independence in r× c Contingency Table Homogeneity Yates' Continuity Correction

Topic Overview — What & Why

Unit V develops the formal framework for deciding between competing claims about a population using sample data. Estimation tells us what the parameter is; testing tells us whether a hypothesised value is plausible.

  • Basic concepts: hypotheses, type I/II errors, size, power, $p$-value — the vocabulary every test-of-significance shares.
  • Neyman-Pearson lemma: for simple-vs-simple testing, the likelihood-ratio test is most powerful. Foundation of optimal testing.
  • Monotone likelihood ratio & Karlin-Rubin: when the likelihood ratio is monotone in a sufficient statistic, UMP tests for one-sided composite hypotheses exist.
  • UMP, UMPU, UMPI tests: hierarchy of optimality concepts — UMP is rare, UMPU/UMPI broaden the class of "best" tests when UMP fails (e.g., two-sided tests).
  • Likelihood ratio test (LRT): general-purpose construction; under regularity, $-2\log\Lambda$ is asymptotically $\chi^2$ (Wilks).
  • Wald's SPRT: sequential testing — sample size random; on average uses fewer observations than fixed-sample tests for the same error rates.
  • Chi-square tests: goodness-of-fit, independence, homogeneity. The workhorse of categorical data analysis.
  • Sign & Wilcoxon signed-rank: distribution-free alternatives to the one-sample $t$-test for medians.
  • Mann-Whitney U: nonparametric two-sample test; robust to non-normality and outliers.
  • Linear rank tests: general framework $T=\sum c_i a(R_i)$ — Wilcoxon, normal scores, Ansari-Bradley, Mood are all special cases.
  • Kruskal-Wallis: distribution-free analogue of one-way ANOVA for $k>2$ groups.

Where the derivations are

This unit is written to be read straight through as exam preparation: the statement, the condition, a worked example, and the MCQs at the end. The proofs behind it are written out elsewhere on this site, one step at a time, and they are worth reading once before you rely on the summaries here.

Where this unit is examined by other papers, see the ISS Paper II map and the CSIR NET map.

1. Basic Concepts

Why this section? Without precise definitions of error rates and power, all subsequent tests would be "shooting in the dark". Mastery of these terms makes every subsequent test transparent.

Null hypothesis $H_0$ vs alternative $H_1$.
Type I error: reject $H_0$ when true. Probability $\alpha$ (size).
Type II error: accept $H_0$ when false. Probability $\beta$.
Power: $1-\beta=P(\text{reject }H_0\mid H_1).$
p-value: probability under $H_0$ of obtaining a test statistic as or more extreme than observed.

Common misconception. The $p$-value is not $P(H_0\text{ is true})$ and it is not the probability the result is due to chance. It is computed assuming $H_0$ holds: it answers “if $H_0$ were true, how surprising would data this extreme be?” A small $p$-value means the observed data sit far in the tail of the null distribution.

critical value c H₀ true H₁ true α β power = 1−β accept H₀ reject H₀ → Type I (α) and Type II (β) errors
Two error types. Rejecting H₀ when it is true has probability α (red); failing to reject when H₁ holds has probability β (grey). Moving c trades α against β; the test's power is 1−β (orange).

Test Function

A test is a function $\phi(x)\in[0,1]$: probability of rejection given data. $E_\theta[\phi(X)]=$ power function.

Simple vs Composite

$H_0:\theta=\theta_0$ is simple; $H_0:\theta\le\theta_0$ is composite.
EXAMPLE 1 $X\sim N(\mu,1),n=25.$ Test $H_0:\mu=0$ vs $H_1:\mu>0$ at $\alpha=0.05.$ Reject if $\bar X>z_{0.05}/\sqrt n=0.329.$ Power at $\mu=0.5$: $P(\bar X>0.329\mid\mu=0.5)=P(Z>(0.329-0.5)\cdot 5)=P(Z>-0.855)\approx 0.804.$
EXAMPLE 2 Coin: 100 tosses, 60 heads. Test $H_0:p=0.5$ vs $H_1:p\ne 0.5.$ $z=(0.6-0.5)/\sqrt{0.0025}=2.0$, $p$-value $=2(1-\Phi(2))=0.0455$ → reject at $\alpha=0.05.$

🌍 Where it's used in real life

  1. Deciding whether a new drug really works.
  2. Checking if a coin or machine is fair.
  3. Did a website change lift sales?
  4. Is average wait time above target?
  5. Detecting bias in a process.

2. Neyman–Pearson Lemma

For testing simple $H_0:\theta=\theta_0$ vs simple $H_1:\theta=\theta_1$ at size $\alpha$, the most powerful (MP) test has rejection region $$R=\left\{x:\frac{L(x;\theta_1)}{L(x;\theta_0)}\ge k\right\},$$ with $k$ chosen so that $P_{\theta_0}(R)=\alpha.$

Intuition. To maximise power at a fixed size, spend your allowed Type-I “budget” on the sample outcomes where the data are relatively far more likely under $H_1$ than under $H_0$. Ranking outcomes by the likelihood ratio and rejecting the top ones until the size reaches $\alpha$ is provably the most powerful strategy — every unit of $\alpha$ is bought where it buys the most power.

EXAMPLE 1 $X_i\sim N(\mu,1).$ $H_0:\mu=0$ vs $H_1:\mu=1.$ $L_1/L_0=\exp\{\sum X_i-n/2\}\ge k\iff\bar X\ge c.$ MP test rejects when $\bar X\ge z_\alpha/\sqrt n.$
EXAMPLE 2 Bernoulli: $H_0:p=0.5,H_1:p=0.7,n=10.$ $L_1/L_0=(0.7/0.5)^{\sum X}(0.3/0.5)^{n-\sum X}.$ MP test based on $\sum X\ge c.$

🌍 Where it's used in real life

  1. Radar: is a signal present or absent?
  2. Setting a medical test's cut-off.
  3. Optimal spam vs not-spam rule.
  4. Accept/reject decisions in quality control.
  5. Fraud-detection thresholds.

3. Monotone Likelihood Ratio (MLR)

A family $\{f(x;\theta)\}$ has MLR in $T(x)$ if for $\theta_1<\theta_2$, $f(x;\theta_2)/f(x;\theta_1)$ is a non-decreasing function of $T(x).$

One-parameter exponential family has MLR in the natural sufficient statistic.

Karlin–Rubin Theorem

If MLR in $T$, the UMP test for $H_0:\theta\le\theta_0$ vs $H_1:\theta>\theta_0$ rejects when $T\ge c.$
EXAMPLE 1 Exp$(\lambda)$ has MLR in $-\sum X_i$ (or $\sum X_i$ for $1/\lambda$). UMP test for $H_0:\lambda\le\lambda_0$ vs $H_1:\lambda>\lambda_0$ rejects when $\sum X_i\le c.$
EXAMPLE 2 Bin$(n,p)$ has MLR in $\sum X_i.$ UMP test for $H_0:p\le p_0$ vs $H_1:p>p_0$ rejects when $\sum X_i\ge c.$

🌍 Where it's used in real life

  1. One-sided quality tests (is the defect rate too high?).
  2. Is mean strength above the specification?
  3. Detecting a rising failure rate.
  4. One-sided drug-improvement tests.
  5. Watching for increasing pollution levels.

4. UMP, UMPU, UMPI Tests

Two-sided tests $H_0:\theta=\theta_0$ vs $H_1:\theta\ne\theta_0$ generally have no UMP; UMPU exists.

EXAMPLE 1 $N(\mu,1)$, two-sided test $H_0:\mu=\mu_0$ vs $H_1:\mu\ne\mu_0.$ UMPU: reject if $|\bar X-\mu_0|>z_{\alpha/2}/\sqrt n.$
EXAMPLE 2 $N(\mu,\sigma^2)$ with $\sigma$ unknown: $t$-test is UMPU for $\mu$ and UMPI under translation–scale group.

🌍 Where it's used in real life

  1. Picking the most powerful test available.
  2. Two-sided tests around a target value.
  3. Tests that don't depend on the units used.
  4. Best process-monitoring tests.
  5. Optimal clinical-trial test design.

5. Likelihood Ratio Test (LRT)

$$\Lambda(x)=\frac{\sup_{\theta\in\Theta_0}L(\theta;x)}{\sup_{\theta\in\Theta}L(\theta;x)}.$$ Reject $H_0$ when $\Lambda\le c$ (or $-2\log\Lambda\ge\chi^2$ critical).

Wilks' Theorem

Under regularity, $-2\log\Lambda\xrightarrow{d}\chi^2_r$ under $H_0$, where $r=\dim\Theta-\dim\Theta_0.$

Important LRTs

EXAMPLE 1 $X\sim N(\mu,\sigma^2)$ both unknown. $H_0:\mu=\mu_0$ gives LRT statistic monotone in $|t|=|\bar X-\mu_0|/(S/\sqrt n)$ — equivalent to $t$-test.
EXAMPLE 2 For two normal samples with $H_0:\sigma_1=\sigma_2$, LRT yields $F=S_1^2/S_2^2\sim F_{n_1-1,n_2-1}$ (under $H_0$).

🌍 Where it's used in real life

  1. Comparing nested regression models.
  2. Testing genetic linkage models.
  3. Testing equality of several means.
  4. Model selection in machine learning.
  5. Testing whether data fit a distribution.

6. Wald's Sequential Probability Ratio Test (SPRT)

Test $H_0:\theta=\theta_0$ vs $H_1:\theta=\theta_1$ sequentially. After observing $X_1,\ldots,X_n$, compute likelihood ratio: $$\lambda_n=\frac{L_1}{L_0}=\prod_{i=1}^n\frac{f(x_i;\theta_1)}{f(x_i;\theta_0)}.$$ Define bounds $A=(1-\beta)/\alpha,\,B=\beta/(1-\alpha)$:

Operating Characteristic (OC)

$L(\theta)=P_\theta(\text{accept }H_0)$ — Wald's approximation: $$L(\theta)\approx\frac{A^{h(\theta)}-1}{A^{h(\theta)}-B^{h(\theta)}}.$$

Average Sample Number (ASN)

$$E_\theta(N)\approx\frac{L(\theta)\log B+(1-L(\theta))\log A}{E_\theta\log[f(X;\theta_1)/f(X;\theta_0)]}.$$

Wald's Identity

For SPRT: $E(\sum_{i=1}^N Z_i)=E(N)E(Z),$ where $Z_i=\log[f_1(X_i)/f_0(X_i)].$
EXAMPLE 1 Bernoulli, $H_0:p=0.3$ vs $H_1:p=0.5,\alpha=\beta=0.05.$ $A=19,B=1/19.$ Continue while $1/19<\lambda_n<19.$
EXAMPLE 2 SPRT for normal mean ($\sigma$ known) reduces to $\sum X_i$ vs linear bounds in $n$ — random walk crossing boundaries.

🌍 Where it's used in real life

  1. Acceptance sampling that can stop early.
  2. Online A/B tests that end as soon as clear.
  3. Clinical trials with early stopping for safety.
  4. Real-time fault detection.
  5. Inspection that minimises the number sampled.

7. Chi-square Tests

Goodness of Fit

Observed $O_i$ vs expected $E_i$ for $k$ categories: $$\chi^2=\sum_{i=1}^k\frac{(O_i-E_i)^2}{E_i}\sim\chi^2_{k-1-r},$$ where $r$ = number of parameters estimated from data.

Independence in $r\times c$ Contingency Table

$E_{ij}=R_iC_j/N$: $$\chi^2=\sum\sum\frac{(O_{ij}-E_{ij})^2}{E_{ij}}\sim\chi^2_{(r-1)(c-1)}.$$

Homogeneity

Same statistic; tests whether several populations have the same distribution across $c$ categories.

Yates' Continuity Correction

For $2\times 2$ tables: $\chi^2=\sum\frac{(|O-E|-0.5)^2}{E}.$
EXAMPLE 1 Die rolls 600 times: observed $(95,110,100,90,105,100).$ Expected 100 each. $\chi^2=\frac{25+100+0+100+25+0}{100}=2.5,\,df=5.$ Not significant.
EXAMPLE 2 $2\times 2$ table — drug effect: $a=40,b=10,c=20,d=30.$ $\chi^2=\frac{N(ad-bc)^2}{(a+b)(c+d)(a+c)(b+d)}=\frac{100(1200-200)^2}{50\cdot 50\cdot 60\cdot 40}=\frac{10^8}{6\times 10^6}\approx 16.67$, $df=1$, highly significant.

🌍 Where it's used in real life

  1. Is a die or roulette wheel fair? (goodness of fit).
  2. Gender vs product preference (independence).
  3. Survey response vs region.
  4. Genetics — observed vs expected ratios.
  5. Comparing defect patterns across machines.

8. Sign Test

For $H_0$: median = $m_0$. Let $S^+=\#(X_i>m_0),\,S^-=\#(X_i<m_0).$ Under $H_0$, $S^+\sim\text{Bin}(n,1/2)$ (after deleting ties).

Two-sample Paired Sign Test

Apply to differences $D_i=X_i-Y_i.$
EXAMPLE 1 $n=12,$ 9 differences positive, 3 negative. $P(S^+\ge 9\mid p=0.5)=\binom{12}{9}(0.5)^{12}+\cdots\approx 0.073.$ Two-sided $p=0.146$ — fail to reject at $5\%.$
EXAMPLE 2 For $n=20$ with $S^+=15$: $z=\frac{15-10}{\sqrt 5}=2.236,p\approx 0.025$ (two-sided): significant.

🌍 Where it's used in real life

  1. Did scores improve after training (before/after)?
  2. Taste test: do people prefer A or B?
  3. Is the median house price above a claim?
  4. A quick paired test needing no normality.
  5. Did a treatment change symptoms?

9. Wilcoxon Signed Rank Test

For symmetric distribution; tests $H_0:$ median = $m_0.$

  1. Compute $D_i=X_i-m_0.$
  2. Rank $|D_i|$, attach signs.
  3. $W^+=\sum$ ranks of positive $D_i.$
Under $H_0$: $E(W^+)=n(n+1)/4,\,V(W^+)=n(n+1)(2n+1)/24.$ Use normal approx for $n\ge 20$ or table for small $n.$
EXAMPLE 1 $n=10,W^+=45.$ $E=27.5,V=96.25,SD=9.81.$ $z=(45-27.5)/9.81=1.78,$ two-sided $p\approx 0.075.$
EXAMPLE 2 For paired data $(X_i,Y_i)$, set $D_i=X_i-Y_i$ and apply Wilcoxon to test $H_0:$ median of $D=0.$

🌍 Where it's used in real life

  1. Before/after weight change on non-normal data.
  2. Pain scores before and after treatment.
  3. Matched-pair product comparisons.
  4. Median tests on small samples.
  5. Rating changes after a policy.

10. Mann–Whitney U Test (Wilcoxon Rank-Sum)

Two independent samples of sizes $m,n.$ Combine, rank from 1 to $m+n.$ Let $R_X$ = sum of ranks of $X$-sample. $$U_X=R_X-\tfrac{m(m+1)}{2},\,U_Y=mn-U_X.$$ Under $H_0$ (same distribution): $E(U)=mn/2,\,V(U)=mn(m+n+1)/12.$ Large samples: $z=(U-mn/2)/\sqrt{V}\xrightarrow{d}N(0,1).$

EXAMPLE 1 $m=5,n=5,$ all $X_i$ smaller than all $Y_i$: ranks 1-5 for $X$, sum $R_X=15.$ $U_X=15-15=0$ — extreme; reject $H_0.$
EXAMPLE 2 $m=20,n=25,U=160.$ $E=250,V=20\cdot 25\cdot 46/12=1916.67,SD=43.78.$ $z=(160-250)/43.78=-2.06,p\approx 0.04.$

🌍 Where it's used in real life

  1. Comparing two groups' incomes (skewed data).
  2. Test scores from two teaching methods.
  3. Recovery times: drug vs placebo (non-normal).
  4. Satisfaction across two stores.
  5. Comparing rankings of two products.

11. Linear Rank Tests for Location and Scale

General form: $T=\sum_{i=1}^n c_i a(R_i),$ with constants $c_i$ and score function $a(\cdot).$

Common Scores

Asymptotic Distribution

$T$ is asymptotically normal under $H_0$, with $E$ and $V$ in closed form depending on scores.
EXAMPLE 1 Wilcoxon rank-sum: location-shift alternative, scores $a(i)=i.$
EXAMPLE 2 Ansari–Bradley: tests $H_0:\sigma_X=\sigma_Y$ when medians equal; assigns higher scores to extreme ranks.

🌍 Where it's used in real life

  1. Detecting shifts in location without assuming normality.
  2. Comparing the spread of two processes (Ansari–Bradley).
  3. Robust two-sample comparisons.
  4. Comparing environmental measurements.
  5. Reliability data with outliers.

12. Kruskal–Wallis Test

Nonparametric analogue of one-way ANOVA for $k$ samples; tests $H_0$: all $k$ populations have the same continuous distribution.

Statistic

Combine all $N=\sum n_i$ observations and rank. $R_i$ = sum of ranks in group $i.$ $$H=\frac{12}{N(N+1)}\sum_{i=1}^k\frac{R_i^2}{n_i}-3(N+1).$$ Under $H_0$: $H\xrightarrow{d}\chi^2_{k-1}$ for $n_i\to\infty.$

Tie Correction

$H_{\text{adj}}=H/[1-\frac{\sum(t_j^3-t_j)}{N^3-N}].$
EXAMPLE 1 3 groups, $n_1=n_2=n_3=4,N=12.$ $R_1=10,R_2=20,R_3=48.$ $H=\frac{12}{12\cdot 13}\Big(\frac{100}{4}+\frac{400}{4}+\frac{2304}{4}\Big)-3\cdot 13=\frac{1}{13}(701)-39=53.92-39=14.92.\,p<0.001.$
EXAMPLE 2 For $k=2$: Kruskal–Wallis equivalent to Mann–Whitney; $H=Z^2.$

🌍 Where it's used in real life

  1. Comparing 3+ groups on satisfaction ratings.
  2. Yields across several varieties (non-normal).
  3. Salaries across departments.
  4. Pain relief across several drugs.
  5. Ratings across multiple websites.