Hypothesis testing is a process of testing the significance regarding a population parameter on the basis of a sample. A statistic is computed from the sample observations, and on the strength of that statistic we judge whether the sample was drawn from a parent population with certain specified characteristics.
The computed value of the statistic will almost never equal the hypothetical value of the parameter. The whole question is how to read that gap:
So testing of hypothesis is the procedure of examining whether the difference between the computed statistic (from the sample) and the hypothetical parameter (from the population) is significant or not.
Another definition. Testing of hypothesis is a process of decision making, using various statistical methods and the theory of modern probability.
A statistical hypothesis is a certain statement about the probability distribution of a random variable — equivalently, a certain statement about the population. A statistical hypothesis is denoted by \(H\).
The printed page states example 5 as "the variation … is 10 hrs" and then writes \(\sigma_1^2-\sigma_2^2=2\). The prose and the formula disagree; the formula is the form the rest of the chapter uses, so the sentence is given as 2 above.
Simple statistical hypothesis. If a statistical hypothesis specifies the population completely — that is, the probability distribution is fully known — it is called a simple statistical hypothesis. In a simple hypothesis all the parameters of the population are given clear values.
Examples. Let \(X\sim N(\mu,\sigma^2)\), then \(H:\mu = 1200\) kms, \(\sigma^2 = 2\). Let \(X\sim B(n,P)\), then \(H:P=\tfrac13\).
Composite statistical hypothesis. If a statistical hypothesis does not specify the population completely, it is called a composite statistical hypothesis.
Examples. \(H:\mu < 20\); \(H:\mu_1 > \mu_2\); \(H:\sigma^2 = 2,\ \mu > 2\).
The third one is composite even though it pins \(\sigma^2\) down: leaving any parameter unspecified is enough.
There are two kinds of hypothesis essential to conducting a test procedure.
1. Null hypothesis. A statistical hypothesis with no difference, or with a null attitude, is called a null hypothesis. It is denoted by \(H_0\).
According to R. A. Fisher: "Null hypothesis is the hypothesis which is tested for possible rejection under the assumption that it is true."
Examples. The average height of the competitors in a game is 160 cms, i.e. \(H_0:\mu=160\) cms. The average life time of electrical bulbs manufactured by a company is 1800 hours, i.e. \(H_0:\mu=1800\) hrs.
2. Alternative hypothesis. A statistical hypothesis which is complementary to the null hypothesis is called an alternative hypothesis. It is denoted by \(H_1\).
It is clear that a null hypothesis is meaningful only when an alternative hypothesis has been formulated alongside it.
1. If the null hypothesis is that the average height of the competitors in a game is 160 cms, i.e. \(H_0:\mu=160\) cms, then the alternative hypothesis may be formulated as
The alternative hypothesis in (i) gives a two-tailed test; those in (ii) and (iii) give one-tailed tests.
2. If the null hypothesis is that the average life time of electrical bulbs in a company is 1800 hours, the alternative hypothesis may be considered as follows:
| Form | Type | Example |
|---|---|---|
| \(H_1: \theta \ne \theta_0\) | Two-tailed | μ ≠ 50 |
| \(H_1: \theta > \theta_0\) | Right-tailed | μ > 50 |
| \(H_1: \theta < \theta_0\) | Left-tailed | μ < 50 |
The critical region \(W\) is the set of sample-space outcomes for which we reject \(H_0\). Its complement is the acceptance region.
For a continuous test statistic \(T\) with a critical value \(c\):
| H₀ True | H₀ False | |
|---|---|---|
| Reject H₀ | Type I error (probability α) | Correct decision (probability 1−β) |
| Accept H₀ | Correct decision (probability 1−α) | Type II error (probability β) |
Reducing α typically increases β and vice versa. Increasing sample size \(n\) is the proper way to reduce both.
A decision — whether \(H_0\) is to be accepted or rejected — is made on the information supplied by the sample data, so there is always a chance of a good decision or of an error. Written over the regions of §2:
Type I error is the error of rejecting \(H_0\) when \(H_0\) is true:
\[ \alpha = P(\text{Type I error}) = P(\text{rejecting } H_0 \text{ when } H_0 \text{ is true}) = P(x\in W \mid H_0) = \int_{W} L_0\,dx, \]where \(L_0\) is the likelihood function of the sample observations \(x_1,x_2,\dots,x_n\) under \(H_0\).
Type II error is the error of accepting \(H_0\) when \(H_0\) is false:
\[ \beta = P(\text{Type II error}) = P(\text{accepting } H_0 \text{ when } H_0 \text{ is false}) = P(x\in \bar W \mid H_1) = \int_{\bar W} L_1\,dx, \]where \(L_1\) is the likelihood function under \(H_1\). The two errors are integrals of different likelihoods over complementary regions, which is why shrinking one enlarges the other.
The printed page writes this second integral as \(\int_{\bar W} L_0\,dx\), then says on the very next line that \(L_1\) is the likelihood under \(H_1\). It is \(L_1\), as above.
The probability of a Type I error, \(\alpha\), is known as the level of significance. It is also called the size of the critical region — a name worth keeping in mind, because it says plainly that choosing \(\alpha\) is choosing how big \(W\) is allowed to be.
In statistical quality control the two errors have names taken from who pays for them: \(\alpha\) is the producer's risk (good batches rejected) and \(\beta\) is the consumer's risk (bad batches accepted).
A factory's quality engineer rejects a defect-free batch (Type I error) — money lost on rework. With α = 0.05, this happens 5 % of the time even though the batch was fine.
A medical test fails to detect a sick patient (Type II error). If β = 0.20, only 80 % of truly ill patients are correctly identified — power = 0.80.
The p-value is the probability, assuming \(H_0\) is true, of observing a test statistic as extreme as (or more extreme than) the one observed.
For an observed Z = 2.10 in a two-tailed test, p = 2(1 − Φ(2.10)) = 2(0.0179) = 0.0358. Since p < 0.05, reject \(H_0\) at 5 %.
p = 0.18 in any test ⇒ data are consistent with \(H_0\); no significant evidence to reject.
Power = \(1 - \beta\) = P(reject \(H_0\) | \(H_1\) true) — the probability of correctly detecting an effect.
Power function \(\pi(\theta) = P(\text{reject } H_0 \mid \theta)\) is a function of the true \(\theta\). At \(\theta = \theta_0\), \(\pi = \alpha\); for \(\theta\) far from \(\theta_0\), \(\pi \to 1\).
In the notation of §2 and §3,
\[ 1-\beta = P(x\in W \mid H_1) \]is the probability of rejecting \(H_0\) when \(H_0\) is false. This is called the power function of testing the hypothesis, and its value is the power of the test. Compare it with \(\alpha=\int_W L_0\,dx\): the same region \(W\), scored under the other hypothesis.
Take the problem of testing a simple null hypothesis \(H_0:\theta=\theta_0\) against a simple alternative \(H_1:\theta=\theta_1\). The critical region \(W\) is the most powerful critical region of size \(\alpha\) for testing \(H_0\) against \(H_1\) if
\[ P(x\in W \mid H_0) = \int_{W} L_0\,dx = \alpha \tag{1} \]and
\[ P(x\in W \mid H_1) \;\ge\; P(x\in W_1 \mid H_1) \tag{2} \]for every other critical region \(W_1\) satisfying (1). The corresponding test is called the most powerful test.
Read plainly: among all regions that make the same number of Type I errors, \(W\) catches the most false nulls. §7 names that region.
| Aspect | One-tailed | Two-tailed |
|---|---|---|
| Form of \(H_1\) | \(\theta > \theta_0\) or \(\theta < \theta_0\) | \(\theta \ne \theta_0\) |
| Critical region | One side of distribution | Both sides |
| Critical value at α = 0.05 (Z) | 1.645 | 1.96 |
| Use | Direction known a priori | Direction unknown / either side matters |
Let \(x_1,x_2,\dots,x_n\) be a random sample of size \(n\) drawn from a population with density function \(f(x,\theta)\). Let \(K>0\) be a constant, and let \(W\) be the most powerful critical region of size \(\alpha\) for testing a simple null hypothesis \(H_0:\theta=\theta_0\) against a simple alternative \(H_1:\theta=\theta_1\), such that
\[ W = \left\{ x\in S : \frac{f(x,\theta_1)}{f(x,\theta_0)} > K \right\} \]i.e.
\[ W = \left\{ x\in S : \frac{L_1}{L_0} > K \right\} \tag{I} \]and
\[ \bar W = \left\{ x\in S : \frac{L_1}{L_0} \le K \right\} \tag{II} \]where \(L_0\) and \(L_1\) are the likelihood functions of the sample observations \(x_1,x_2,\dots,x_n\) under \(H_0\) and \(H_1\) respectively.
The lemma is the foundation of optimal hypothesis testing: it does not merely offer a good test, it says that no test of the same size can do better. For composite hypotheses, generalised likelihood ratio tests (GLRT, §8) extend the idea.
The size of the critical region is
\[ P(x\in W \mid H_0) = \int_{W} L_0\,dx = \alpha \tag{1} \]and the power of the region is
\[ P(x\in W \mid H_1) = \int_{W} L_1\,dx = 1-\beta . \tag{2} \]To prove the lemma we have to show that there exists no other critical region, of size less than or equal to \(\alpha\), which is more powerful than \(W\). Let \(W_1\) be another critical region, of size \(\alpha_1\le\alpha\) and power \(1-\beta_1\):
\[ P(x\in W_1 \mid H_0) = \int_{W_1} L_0\,dx = \alpha_1 \tag{3} \] \[ P(x\in W_1 \mid H_1) = \int_{W_1} L_1\,dx = 1-\beta_1 . \tag{4} \]We have to prove that the power is greater for \(W\), i.e. that \(1-\beta\) is the larger.
Let \(W = A\cup C\) and \(W_1 = B\cup C\). Consider
\[ \alpha_1 \le \alpha \;\Longrightarrow\; \int_{W_1} L_0\,dx \le \int_{W} L_0\,dx \] \[ \Longrightarrow\; \int_{B\cup C} L_0\,dx \le \int_{A\cup C} L_0\,dx \;\Longrightarrow\; \int_{B} L_0\,dx \le \int_{A} L_0\,dx \] \[ \Longrightarrow\; \int_{A} L_0\,dx \ge \int_{B} L_0\,dx . \tag{5} \]From (I) in the statement, for \(x\in W\) we have
\[ \frac{L_1}{L_0} > K \;\Longrightarrow\; L_1 > K L_0 \;\Longrightarrow\; \int_{W} L_1\,dx > K\int_{W} L_0\,dx . \]Since \(A\subset W\),
\[ \int_{A} L_1\,dx > K\int_{A} L_0\,dx . \tag{6} \]Multiplying equation (5) by \(K\),
\[ K\int_{A} L_0\,dx \ge K\int_{B} L_0\,dx . \tag{7} \]From (6) and (7) we get
\[ \int_{A} L_1\,dx > K\int_{B} L_0\,dx \qquad\text{i.e.}\qquad K\int_{B} L_0\,dx \le \int_{A} L_1\,dx . \tag{8} \]From (II) in the statement, for \(x\in\bar W\) we have
\[ \frac{L_1}{L_0} \le K \;\Longrightarrow\; L_1 \le K L_0 \;\Longrightarrow\; \int_{\bar W} L_1\,dx \le K\int_{\bar W} L_0\,dx . \]Since \(B\subset\bar W\),
\[ \int_{B} L_1\,dx \le K\int_{B} L_0\,dx . \tag{9} \]From (8) and (9) we get
\[ \int_{B} L_1\,dx \le \int_{A} L_1\,dx . \]By adding \(\int_{C} L_1\,dx\) to both sides,
\[ \int_{B\cup C} L_1\,dx \le \int_{A\cup C} L_1\,dx \;\Longrightarrow\; \int_{W_1} L_1\,dx \le \int_{W} L_1\,dx \] \[ \Longrightarrow\; 1-\beta_1 \le 1-\beta \qquad\text{i.e.}\qquad 1-\beta \ge 1-\beta_1 . \quad\blacksquare \]Hence no critical region of size \(\le\alpha\) has greater power than \(W\), which is the lemma.
One transcription note: the printed page writes step (9) as \(\int_{B} L_1\,dx \le K\int_{P} L_0\,dx\). There is no region \(P\) in the proof; the subscript is a smudged \(B\), as written above.
Four applications of the lemma. Each follows the same three moves: write \(L_1/L_0\), take logarithms to turn the product into a sum in \(\sum x_i\), then divide — and watch the inequality flip when the divisor is negative. That flip is the whole reason each problem has two cases.
Obtain the best critical region for testing \(H_0:p=p_0\) against \(H_1:p=p_1\).
Solution. Let \(x_1,x_2,\dots,x_m\) be a random sample of size \(m\) from a binomial population, whose probability mass function is
\[ f(x;n,p) = {}^{n}C_{x}\,p^{x}q^{\,n-x},\quad x=0,1,2,\dots,n, \qquad q = 1-p. \]The likelihood function of \(x_1,x_2,\dots,x_m\) is
\[ L = f(x_1;n,p)\,f(x_2;n,p)\cdots f(x_m;n,p) = \left({}^{n}C_{x_1}\,{}^{n}C_{x_2}\cdots{}^{n}C_{x_m}\right) p^{\sum_{i=1}^{m} x_i}\,(1-p)^{\,mn-\sum_{i=1}^{m} x_i}. \]Hence, under the two hypotheses,
\[ L_1 = \left({}^{n}C_{x_1}\cdots{}^{n}C_{x_m}\right) p_1^{\sum x_i}(1-p_1)^{\,mn-\sum x_i}, \qquad L_0 = \left({}^{n}C_{x_1}\cdots{}^{n}C_{x_m}\right) p_0^{\sum x_i}(1-p_0)^{\,mn-\sum x_i}. \]Using the N–P lemma the best critical region is obtained from \(L_1/L_0 \ge K\):
\[ \left(\frac{p_1}{p_0}\right)^{\sum_{i=1}^{m} x_i} \left(\frac{1-p_1}{1-p_0}\right)^{\,mn-\sum_{i=1}^{m} x_i} \ge K \]Taking logarithms,
\[ \sum_{i=1}^{m} x_i \log\!\left(\frac{p_1}{p_0}\right) + \left(mn - \sum_{i=1}^{m} x_i\right)\log\!\left(\frac{1-p_1}{1-p_0}\right) \ge \log K = K_1 \ (\text{say}) \] \[ \Longrightarrow\; \sum_{i=1}^{m} x_i\left(\log\frac{p_1}{p_0} - \log\frac{1-p_1}{1-p_0}\right) \ge K_1 - mn\log\frac{1-p_1}{1-p_0} = K_2 \ (\text{say}) \] \[ \Longrightarrow\; m\bar x\left(\log\frac{p_1}{p_0} - \log\frac{1-p_1}{1-p_0}\right) \ge K_2 \qquad\Longrightarrow\qquad \bar x \ge \frac{K_2}{m\left(\log\dfrac{p_1}{p_0} - \log\dfrac{1-p_1}{1-p_0}\right)} \]Case 1. If \(p_1 > p_0\), then \(\log\dfrac{p_1}{p_0} - \log\dfrac{1-p_1}{1-p_0} > 0\), so
\[ \frac{K_2}{m\left(\log\dfrac{p_1}{p_0}-\log\dfrac{1-p_1}{1-p_0}\right)} = K_3 \ (\text{say}) \]and the best critical region of size \(\alpha\) is \(\ W = \{x\in S : \bar x \ge K_3\}\).
Case 2. If \(p_1 < p_0\), then that bracket is negative, dividing by it reverses the inequality, and
\[ \frac{K_2}{m\left(\log\dfrac{p_1}{p_0}-\log\dfrac{1-p_1}{1-p_0}\right)} = K_4 \ (\text{say}) \]so the best critical region of size \(\alpha\) is \(\ W = \{x\in S : \bar x \le K_4\}\).
The sign of the constant is not known, and is not needed. \(K_2\) is an arbitrary constant, so nothing in the working makes \(K_3\) positive or \(K_4\) negative — and a negative \(K_4\) would make \(\{\bar x \le K_4\}\) empty, since \(\bar x \ge 0\) here. What each case fixes is the direction of the region. The constant itself comes from the size condition, \(P(\bar x \ge K_3 \mid H_0) = \alpha\) in Case 1 and \(P(\bar x \le K_4 \mid H_0) = \alpha\) in Case 2. The same holds for the three problems below.
The printed page opens Case 1 with "if \(p_1 \ge p_0\)". At \(p_1=p_0\) the bracket is exactly zero and the division is undefined, so the condition is strict, as written above. The same correction applies to Problem 3.
Obtain the best critical region for testing \(H_0:\lambda=\lambda_0\) against \(H_1:\lambda=\lambda_1\).
Solution. Let \(x_1,x_2,\dots,x_n\) be a random sample of size \(n\) drawn from a Poisson population, with probability mass function
\[ f(x,\lambda) = \frac{e^{-\lambda}\lambda^{x}}{x!}, \qquad x=0,1,2,\dots,\ \lambda>0. \]The likelihood function is
\[ L = f(x_1;\lambda)\,f(x_2;\lambda)\cdots f(x_n;\lambda) = \frac{e^{-n\lambda}\,\lambda^{\sum_{i=1}^{n} x_i}}{x_1!\,x_2!\cdots x_n!}, \]so under \(H_0:\lambda=\lambda_0\) and \(H_1:\lambda=\lambda_1\),
\[ L_0 = \frac{e^{-n\lambda_0}\,\lambda_0^{\sum x_i}}{x_1!\cdots x_n!}, \qquad L_1 = \frac{e^{-n\lambda_1}\,\lambda_1^{\sum x_i}}{x_1!\cdots x_n!}. \]The best critical region according to the N–P lemma is \(L_1/L_0 \ge K\):
\[ \frac{e^{-n\lambda_1}\lambda_1^{\sum x_i}}{e^{-n\lambda_0}\lambda_0^{\sum x_i}} \ge K \qquad\Longrightarrow\qquad e^{-n(\lambda_1-\lambda_0)}\left(\frac{\lambda_1}{\lambda_0}\right)^{\sum_{i=1}^{n} x_i} \ge K \] \[ \Longrightarrow\; \left(\frac{\lambda_1}{\lambda_0}\right)^{\sum x_i} \ge K\,e^{\,n(\lambda_1-\lambda_0)} = K_1 \ (\text{say}) \] \[ \Longrightarrow\; \sum_{i=1}^{n} x_i \log\!\left(\frac{\lambda_1}{\lambda_0}\right) \ge \log K_1 = K_2 \ (\text{say}) \qquad\Longrightarrow\qquad n\bar x\,(\log\lambda_1 - \log\lambda_0) \ge K_2 \] \[ \Longrightarrow\; \bar x \ge \frac{K_2}{n(\log\lambda_1 - \log\lambda_0)} \]Case 1. If \(\lambda_1 > \lambda_0\), then \(\dfrac{K_2}{n(\log\lambda_1-\log\lambda_0)} = K_3 \ (\text{say})\), and the best critical region of size \(\alpha\) is \(\ W = \{x\in S : \bar x \ge K_3\}\).
Case 2. If \(\lambda_1 < \lambda_0\), then \(\dfrac{K_2}{n(\log\lambda_1-\log\lambda_0)} = K_4 \ (\text{say})\), and the best critical region of size \(\alpha\) is \(\ W = \{x\in S : \bar x \le K_4\}\).
Obtain the best critical region for testing \(H_0:\mu=\mu_0\) against \(H_1:\mu=\mu_1\).
Solution. Let \(x_1,x_2,\dots,x_n\) be a random sample drawn from a normal population with probability density function
\[ f(x;\mu,\sigma^2) = \frac{1}{\sigma\sqrt{2\pi}}\, e^{-\frac12\left(\frac{x-\mu}{\sigma}\right)^{2}}, \qquad -\infty<x<\infty . \]The likelihood function is
\[ L = f(x_1;\mu,\sigma^2)\cdots f(x_n;\mu,\sigma^2) = \frac{1}{(\sigma\sqrt{2\pi})^{n}}\, e^{-\frac{1}{2\sigma^{2}}\sum_{i=1}^{n}(x_i-\mu)^{2}}, \]so
\[ L_0 = \frac{1}{(\sigma\sqrt{2\pi})^{n}}e^{-\frac{1}{2\sigma^{2}}\sum(x_i-\mu_0)^{2}}, \qquad L_1 = \frac{1}{(\sigma\sqrt{2\pi})^{n}}e^{-\frac{1}{2\sigma^{2}}\sum(x_i-\mu_1)^{2}}. \]Now the best critical region using the N–P lemma, \(L_1/L_0 \ge K\):
\[ e^{-\frac{1}{2\sigma^{2}}\left[\sum(x_i-\mu_1)^{2}-\sum(x_i-\mu_0)^{2}\right]} \ge K \] \[ \Longrightarrow\; \frac{-1}{2\sigma^{2}} \left[\sum_{i=1}^{n}(x_i-\mu_1)^{2}-\sum_{i=1}^{n}(x_i-\mu_0)^{2}\right] \ge \log K = K_1 \ (\text{say}) \]Multiply by \(-2\sigma^2\). That reverses the inequality, and writing the difference the other way round reverses it back, so the \(\ge\) survives:
\[ \sum_{i=1}^{n}(x_i-\mu_0)^{2} - \sum_{i=1}^{n}(x_i-\mu_1)^{2} \ge K_1(2\sigma^{2}) = K_2 \ (\text{say}) \]Expanding both squares, the \(\sum x_i^2\) terms cancel:
\[ \sum x_i^{2} - 2\mu_0\sum x_i + n\mu_0^{2} - \sum x_i^{2} + 2\mu_1\sum x_i - n\mu_1^{2} \ge K_2 \] \[ \Longrightarrow\; 2(\mu_1-\mu_0)\sum_{i=1}^{n} x_i - n(\mu_1^{2}-\mu_0^{2}) \ge K_2 \] \[ \Longrightarrow\; (\mu_1-\mu_0)\sum_{i=1}^{n} x_i \ge \frac{K_2 + n(\mu_1^{2}-\mu_0^{2})}{2} = K_3 \ (\text{say}) \] \[ \Longrightarrow\; (\mu_1-\mu_0)\,n\bar x \ge K_3 \qquad\Longrightarrow\qquad \bar x \ge \frac{K_3}{n(\mu_1-\mu_0)} \]Case 1. If \(\mu_1 > \mu_0\), then \(\dfrac{K_3}{n(\mu_1-\mu_0)} = K_4 \ (\text{say})\), and the best critical region of size \(\alpha\) is \(\ W = \{x\in S : \bar x \ge K_4\}\).
Case 2. If \(\mu_1 < \mu_0\), then \(\dfrac{K_3}{n(\mu_1-\mu_0)} = K_5 \ (\text{say})\), and the best critical region of size \(\alpha\) is \(\ W = \{x\in S : \bar x \le K_5\}\).
This is where the familiar \(Z\)-test comes from: the optimal test of a normal mean rejects for extreme \(\bar x\), and nothing else. Unit 3 supplies the constant.
Obtain the best critical region for testing \(H_0:\lambda=\lambda_0\) against \(H_1:\lambda=\lambda_1\).
Solution. Let \(x_1,x_2,\dots,x_n\) be a random sample drawn from the exponential population with probability density function
\[ f(x;\lambda) = \lambda e^{-\lambda x}, \qquad x\ge 0,\ \lambda>0 . \]The likelihood function is
\[ L = f(x_1;\lambda)\,f(x_2;\lambda)\cdots f(x_n;\lambda) = \lambda^{n} e^{-\lambda\sum_{i=1}^{n} x_i}, \]so
\[ L_0 = \lambda_0^{n} e^{-\lambda_0\sum x_i}, \qquad L_1 = \lambda_1^{n} e^{-\lambda_1\sum x_i}. \]The best critical region obtained by the N–P lemma, \(L_1/L_0 \ge K\):
\[ \left(\frac{\lambda_1}{\lambda_0}\right)^{n} e^{-(\lambda_1-\lambda_0)\sum_{i=1}^{n} x_i} \ge K \] \[ \Longrightarrow\; e^{-(\lambda_1-\lambda_0)\sum x_i} \ge K\left(\frac{\lambda_0}{\lambda_1}\right)^{n} = K_1 \ (\text{say}) \qquad\Longrightarrow\qquad e^{(\lambda_0-\lambda_1)\sum x_i} \ge K_1 \] \[ \Longrightarrow\; (\lambda_0-\lambda_1)\sum_{i=1}^{n} x_i \ge \log K_1 = K_2 \ (\text{say}) \] \[ \Longrightarrow\; (\lambda_0-\lambda_1)\,n\bar x \ge K_2 \qquad\Longrightarrow\qquad \bar x \ge \frac{K_2}{n(\lambda_0-\lambda_1)} \]Case 1. If \(\lambda_0 > \lambda_1\), then \(\dfrac{K_2}{n(\lambda_0-\lambda_1)} = K_3 \ (\text{say})\), and the best critical region of size \(\alpha\) is \(\ W = \{x\in S : \bar x \ge K_3\}\).
Case 2. If \(\lambda_0 < \lambda_1\), then \(\dfrac{K_2}{n(\lambda_0-\lambda_1)} = K_4 \ (\text{say})\), and the best critical region of size \(\alpha\) is \(\ W = \{x\in S : \bar x \le K_4\}\).
The printed page writes the middle step as \(e^{-(\lambda_0-\lambda_1)\sum x_i} \ge K_1\). That sign contradicts both the line above it and the line below it, which take \(\log\) to get \((\lambda_0-\lambda_1)\sum x_i \ge \log K_1\); the exponent is positive, as written above.
Notice how every one of the four ends in the same place: a statement about \(\bar x\) alone. Four different populations, four different likelihood ratios, one test statistic — which is why the rest of this course can talk about \(\bar x\) and never mention likelihood again.
Test \(H_0: \mu = 100\) vs \(H_1: \mu = 105\) at α = 0.05 with σ = 10, n = 25.
Reject when \(\bar X > 100 + 1.645(10/5) = 103.29\). Power at \(\mu_1 = 105\): \(P(\bar X > 103.29 \mid \mu = 105) = P(Z > -0.855) = 0.804\).
Test \(H_0: p = 0.5\) vs \(H_1: p = 0.6\) with n = 100. MPT rejects when \(X \ge c\). For a size of at most 0.05, \(c = 59\): the exact size is \(P(X \ge 59 \mid p = 0.5) = 0.044\), while \(c = 58\) would give 0.067. Power at \(p_1 = 0.6\) is \(P(X \ge 59 \mid p = 0.6) = 0.62\).
The Neyman–Pearson lemma gives the most powerful test only for a simple-vs-simple hypothesis. The likelihood ratio test extends the idea to composite hypotheses and multiple parameters. To test \(H_0:\theta\in\Theta_0\) against \(H_1:\theta\in\Theta\), form
The numerator maximises the likelihood under the restriction \(H_0\); the denominator maximises it without restriction. Small \(\lambda\) means the data fit far better under \(H_1\), so we reject \(H_0\) when \(\lambda \le \lambda_0\) (equivalently \(-2\ln\lambda\) large).
where \(r\) is the number of independent restrictions imposed by \(H_0\). This provides the critical value when the exact distribution of \(\lambda\) is intractable.
Example. For a normal sample with unknown mean and variance, the LRT of \(H_0:\mu=\mu_0\) reduces exactly to the one-sample t-test — the familiar test is the likelihood ratio test in disguise.
Where this came from, and where it goes. Unit 1 estimated a parameter and put an interval around it; this unit turned that into a decision procedure. But it left one thing deliberately open: which distribution the test statistic follows. That is not settled by theory — it is settled by the sample. Unit 3 takes the easiest case first: when \(n\) is large, the Central Limit Theorem makes the statistic normal whatever the population looks like, so every test in it is a \(Z\)-test.