Skip to the content

Topics Covered

Null & Alternative Critical Region Type I & II Errors Significance Level p-value Power of a Test One- vs Two-tailed Neyman–Pearson Most Powerful Test Best Critical Region Producer's Risk
On this page
  1. 1. Statistical Hypothesis
  2. 2. Critical Region (Rejection Region)
  3. 3. Two Types of Errors
  4. 4. p-value
  5. 5. Power of a Test
  6. 6. One-tailed vs Two-tailed Tests
  7. 7. Neyman–Pearson Lemma
  8. 8. Generalized Likelihood Ratio Test (LRT)
  9. General Procedure for Hypothesis Testing
  10. Key Take-aways

1. Statistical Hypothesis

WHAT TESTING OF HYPOTHESIS IS

Hypothesis testing is a process of testing the significance regarding a population parameter on the basis of a sample. A statistic is computed from the sample observations, and on the strength of that statistic we judge whether the sample was drawn from a parent population with certain specified characteristics.

The computed value of the statistic will almost never equal the hypothetical value of the parameter. The whole question is how to read that gap:

So testing of hypothesis is the procedure of examining whether the difference between the computed statistic (from the sample) and the hypothetical parameter (from the population) is significant or not.

Another definition. Testing of hypothesis is a process of decision making, using various statistical methods and the theory of modern probability.

DEFINITION

A statistical hypothesis is a certain statement about the probability distribution of a random variable — equivalently, a certain statement about the population. A statistical hypothesis is denoted by \(H\).

EXAMPLES — writing a hypothesis symbolically
  1. The average life time of an electrical bulb manufactured by a company is 1200 hrs.
    i.e. \(H:\mu = 1200\) hrs.
  2. The average height of competitors in a game is 160 cms.
    i.e. \(H:\mu = 160\) cms.
  3. The average marks in the subject mathematics \((\mu_1)\) is more than that of marks in computers \((\mu_2)\).
    i.e. \(H:\mu_1 > \mu_2\).
  4. Machine X has an effective life period 20 years less than any other machine in the company.
    i.e. \(H:\mu < 20\).
  5. The variation between the products of two companies is 2 hrs in their life time.
    i.e. \(H:\sigma_1^2 - \sigma_2^2 = 2\).

The printed page states example 5 as "the variation … is 10 hrs" and then writes \(\sigma_1^2-\sigma_2^2=2\). The prose and the formula disagree; the formula is the form the rest of the chapter uses, so the sentence is given as 2 above.

Simple and Composite Hypotheses

Simple statistical hypothesis. If a statistical hypothesis specifies the population completely — that is, the probability distribution is fully known — it is called a simple statistical hypothesis. In a simple hypothesis all the parameters of the population are given clear values.

Examples. Let \(X\sim N(\mu,\sigma^2)\), then \(H:\mu = 1200\) kms, \(\sigma^2 = 2\). Let \(X\sim B(n,P)\), then \(H:P=\tfrac13\).

Composite statistical hypothesis. If a statistical hypothesis does not specify the population completely, it is called a composite statistical hypothesis.

Examples. \(H:\mu < 20\); \(H:\mu_1 > \mu_2\); \(H:\sigma^2 = 2,\ \mu > 2\).

The third one is composite even though it pins \(\sigma^2\) down: leaving any parameter unspecified is enough.

Null and Alternative Hypotheses

There are two kinds of hypothesis essential to conducting a test procedure.

1. Null hypothesis. A statistical hypothesis with no difference, or with a null attitude, is called a null hypothesis. It is denoted by \(H_0\).

According to R. A. Fisher: "Null hypothesis is the hypothesis which is tested for possible rejection under the assumption that it is true."

Examples. The average height of the competitors in a game is 160 cms, i.e. \(H_0:\mu=160\) cms. The average life time of electrical bulbs manufactured by a company is 1800 hours, i.e. \(H_0:\mu=1800\) hrs.

2. Alternative hypothesis. A statistical hypothesis which is complementary to the null hypothesis is called an alternative hypothesis. It is denoted by \(H_1\).

It is clear that a null hypothesis is meaningful only when an alternative hypothesis has been formulated alongside it.

EXAMPLES — one null hypothesis, three alternatives

1. If the null hypothesis is that the average height of the competitors in a game is 160 cms, i.e. \(H_0:\mu=160\) cms, then the alternative hypothesis may be formulated as

The alternative hypothesis in (i) gives a two-tailed test; those in (ii) and (iii) give one-tailed tests.

2. If the null hypothesis is that the average life time of electrical bulbs in a company is 1800 hours, the alternative hypothesis may be considered as follows:

The Three Forms of the Alternative

FormTypeExample
\(H_1: \theta \ne \theta_0\)Two-tailedμ ≠ 50
\(H_1: \theta > \theta_0\)Right-tailedμ > 50
\(H_1: \theta < \theta_0\)Left-tailedμ < 50

2. Critical Region (Rejection Region)

The critical region \(W\) is the set of sample-space outcomes for which we reject \(H_0\). Its complement is the acceptance region.

For a continuous test statistic \(T\) with a critical value \(c\):

c α Right-tailed Reject if T > c c α Left-tailed Reject if T < c −c +c α/2 α/2 Two-tailed Reject if |T| > c
Fig 2.1 — Three types of critical regions

3. Two Types of Errors

H₀ TrueH₀ False
Reject H₀Type I error (probability α)Correct decision (probability 1−β)
Accept H₀Correct decision (probability 1−α)Type II error (probability β)

Reducing α typically increases β and vice versa. Increasing sample size \(n\) is the proper way to reduce both.

THE TWO ERRORS AS INTEGRALS OVER THE CRITICAL REGION

A decision — whether \(H_0\) is to be accepted or rejected — is made on the information supplied by the sample data, so there is always a chance of a good decision or of an error. Written over the regions of §2:

Type I error is the error of rejecting \(H_0\) when \(H_0\) is true:

\[ \alpha = P(\text{Type I error}) = P(\text{rejecting } H_0 \text{ when } H_0 \text{ is true}) = P(x\in W \mid H_0) = \int_{W} L_0\,dx, \]

where \(L_0\) is the likelihood function of the sample observations \(x_1,x_2,\dots,x_n\) under \(H_0\).

Type II error is the error of accepting \(H_0\) when \(H_0\) is false:

\[ \beta = P(\text{Type II error}) = P(\text{accepting } H_0 \text{ when } H_0 \text{ is false}) = P(x\in \bar W \mid H_1) = \int_{\bar W} L_1\,dx, \]

where \(L_1\) is the likelihood function under \(H_1\). The two errors are integrals of different likelihoods over complementary regions, which is why shrinking one enlarges the other.

The printed page writes this second integral as \(\int_{\bar W} L_0\,dx\), then says on the very next line that \(L_1\) is the likelihood under \(H_1\). It is \(L_1\), as above.

LEVEL OF SIGNIFICANCE

The probability of a Type I error, \(\alpha\), is known as the level of significance. It is also called the size of the critical region — a name worth keeping in mind, because it says plainly that choosing \(\alpha\) is choosing how big \(W\) is allowed to be.

In statistical quality control the two errors have names taken from who pays for them: \(\alpha\) is the producer's risk (good batches rejected) and \(\beta\) is the consumer's risk (bad batches accepted).

H₀ distribution H₁ distribution α β c (cutoff) Type I (α) and Type II (β) Errors
Fig 2.2 — α is the H₀ tail beyond the cutoff; β is the H₁ tail before it. Moving c right shrinks α but enlarges β; larger n separates the curves and reduces both.
EXAMPLE 1

A factory's quality engineer rejects a defect-free batch (Type I error) — money lost on rework. With α = 0.05, this happens 5 % of the time even though the batch was fine.

EXAMPLE 2

A medical test fails to detect a sick patient (Type II error). If β = 0.20, only 80 % of truly ill patients are correctly identified — power = 0.80.

4. p-value

DEFINITION

The p-value is the probability, assuming \(H_0\) is true, of observing a test statistic as extreme as (or more extreme than) the one observed.

Decision Rule

EXAMPLE 1

For an observed Z = 2.10 in a two-tailed test, p = 2(1 − Φ(2.10)) = 2(0.0179) = 0.0358. Since p < 0.05, reject \(H_0\) at 5 %.

EXAMPLE 2

p = 0.18 in any test ⇒ data are consistent with \(H_0\); no significant evidence to reject.

5. Power of a Test

Power = \(1 - \beta\) = P(reject \(H_0\) | \(H_1\) true) — the probability of correctly detecting an effect.

Factors that Increase Power

  1. Larger sample size \(n\).
  2. Larger effect size (true value of \(\theta\) far from \(\theta_0\)).
  3. Larger α (less stringent test).
  4. Smaller population variance.
  5. One-tailed instead of two-tailed (when direction is known).

Power function \(\pi(\theta) = P(\text{reject } H_0 \mid \theta)\) is a function of the true \(\theta\). At \(\theta = \theta_0\), \(\pi = \alpha\); for \(\theta\) far from \(\theta_0\), \(\pi \to 1\).

POWER OVER THE CRITICAL REGION

In the notation of §2 and §3,

\[ 1-\beta = P(x\in W \mid H_1) \]

is the probability of rejecting \(H_0\) when \(H_0\) is false. This is called the power function of testing the hypothesis, and its value is the power of the test. Compare it with \(\alpha=\int_W L_0\,dx\): the same region \(W\), scored under the other hypothesis.

Most Powerful Test

Take the problem of testing a simple null hypothesis \(H_0:\theta=\theta_0\) against a simple alternative \(H_1:\theta=\theta_1\). The critical region \(W\) is the most powerful critical region of size \(\alpha\) for testing \(H_0\) against \(H_1\) if

\[ P(x\in W \mid H_0) = \int_{W} L_0\,dx = \alpha \tag{1} \]

and

\[ P(x\in W \mid H_1) \;\ge\; P(x\in W_1 \mid H_1) \tag{2} \]

for every other critical region \(W_1\) satisfying (1). The corresponding test is called the most powerful test.

Read plainly: among all regions that make the same number of Type I errors, \(W\) catches the most false nulls. §7 names that region.

6. One-tailed vs Two-tailed Tests

AspectOne-tailedTwo-tailed
Form of \(H_1\)\(\theta > \theta_0\) or \(\theta < \theta_0\)\(\theta \ne \theta_0\)
Critical regionOne side of distributionBoth sides
Critical value at α = 0.05 (Z)1.6451.96
UseDirection known a prioriDirection unknown / either side matters

7. Neyman–Pearson Lemma

STATEMENT

Let \(x_1,x_2,\dots,x_n\) be a random sample of size \(n\) drawn from a population with density function \(f(x,\theta)\). Let \(K>0\) be a constant, and let \(W\) be the most powerful critical region of size \(\alpha\) for testing a simple null hypothesis \(H_0:\theta=\theta_0\) against a simple alternative \(H_1:\theta=\theta_1\), such that

\[ W = \left\{ x\in S : \frac{f(x,\theta_1)}{f(x,\theta_0)} > K \right\} \]

i.e.

\[ W = \left\{ x\in S : \frac{L_1}{L_0} > K \right\} \tag{I} \]

and

\[ \bar W = \left\{ x\in S : \frac{L_1}{L_0} \le K \right\} \tag{II} \]

where \(L_0\) and \(L_1\) are the likelihood functions of the sample observations \(x_1,x_2,\dots,x_n\) under \(H_0\) and \(H_1\) respectively.

The lemma is the foundation of optimal hypothesis testing: it does not merely offer a good test, it says that no test of the same size can do better. For composite hypotheses, generalised likelihood ratio tests (GLRT, §8) extend the idea.

Proof of the Lemma

The size of the critical region is

\[ P(x\in W \mid H_0) = \int_{W} L_0\,dx = \alpha \tag{1} \]

and the power of the region is

\[ P(x\in W \mid H_1) = \int_{W} L_1\,dx = 1-\beta . \tag{2} \]

To prove the lemma we have to show that there exists no other critical region, of size less than or equal to \(\alpha\), which is more powerful than \(W\). Let \(W_1\) be another critical region, of size \(\alpha_1\le\alpha\) and power \(1-\beta_1\):

\[ P(x\in W_1 \mid H_0) = \int_{W_1} L_0\,dx = \alpha_1 \tag{3} \] \[ P(x\in W_1 \mid H_1) = \int_{W_1} L_1\,dx = 1-\beta_1 . \tag{4} \]

We have to prove that the power is greater for \(W\), i.e. that \(1-\beta\) is the larger.

S W W₁ A C B only W both only W₁
Fig 2.3 — The two candidate regions share \(C\), so \(W=A\cup C\) and \(W_1=B\cup C\). Everything below compares \(A\) against \(B\); the shared part cancels.

Let \(W = A\cup C\) and \(W_1 = B\cup C\). Consider

\[ \alpha_1 \le \alpha \;\Longrightarrow\; \int_{W_1} L_0\,dx \le \int_{W} L_0\,dx \] \[ \Longrightarrow\; \int_{B\cup C} L_0\,dx \le \int_{A\cup C} L_0\,dx \;\Longrightarrow\; \int_{B} L_0\,dx \le \int_{A} L_0\,dx \] \[ \Longrightarrow\; \int_{A} L_0\,dx \ge \int_{B} L_0\,dx . \tag{5} \]

From (I) in the statement, for \(x\in W\) we have

\[ \frac{L_1}{L_0} > K \;\Longrightarrow\; L_1 > K L_0 \;\Longrightarrow\; \int_{W} L_1\,dx > K\int_{W} L_0\,dx . \]

Since \(A\subset W\),

\[ \int_{A} L_1\,dx > K\int_{A} L_0\,dx . \tag{6} \]

Multiplying equation (5) by \(K\),

\[ K\int_{A} L_0\,dx \ge K\int_{B} L_0\,dx . \tag{7} \]

From (6) and (7) we get

\[ \int_{A} L_1\,dx > K\int_{B} L_0\,dx \qquad\text{i.e.}\qquad K\int_{B} L_0\,dx \le \int_{A} L_1\,dx . \tag{8} \]

From (II) in the statement, for \(x\in\bar W\) we have

\[ \frac{L_1}{L_0} \le K \;\Longrightarrow\; L_1 \le K L_0 \;\Longrightarrow\; \int_{\bar W} L_1\,dx \le K\int_{\bar W} L_0\,dx . \]

Since \(B\subset\bar W\),

\[ \int_{B} L_1\,dx \le K\int_{B} L_0\,dx . \tag{9} \]

From (8) and (9) we get

\[ \int_{B} L_1\,dx \le \int_{A} L_1\,dx . \]

By adding \(\int_{C} L_1\,dx\) to both sides,

\[ \int_{B\cup C} L_1\,dx \le \int_{A\cup C} L_1\,dx \;\Longrightarrow\; \int_{W_1} L_1\,dx \le \int_{W} L_1\,dx \] \[ \Longrightarrow\; 1-\beta_1 \le 1-\beta \qquad\text{i.e.}\qquad 1-\beta \ge 1-\beta_1 . \quad\blacksquare \]

Hence no critical region of size \(\le\alpha\) has greater power than \(W\), which is the lemma.

One transcription note: the printed page writes step (9) as \(\int_{B} L_1\,dx \le K\int_{P} L_0\,dx\). There is no region \(P\) in the proof; the subscript is a smudged \(B\), as written above.

Best Critical Regions in Standard Distributions

Four applications of the lemma. Each follows the same three moves: write \(L_1/L_0\), take logarithms to turn the product into a sum in \(\sum x_i\), then divide — and watch the inequality flip when the divisor is negative. That flip is the whole reason each problem has two cases.

PROBLEM 1 — best critical region in a binomial population

Obtain the best critical region for testing \(H_0:p=p_0\) against \(H_1:p=p_1\).

Solution. Let \(x_1,x_2,\dots,x_m\) be a random sample of size \(m\) from a binomial population, whose probability mass function is

\[ f(x;n,p) = {}^{n}C_{x}\,p^{x}q^{\,n-x},\quad x=0,1,2,\dots,n, \qquad q = 1-p. \]

The likelihood function of \(x_1,x_2,\dots,x_m\) is

\[ L = f(x_1;n,p)\,f(x_2;n,p)\cdots f(x_m;n,p) = \left({}^{n}C_{x_1}\,{}^{n}C_{x_2}\cdots{}^{n}C_{x_m}\right) p^{\sum_{i=1}^{m} x_i}\,(1-p)^{\,mn-\sum_{i=1}^{m} x_i}. \]

Hence, under the two hypotheses,

\[ L_1 = \left({}^{n}C_{x_1}\cdots{}^{n}C_{x_m}\right) p_1^{\sum x_i}(1-p_1)^{\,mn-\sum x_i}, \qquad L_0 = \left({}^{n}C_{x_1}\cdots{}^{n}C_{x_m}\right) p_0^{\sum x_i}(1-p_0)^{\,mn-\sum x_i}. \]

Using the N–P lemma the best critical region is obtained from \(L_1/L_0 \ge K\):

\[ \left(\frac{p_1}{p_0}\right)^{\sum_{i=1}^{m} x_i} \left(\frac{1-p_1}{1-p_0}\right)^{\,mn-\sum_{i=1}^{m} x_i} \ge K \]

Taking logarithms,

\[ \sum_{i=1}^{m} x_i \log\!\left(\frac{p_1}{p_0}\right) + \left(mn - \sum_{i=1}^{m} x_i\right)\log\!\left(\frac{1-p_1}{1-p_0}\right) \ge \log K = K_1 \ (\text{say}) \] \[ \Longrightarrow\; \sum_{i=1}^{m} x_i\left(\log\frac{p_1}{p_0} - \log\frac{1-p_1}{1-p_0}\right) \ge K_1 - mn\log\frac{1-p_1}{1-p_0} = K_2 \ (\text{say}) \] \[ \Longrightarrow\; m\bar x\left(\log\frac{p_1}{p_0} - \log\frac{1-p_1}{1-p_0}\right) \ge K_2 \qquad\Longrightarrow\qquad \bar x \ge \frac{K_2}{m\left(\log\dfrac{p_1}{p_0} - \log\dfrac{1-p_1}{1-p_0}\right)} \]

Case 1. If \(p_1 > p_0\), then \(\log\dfrac{p_1}{p_0} - \log\dfrac{1-p_1}{1-p_0} > 0\), so

\[ \frac{K_2}{m\left(\log\dfrac{p_1}{p_0}-\log\dfrac{1-p_1}{1-p_0}\right)} = K_3 \ (\text{say}) \]

and the best critical region of size \(\alpha\) is \(\ W = \{x\in S : \bar x \ge K_3\}\).

Case 2. If \(p_1 < p_0\), then that bracket is negative, dividing by it reverses the inequality, and

\[ \frac{K_2}{m\left(\log\dfrac{p_1}{p_0}-\log\dfrac{1-p_1}{1-p_0}\right)} = K_4 \ (\text{say}) \]

so the best critical region of size \(\alpha\) is \(\ W = \{x\in S : \bar x \le K_4\}\).

The sign of the constant is not known, and is not needed. \(K_2\) is an arbitrary constant, so nothing in the working makes \(K_3\) positive or \(K_4\) negative — and a negative \(K_4\) would make \(\{\bar x \le K_4\}\) empty, since \(\bar x \ge 0\) here. What each case fixes is the direction of the region. The constant itself comes from the size condition, \(P(\bar x \ge K_3 \mid H_0) = \alpha\) in Case 1 and \(P(\bar x \le K_4 \mid H_0) = \alpha\) in Case 2. The same holds for the three problems below.

The printed page opens Case 1 with "if \(p_1 \ge p_0\)". At \(p_1=p_0\) the bracket is exactly zero and the division is undefined, so the condition is strict, as written above. The same correction applies to Problem 3.

PROBLEM 2 — best critical region in a Poisson population

Obtain the best critical region for testing \(H_0:\lambda=\lambda_0\) against \(H_1:\lambda=\lambda_1\).

Solution. Let \(x_1,x_2,\dots,x_n\) be a random sample of size \(n\) drawn from a Poisson population, with probability mass function

\[ f(x,\lambda) = \frac{e^{-\lambda}\lambda^{x}}{x!}, \qquad x=0,1,2,\dots,\ \lambda>0. \]

The likelihood function is

\[ L = f(x_1;\lambda)\,f(x_2;\lambda)\cdots f(x_n;\lambda) = \frac{e^{-n\lambda}\,\lambda^{\sum_{i=1}^{n} x_i}}{x_1!\,x_2!\cdots x_n!}, \]

so under \(H_0:\lambda=\lambda_0\) and \(H_1:\lambda=\lambda_1\),

\[ L_0 = \frac{e^{-n\lambda_0}\,\lambda_0^{\sum x_i}}{x_1!\cdots x_n!}, \qquad L_1 = \frac{e^{-n\lambda_1}\,\lambda_1^{\sum x_i}}{x_1!\cdots x_n!}. \]

The best critical region according to the N–P lemma is \(L_1/L_0 \ge K\):

\[ \frac{e^{-n\lambda_1}\lambda_1^{\sum x_i}}{e^{-n\lambda_0}\lambda_0^{\sum x_i}} \ge K \qquad\Longrightarrow\qquad e^{-n(\lambda_1-\lambda_0)}\left(\frac{\lambda_1}{\lambda_0}\right)^{\sum_{i=1}^{n} x_i} \ge K \] \[ \Longrightarrow\; \left(\frac{\lambda_1}{\lambda_0}\right)^{\sum x_i} \ge K\,e^{\,n(\lambda_1-\lambda_0)} = K_1 \ (\text{say}) \] \[ \Longrightarrow\; \sum_{i=1}^{n} x_i \log\!\left(\frac{\lambda_1}{\lambda_0}\right) \ge \log K_1 = K_2 \ (\text{say}) \qquad\Longrightarrow\qquad n\bar x\,(\log\lambda_1 - \log\lambda_0) \ge K_2 \] \[ \Longrightarrow\; \bar x \ge \frac{K_2}{n(\log\lambda_1 - \log\lambda_0)} \]

Case 1. If \(\lambda_1 > \lambda_0\), then \(\dfrac{K_2}{n(\log\lambda_1-\log\lambda_0)} = K_3 \ (\text{say})\), and the best critical region of size \(\alpha\) is \(\ W = \{x\in S : \bar x \ge K_3\}\).

Case 2. If \(\lambda_1 < \lambda_0\), then \(\dfrac{K_2}{n(\log\lambda_1-\log\lambda_0)} = K_4 \ (\text{say})\), and the best critical region of size \(\alpha\) is \(\ W = \{x\in S : \bar x \le K_4\}\).

PROBLEM 3 — best critical region in a normal population

Obtain the best critical region for testing \(H_0:\mu=\mu_0\) against \(H_1:\mu=\mu_1\).

Solution. Let \(x_1,x_2,\dots,x_n\) be a random sample drawn from a normal population with probability density function

\[ f(x;\mu,\sigma^2) = \frac{1}{\sigma\sqrt{2\pi}}\, e^{-\frac12\left(\frac{x-\mu}{\sigma}\right)^{2}}, \qquad -\infty<x<\infty . \]

The likelihood function is

\[ L = f(x_1;\mu,\sigma^2)\cdots f(x_n;\mu,\sigma^2) = \frac{1}{(\sigma\sqrt{2\pi})^{n}}\, e^{-\frac{1}{2\sigma^{2}}\sum_{i=1}^{n}(x_i-\mu)^{2}}, \]

so

\[ L_0 = \frac{1}{(\sigma\sqrt{2\pi})^{n}}e^{-\frac{1}{2\sigma^{2}}\sum(x_i-\mu_0)^{2}}, \qquad L_1 = \frac{1}{(\sigma\sqrt{2\pi})^{n}}e^{-\frac{1}{2\sigma^{2}}\sum(x_i-\mu_1)^{2}}. \]

Now the best critical region using the N–P lemma, \(L_1/L_0 \ge K\):

\[ e^{-\frac{1}{2\sigma^{2}}\left[\sum(x_i-\mu_1)^{2}-\sum(x_i-\mu_0)^{2}\right]} \ge K \] \[ \Longrightarrow\; \frac{-1}{2\sigma^{2}} \left[\sum_{i=1}^{n}(x_i-\mu_1)^{2}-\sum_{i=1}^{n}(x_i-\mu_0)^{2}\right] \ge \log K = K_1 \ (\text{say}) \]

Multiply by \(-2\sigma^2\). That reverses the inequality, and writing the difference the other way round reverses it back, so the \(\ge\) survives:

\[ \sum_{i=1}^{n}(x_i-\mu_0)^{2} - \sum_{i=1}^{n}(x_i-\mu_1)^{2} \ge K_1(2\sigma^{2}) = K_2 \ (\text{say}) \]

Expanding both squares, the \(\sum x_i^2\) terms cancel:

\[ \sum x_i^{2} - 2\mu_0\sum x_i + n\mu_0^{2} - \sum x_i^{2} + 2\mu_1\sum x_i - n\mu_1^{2} \ge K_2 \] \[ \Longrightarrow\; 2(\mu_1-\mu_0)\sum_{i=1}^{n} x_i - n(\mu_1^{2}-\mu_0^{2}) \ge K_2 \] \[ \Longrightarrow\; (\mu_1-\mu_0)\sum_{i=1}^{n} x_i \ge \frac{K_2 + n(\mu_1^{2}-\mu_0^{2})}{2} = K_3 \ (\text{say}) \] \[ \Longrightarrow\; (\mu_1-\mu_0)\,n\bar x \ge K_3 \qquad\Longrightarrow\qquad \bar x \ge \frac{K_3}{n(\mu_1-\mu_0)} \]

Case 1. If \(\mu_1 > \mu_0\), then \(\dfrac{K_3}{n(\mu_1-\mu_0)} = K_4 \ (\text{say})\), and the best critical region of size \(\alpha\) is \(\ W = \{x\in S : \bar x \ge K_4\}\).

Case 2. If \(\mu_1 < \mu_0\), then \(\dfrac{K_3}{n(\mu_1-\mu_0)} = K_5 \ (\text{say})\), and the best critical region of size \(\alpha\) is \(\ W = \{x\in S : \bar x \le K_5\}\).

This is where the familiar \(Z\)-test comes from: the optimal test of a normal mean rejects for extreme \(\bar x\), and nothing else. Unit 3 supplies the constant.

PROBLEM 4 — best critical region in an exponential population

Obtain the best critical region for testing \(H_0:\lambda=\lambda_0\) against \(H_1:\lambda=\lambda_1\).

Solution. Let \(x_1,x_2,\dots,x_n\) be a random sample drawn from the exponential population with probability density function

\[ f(x;\lambda) = \lambda e^{-\lambda x}, \qquad x\ge 0,\ \lambda>0 . \]

The likelihood function is

\[ L = f(x_1;\lambda)\,f(x_2;\lambda)\cdots f(x_n;\lambda) = \lambda^{n} e^{-\lambda\sum_{i=1}^{n} x_i}, \]

so

\[ L_0 = \lambda_0^{n} e^{-\lambda_0\sum x_i}, \qquad L_1 = \lambda_1^{n} e^{-\lambda_1\sum x_i}. \]

The best critical region obtained by the N–P lemma, \(L_1/L_0 \ge K\):

\[ \left(\frac{\lambda_1}{\lambda_0}\right)^{n} e^{-(\lambda_1-\lambda_0)\sum_{i=1}^{n} x_i} \ge K \] \[ \Longrightarrow\; e^{-(\lambda_1-\lambda_0)\sum x_i} \ge K\left(\frac{\lambda_0}{\lambda_1}\right)^{n} = K_1 \ (\text{say}) \qquad\Longrightarrow\qquad e^{(\lambda_0-\lambda_1)\sum x_i} \ge K_1 \] \[ \Longrightarrow\; (\lambda_0-\lambda_1)\sum_{i=1}^{n} x_i \ge \log K_1 = K_2 \ (\text{say}) \] \[ \Longrightarrow\; (\lambda_0-\lambda_1)\,n\bar x \ge K_2 \qquad\Longrightarrow\qquad \bar x \ge \frac{K_2}{n(\lambda_0-\lambda_1)} \]

Case 1. If \(\lambda_0 > \lambda_1\), then \(\dfrac{K_2}{n(\lambda_0-\lambda_1)} = K_3 \ (\text{say})\), and the best critical region of size \(\alpha\) is \(\ W = \{x\in S : \bar x \ge K_3\}\).

Case 2. If \(\lambda_0 < \lambda_1\), then \(\dfrac{K_2}{n(\lambda_0-\lambda_1)} = K_4 \ (\text{say})\), and the best critical region of size \(\alpha\) is \(\ W = \{x\in S : \bar x \le K_4\}\).

The printed page writes the middle step as \(e^{-(\lambda_0-\lambda_1)\sum x_i} \ge K_1\). That sign contradicts both the line above it and the line below it, which take \(\log\) to get \((\lambda_0-\lambda_1)\sum x_i \ge \log K_1\); the exponent is positive, as written above.

Notice how every one of the four ends in the same place: a statement about \(\bar x\) alone. Four different populations, four different likelihood ratios, one test statistic — which is why the rest of this course can talk about \(\bar x\) and never mention likelihood again.

EXAMPLE 1 (Normal)

Test \(H_0: \mu = 100\) vs \(H_1: \mu = 105\) at α = 0.05 with σ = 10, n = 25.

Reject when \(\bar X > 100 + 1.645(10/5) = 103.29\). Power at \(\mu_1 = 105\): \(P(\bar X > 103.29 \mid \mu = 105) = P(Z > -0.855) = 0.804\).

EXAMPLE 2 (Binomial)

Test \(H_0: p = 0.5\) vs \(H_1: p = 0.6\) with n = 100. MPT rejects when \(X \ge c\). For a size of at most 0.05, \(c = 59\): the exact size is \(P(X \ge 59 \mid p = 0.5) = 0.044\), while \(c = 58\) would give 0.067. Power at \(p_1 = 0.6\) is \(P(X \ge 59 \mid p = 0.6) = 0.62\).

8. Generalized Likelihood Ratio Test (LRT)

The Neyman–Pearson lemma gives the most powerful test only for a simple-vs-simple hypothesis. The likelihood ratio test extends the idea to composite hypotheses and multiple parameters. To test \(H_0:\theta\in\Theta_0\) against \(H_1:\theta\in\Theta\), form

\[ \lambda = \dfrac{\displaystyle\sup_{\theta\in\Theta_0} L(\theta)}{\displaystyle\sup_{\theta\in\Theta} L(\theta)}, \qquad 0 \le \lambda \le 1. \]

The numerator maximises the likelihood under the restriction \(H_0\); the denominator maximises it without restriction. Small \(\lambda\) means the data fit far better under \(H_1\), so we reject \(H_0\) when \(\lambda \le \lambda_0\) (equivalently \(-2\ln\lambda\) large).

WILKS' THEOREM (large samples) \[ -2\ln\lambda \;\xrightarrow{d}\; \chi^2_{r}, \]

where \(r\) is the number of independent restrictions imposed by \(H_0\). This provides the critical value when the exact distribution of \(\lambda\) is intractable.

Example. For a normal sample with unknown mean and variance, the LRT of \(H_0:\mu=\mu_0\) reduces exactly to the one-sample t-test — the familiar test is the likelihood ratio test in disguise.

General Procedure for Hypothesis Testing

  1. State \(H_0\) and \(H_1\).
  2. Choose level of significance α.
  3. Identify the appropriate test statistic and its distribution under \(H_0\).
  4. Determine the critical region (or compute p-value).
  5. Compute the test statistic from data.
  6. Compare and conclude.

Key Take-aways

Where this came from, and where it goes. Unit 1 estimated a parameter and put an interval around it; this unit turned that into a decision procedure. But it left one thing deliberately open: which distribution the test statistic follows. That is not settled by theory — it is settled by the sample. Unit 3 takes the easiest case first: when \(n\) is large, the Central Limit Theorem makes the statistic normal whatever the population looks like, so every test in it is a \(Z\)-test.

→ Unit 3 — Large Sample Tests