Skip to the content

Topics Covered

Test Function Randomized Tests Exact Size Neyman–Pearson Existence Necessity Simple vs Composite
On this page
  1. 1. The Test Function
  2. 2. Why a Discrete Distribution Forces the Issue
  3. 3. The Neyman–Pearson Lemma in Its Complete Form
  4. 4. The Question the Lemma Cannot Answer
Where this unit starts — and it starts past the lemma. Inferential Statistics, Unit 2 already gives, and none of it is written again here: the statistical hypothesis, simple against composite, the critical region, both kinds of error written as integrals, the level and the \(p\) value, the power of a test, one- against two-tailed tests, the Neyman–Pearson lemma with its proof, four best critical regions derived from it, and the generalized likelihood ratio statistic with Wilks' theorem stated.

This unit begins with the thing that treatment cannot do. The lemma as proved there produces a critical region — a set. For a discrete distribution no set has the size you asked for, so the test is either too small or too large and the lemma's optimality claim is quietly out of reach. The repair is to stop thinking in sets and start thinking in functions, and the rest of this course is written in that language.

1. The Test Function

FROM A REGION TO A FUNCTION

A test that rejects on a set \(W\) can be written as its indicator. Let

\[ \varphi(x) = P(\text{reject } H_0 \mid X = x). \]

A test defined by a critical region \(W\) is the special case \(\varphi(x) = 1\) for \(x \in W\) and \(0\) otherwise — a non-randomized test. Allowing \(\varphi\) to take values strictly between \(0\) and \(1\) gives a randomized test: on observing such an \(x\) the experimenter performs an auxiliary experiment, independent of the data, and rejects with probability \(\varphi(x)\).

Every quantity of the previous treatment rewrites immediately:

\[ \text{size} = E_{\theta_0}\!\left[\varphi(X)\right], \qquad \text{power at } \theta = \beta_\varphi(\theta) = E_{\theta}\!\left[\varphi(X)\right]. \]

For a non-randomized test these are \(P_{\theta_0}(X \in W)\) and \(P_\theta(X \in W)\), so nothing is lost; but the class of tests has been enlarged, and the enlargement is exactly what the discrete case needs.

The objection, and the answer to it. Deciding by spinning a wheel offends most people the first time they meet it, and the discomfort is worth taking seriously: two statisticians with the same data can reach opposite conclusions. The answer is not that randomization is a good way to decide. It is that randomized tests are the mathematical completion of the problem — the set of achievable \(\left(\text{size},\ \text{power}\right)\) pairs is convex only when they are admitted, and without that convexity the theory of optimal tests has holes in it. In practice one reports the conservative test and the \(p\) value; in theory one works in the completed class, because that is where the theorems are true.

2. Why a Discrete Distribution Forces the Issue

EXAMPLE 1.1 — NO NON-RANDOMIZED TEST HAS SIZE 0.05

Given. \(X \sim \text{Bin}(10, p)\), testing \(H_0: p = 0.3\) against \(H_1: p = 0.6\) at \(\alpha = 0.05\). Both hypotheses are simple, so the Neyman–Pearson lemma applies and the likelihood ratio

\[ \frac{f_1(x)}{f_0(x)} = \frac{\binom{10}{x}(0.6)^{x}(0.4)^{10-x}} {\binom{10}{x}(0.3)^{x}(0.7)^{10-x}} = \left(\frac{0.6}{0.3}\right)^{x}\left(\frac{0.4}{0.7}\right)^{10-x} \]

is increasing in \(x\). So the most powerful test rejects for large \(x\), and the only question left is how large.

Step 1 — the achievable sizes.

\(k\)\(P(X \ge k \mid H_0)\)\(P(X \ge k \mid H_1)\)
40.3503890.945238
50.1502680.833761
60.0473490.633103
70.0105920.382281
80.0015900.167290

Step 2 — read the gap. The size jumps from \(0.150268\) at \(k = 5\) straight to \(0.047349\) at \(k = 6\). The value \(0.05\) lies in the gap. There is no set \(W\) of the required form with \(P_{H_0}(X \in W) = 0.05\); the sizes available are a finite list, and \(0.05\) is not on it.

So a non-randomized test must either reject at \(X \ge 6\), which is a test of size \(0.047349\) and not \(0.05\), or reject at \(X \ge 5\), which has size \(0.150268\) and is not a level-\(0.05\) test at all. The first is what is done in practice and it is not the most powerful level-\(0.05\) test, because it is not a level-\(0.05\) test — it is a level-\(0.047349\) test, and it forfeits the difference.

Step 3 — the randomized test. Put

\[ \varphi(x) = \begin{cases} 1, & x \ge 6, \\ \gamma, & x = 5, \\ 0, & x \le 4, \end{cases} \]

and choose \(\gamma\) so that the size is exactly \(\alpha\):

\[ E_{H_0}\varphi(X) = P_{H_0}(X \ge 6) + \gamma\,P_{H_0}(X = 5) = \alpha, \] \[ \gamma = \frac{\alpha - P_{H_0}(X \ge 6)}{P_{H_0}(X = 5)} = \frac{0.05 - 0.047349}{0.102919} = \frac{13255063}{514596726} = 0.025758. \]

The fraction is exact: with \(p = 3/10\) every probability in sight is rational, and so is \(\gamma\). The size is then \(0.050000\) — not approximately, exactly.

Step 4 — what it buys.

\[ \beta_\varphi(0.6) = P_{H_1}(X \ge 6) + \gamma\,P_{H_1}(X = 5) = 0.633103 + 0.025758 \times 0.200658 = 0.638272, \]

against \(0.633103\) for the conservative test — a gain of \(0.005169\) in power.

Interpretation. The gain is small, and that is the honest summary: on this problem randomization buys about half a percentage point of power. Its importance is not the half point. It is that with randomized tests admitted, the statement “the most powerful level-\(\alpha\) test exists and is given by the likelihood ratio” is true for every \(\alpha\) and every distribution, discrete or not. Without them the statement has to be hedged, and every theorem built on it inherits the hedge.

3. The Neyman–Pearson Lemma in Its Complete Form

STATEMENT

Test \(H_0: \theta = \theta_0\) against \(H_1: \theta = \theta_1\), both simple, with densities \(f_0\) and \(f_1\). Fix \(\alpha \in (0,1)\).

(i) Existence. There exist a constant \(k \ge 0\) and a \(\gamma \in [0,1]\) such that

\[ \varphi(x) = \begin{cases} 1, & f_1(x) > k f_0(x), \\ \gamma, & f_1(x) = k f_0(x), \\ 0, & f_1(x) < k f_0(x), \end{cases} \qquad\text{with}\qquad E_{\theta_0}\varphi(X) = \alpha \]

exactly. (ii) Sufficiency. Any such \(\varphi\) is most powerful at level \(\alpha\). (iii) Necessity. Any most powerful level-\(\alpha\) test has this form almost everywhere, except possibly on the boundary set \(\{f_1 = k f_0\}\).

Inferential Statistics, Unit 2 proves part (ii) for the non-randomized case. What is added here is (i), which is where \(\gamma\) comes from, and (iii), which is what makes the lemma a characterisation rather than merely a recipe.

PROOF OF (i) — WHY A \(k\) AND A \(\gamma\) ALWAYS EXIST

Step 1. Let \(R = f_1(X)/f_0(X)\) and define, for \(c \ge 0\),

\[ G(c) = P_{\theta_0}\!\left(R > c\right). \]

Step 2. \(G\) is non-increasing, right-continuous, with \(G(c) \to 0\) as \(c \to \infty\) and \(G(0^-) = 1\). It is a survival function, so it may jump, and it jumps exactly where \(R\) has an atom.

Step 3. Choose \(k = \inf\{c : G(c) \le \alpha\}\). Right-continuity gives \(G(k) \le \alpha\); the definition of the infimum gives \(G(k^-) \ge \alpha\). So

\[ P_{\theta_0}(R > k) \le \alpha \le P_{\theta_0}(R \ge k). \]

Step 4. If \(P_{\theta_0}(R = k) = 0\) the two bounds coincide, \(\alpha\) is attained by a set, and \(\gamma\) is irrelevant — this is the continuous case. If \(P_{\theta_0}(R = k) > 0\), put

\[ \gamma = \frac{\alpha - P_{\theta_0}(R > k)}{P_{\theta_0}(R = k)}, \]

which Step 3 places in \([0,1]\). Then \(E_{\theta_0}\varphi = P_{\theta_0}(R>k) + \gamma P_{\theta_0}(R=k) = \alpha\) exactly. \(\blacksquare\)

Step 3 is the whole content. The inequality \(P(R > k) \le \alpha \le P(R \ge k)\) says that \(\alpha\) is straddled by the two one-sided probabilities at the jump, and \(\gamma\) is the fraction of the jump needed to close the gap. Example 1.1 is this construction with \(k\) the likelihood ratio at \(x = 5\).

PROOF OF (iii) — NECESSITY

Statement. Let \(\varphi^{*}\) be the test of part (i) and let \(\varphi\) be any most powerful level-\(\alpha\) test. Then \(\varphi = \varphi^{*}\) almost everywhere on \(\{f_1 \ne k f_0\}\).

Step 1. Consider the function

\[ D(x) = \left[\varphi^{*}(x) - \varphi(x)\right]\left[f_1(x) - k f_0(x)\right]. \]

Step 2 — \(D \ge 0\) pointwise. Where \(f_1 > k f_0\) we have \(\varphi^{*} = 1 \ge \varphi\), so both factors are \(\ge 0\). Where \(f_1 < k f_0\) we have \(\varphi^{*} = 0 \le \varphi\), so both are \(\le 0\). Where \(f_1 = k f_0\) the second factor is \(0\). In every case \(D \ge 0\).

Step 3 — its integral is \(\le 0\).

\[ \int D = \underbrace{\left[\beta_{\varphi^{*}}(\theta_1) - \beta_{\varphi}(\theta_1)\right]}_{= 0 \text{, both are most powerful}} - \; k\underbrace{\left[E_{\theta_0}\varphi^{*} - E_{\theta_0}\varphi\right]} _{\ge 0 \text{, since } E_{\theta_0}\varphi^{*} = \alpha \ge E_{\theta_0}\varphi} \;\le\; 0. \]

Step 4. A non-negative function with a non-positive integral is zero almost everywhere. So \(D = 0\) a.e., and since the second factor is non-zero off the boundary set, \(\varphi = \varphi^{*}\) there. \(\blacksquare\)

What necessity is for. Sufficiency says the likelihood ratio test is a best test. Necessity says it is the best test — anything else that achieves the same power is the same test, rewritten. That is what licenses the standard move of Unit 2: to find the best test against a composite alternative, find it against each simple alternative and see whether the answer happens not to depend on which one.

4. The Question the Lemma Cannot Answer

WHERE UNIT 2 BEGINS

The lemma requires \(H_1\) to be simple. Real alternatives are not: one tests \(\mu = \mu_0\) against \(\mu > \mu_0\), a composite family. The lemma still applies pointwise — for each fixed \(\theta_1\) it names the most powerful test against that \(\theta_1\) — and everything then turns on one question:

Does the resulting test depend on \(\theta_1\)?

In Example 1.1 the likelihood ratio was increasing in \(x\) for the particular pair \(0.3\) and \(0.6\). Whether it is increasing in \(x\) for every pair \(p_0 < p_1\) is a property of the binomial family, not of those two numbers, and a family with that property has a name. Unit 2 is about that property and what it buys.