A test \(\varphi^{*}\) of level \(\alpha\) is uniformly most powerful against \(H_1: \theta \in \Theta_1\) if
\[ \beta_{\varphi^{*}}(\theta) \ge \beta_{\varphi}(\theta) \quad\text{for every } \theta \in \Theta_1 \]and every competing level-\(\alpha\) test \(\varphi\).
Read the quantifier. The inequality must hold at every point of \(\Theta_1\) simultaneously, against every rival. That is a very strong demand: it asks one test to win a whole family of separate contests at once. The surprise is not that UMP tests are rare — it is that they exist at all. They exist because for certain families the likelihood ratio happens to be ordered in a way that does not depend on which alternative is being considered.
A family \(\{f_\theta : \theta \in \Theta\}\), \(\Theta \subseteq \mathbb{R}\), has monotone likelihood ratio in the statistic \(T(x)\) if for every pair \(\theta_1 > \theta_2\) the ratio
\[ \frac{f_{\theta_1}(x)}{f_{\theta_2}(x)} \quad\text{is a non-decreasing function of } T(x) \]wherever it is defined.
What the definition is really saying. The likelihood ratio between any two parameter values sorts the sample space in the same order — the order of \(T\) — no matter which two values are chosen. That is precisely the condition under which the Neyman–Pearson critical region \(\{f_1 > kf_0\}\) becomes \(\{T > c\}\) with a \(c\) that does not know about \(\theta_1\).
The one-parameter exponential family has it. If
\[ f_\theta(x) = h(x)\,c(\theta)\,\exp\!\left\{w(\theta)\,T(x)\right\} \]with \(w\) strictly increasing, then
\[ \frac{f_{\theta_1}(x)}{f_{\theta_2}(x)} = \frac{c(\theta_1)}{c(\theta_2)} \exp\!\left\{\left[w(\theta_1) - w(\theta_2)\right]T(x)\right\}, \]and the bracket is positive, so the ratio is increasing in \(T(x)\). Binomial, Poisson, normal with known variance, exponential and gamma with known shape are all of this form, which is why one-sided tests in those families are so well behaved.
The binomial case settles a loose end from Example 1.1: with \(T(x) = x\) and \(w(p) = \log\frac{p}{1-p}\) increasing in \(p\), the ratio is increasing in \(x\) for every pair \(p_2 < p_1\), not merely for \(0.3\) and \(0.6\).
Suppose the family has monotone likelihood ratio in \(T\). To test
\[ H_0: \theta \le \theta_0 \qquad\text{against}\qquad H_1: \theta > \theta_0, \]the test
\[ \varphi^{*}(x) = \begin{cases} 1, & T(x) > c, \\ \gamma, & T(x) = c, \\ 0, & T(x) < c, \end{cases} \qquad E_{\theta_0}\varphi^{*} = \alpha, \]is uniformly most powerful at level \(\alpha\).
Step 1 — it beats every rival at each fixed \(\theta_1 > \theta_0\). Fix such a \(\theta_1\) and consider the simple-versus-simple problem \(\theta_0\) against \(\theta_1\). By monotone likelihood ratio, \(f_{\theta_1}/f_{\theta_0}\) is non-decreasing in \(T\), so the set \(\{f_{\theta_1} > k f_{\theta_0}\}\) is exactly a set of the form \(\{T > c\}\), with the boundary \(\{T = c\}\) carrying the randomization. By the Neyman–Pearson lemma the test \(\varphi^{*}\) is most powerful for that \(\theta_1\).
Step 2 — \(c\) and \(\gamma\) do not depend on \(\theta_1\). They were fixed by the single requirement \(E_{\theta_0}\varphi^{*} = \alpha\), which mentions \(\theta_0\) only. The test of Step 1 is therefore the same test for every \(\theta_1 > \theta_0\) — and a test that is most powerful against each alternative separately, and is one test rather than a family of them, is uniformly most powerful.
Step 3 — it is of level \(\alpha\) on the whole of \(H_0\). This is the step most often skipped, and \(H_0\) is composite, so it is needed. The power function of \(\varphi^{*}\) is non-decreasing in \(\theta\): for \(\theta' < \theta''\), applying Step 1 with \(\theta'\) in the role of the null shows \(\varphi^{*}\) is most powerful there too, and a most powerful test is at least as powerful as the constant test \(\varphi \equiv E_{\theta'}\varphi^{*}\), giving \(E_{\theta''}\varphi^{*} \ge E_{\theta'}\varphi^{*}\). Hence
\[ \sup_{\theta \le \theta_0} E_\theta \varphi^{*} = E_{\theta_0}\varphi^{*} = \alpha, \qquad\blacksquare \]so the test really is of level \(\alpha\) against the composite null, and not merely at its boundary.
Given. \(X_1, \dots, X_9\) independent \(N(\mu, 1)\), so \(\bar X \sim N(\mu, 1/9)\) and the standard error is exactly \(1/3\). The critical values below divide by \(3\) rather than multiplying by the rounded \(0.333333\): that rounding would turn \(0.548285\) into \(0.548284\), wrong in the last digit. Test \(H_0: \mu = 0\) against \(H_1: \mu \ne 0\) at \(\alpha = 0.05\).
By Karlin–Rubin the UMP test against \(\mu > 0\) is
\[ \text{test A: reject when } \bar X > \frac{1.644854}{3} = 0.548285, \]and against \(\mu < 0\) it is
\[ \text{test B: reject when } \bar X < -0.548285. \]Both have size exactly \(0.05\). Their power functions are
\[ \beta_A(\mu) = 1 - \Phi\!\left(\frac{0.548285 - \mu}{0.333333}\right), \qquad \beta_B(\mu) = \Phi\!\left(\frac{-0.548285 - \mu}{0.333333}\right). \]Alongside them, the familiar two-sided test:
\[ \text{test U: reject when } \left|\bar X\right| > \frac{1.959964}{3} = 0.653321. \]| \(\mu\) | power of A | power of B | power of U |
|---|---|---|---|
| −1.0 | 0.000002 | 0.912314 | 0.850839 |
| −0.5 | 0.000831 | 0.442413 | 0.323041 |
| 0.0 | 0.050000 | 0.050000 | 0.050000 |
| 0.5 | 0.442413 | 0.000831 | 0.323041 |
| 1.0 | 0.912314 | 0.000002 | 0.850839 |
Step 1 — the argument, stated once. Suppose a UMP level-\(0.05\) test \(\varphi^{*}\) existed against \(\mu \ne 0\). Then at \(\mu = 1\) it would have to be at least as powerful as A, so \(\beta_{\varphi^{*}}(1) \ge 0.912314\); and at \(\mu = -1\) at least as powerful as B, so \(\beta_{\varphi^{*}}(-1) \ge 0.912314\). But A is the unique most powerful level-\(0.05\) test at \(\mu = 1\) — that is the necessity half of the lemma, proved in Unit 1 — so \(\varphi^{*}\) would have to be A. The same argument at \(\mu = -1\) forces \(\varphi^{*}\) to be B. A test cannot be both, since \(\beta_A(-1) = 0.000002\) and \(\beta_B(-1) = 0.912314\). No UMP test exists. \(\blacksquare\)
Step 2 — the price of insisting. Test A is not merely sub-optimal on the left, it is indefensible there: at \(\mu = -0.5\) its power is \(0.000831\), far below its own level of \(0.05\). A test whose power falls below \(\alpha\) is biased — it is more likely to reject when the null is true than when this particular alternative is. Using A when the alternative is genuinely two-sided is not a small loss of efficiency; it is a test that works backwards on half the parameter space.
Step 3 — the repair. Test U has power \(0.050000\) at \(\mu = 0\) and strictly more than \(0.05\) everywhere else: it is unbiased. It pays for that by losing to A on the right — \(0.323041\) against \(0.442413\) at \(\mu = 0.5\) — which is exactly the point. Restricting attention to unbiased tests throws away the tests that cannot be defended, and inside what remains a best test exists again.
Unbiased. A level-\(\alpha\) test is unbiased if
\[ \beta_\varphi(\theta) \le \alpha \ \text{ for } \theta \in \Theta_0, \qquad \beta_\varphi(\theta) \ge \alpha \ \text{ for } \theta \in \Theta_1. \]A test that is uniformly most powerful within the unbiased class is UMPU. Test U above is UMPU for the two-sided normal problem.
Similar. A test is similar of size \(\alpha\) on a set \(\Lambda\) of parameter values if
\[ E_\theta\varphi(X) = \alpha \quad\text{for every } \theta \in \Lambda, \]with equality, not inequality. The relevance is immediate: if the power function is continuous, an unbiased test must equal \(\alpha\) on the boundary between \(\Theta_0\) and \(\Theta_1\), because it is \(\le \alpha\) on one side and \(\ge \alpha\) on the other. So every unbiased test is similar on the boundary, and the search for a UMPU test can be conducted inside the smaller class of similar tests.
Neyman structure. Let \(T\) be sufficient for the nuisance parameter on the boundary. A test has Neyman structure with respect to \(T\) if
\[ E\!\left[\varphi(X) \mid T = t\right] = \alpha \quad\text{for almost every } t. \]Such a test is similar — average the conditional expectation over \(t\). The converse needs a condition: if the family of distributions of \(T\) on the boundary is complete, every similar test has Neyman structure. The proof is one line and is worth seeing, because it is the same completeness argument as Lehmann–Scheffé: put \(g(T) = E[\varphi \mid T] - \alpha\); similarity gives \(E_\theta\, g(T) = 0\) for every \(\theta\) on the boundary, and completeness then forces \(g \equiv 0\).
Why this matters in practice. It converts an unmanageable search over all tests into a manageable one: condition on the sufficient statistic for the nuisance parameter, and inside each conditional problem the nuisance parameter has gone. The two-sample \(t\) test and Fisher's exact test are both produced this way — the first conditioning on the pooled variance, the second on the margins of the table.
| Class | Exists when | Example |
|---|---|---|
| UMP | MLR family, one-sided alternative | one-sided \(z\) test; one-sided binomial test |
| UMPU | exponential family, two-sided alternative | two-sided \(z\) test; two-sided \(t\) test |
| UMPI (invariant) | the problem has a symmetry group the test should respect | \(F\) test in the linear model, invariant under orthogonal transformation |
| LRT | always constructible — optimality only asymptotic | Unit 3 |
The descent is a descent. Each step down weakens the optimality claim in exchange for existence. The likelihood ratio test at the bottom claims nothing exact at all — it is a construction, not a theorem — and that is the honest reason it is the one used everywhere: it can always be written down, which none of the three above it can.