Skip to the content

Topics Covered

SPRT Wald's Boundaries Wald's Identity OC Function ASN Function Truncation Loss and Risk Bayes and Minimax Rules
On this page
  1. 1. The Sequential Probability Ratio Test
  2. 2. Wald's Identity, and What It Is For
  3. 3. The Operating Characteristic and the Average Sample Number
  4. 4. Where Sequential Testing Is Actually Used
  5. 5. Elements of Decision Theory
Where this unit starts. Everything so far has fixed the sample size in advance and then asked which test is best. This unit lets the sample size be a random variable decided by the data, which is a different problem with a different optimality theorem — and then closes the course with the decision-theoretic frame that holds testing and estimation together.

Loss, risk and Bayes rules are built in Estimation Theory, Unit 4 and are not rebuilt here; section 5 only says what changes when the action is a decision between two hypotheses rather than a number.

1. The Sequential Probability Ratio Test

THE PROCEDURE

Test the simple \(H_0: \theta = \theta_0\) against the simple \(H_1: \theta = \theta_1\). After \(n\) observations form the likelihood ratio

\[ \Lambda_n = \prod_{i=1}^{n}\frac{f_{\theta_1}(x_i)}{f_{\theta_0}(x_i)}, \qquad \text{equivalently}\qquad Z_n = \sum_{i=1}^{n} z_i, \quad z_i = \ln\frac{f_{\theta_1}(x_i)}{f_{\theta_0}(x_i)}. \]

Fix two boundaries \(B < 1 < A\) and at each stage:

Wald's approximations for the boundaries. To achieve error probabilities \(\alpha\) and \(\beta\),

\[ A \approx \frac{1-\beta}{\alpha}, \qquad B \approx \frac{\beta}{1-\alpha}. \]

Where these come from, in one line each. On the event that the test stops by rejecting, \(\Lambda_N \ge A\), so the probability of that event under \(H_1\) is at least \(A\) times its probability under \(H_0\): \(1 - \beta \ge A\alpha\). Symmetrically \(\beta \le B(1-\alpha)\). The approximations replace the inequalities by equalities, which is exact only if the ratio lands precisely on a boundary rather than overshooting it. Since \(\Lambda\) moves in jumps, it always overshoots a little, so the true error probabilities are slightly smaller than \(\alpha\) and \(\beta\) — the approximation errs on the safe side, and \(\alpha' + \beta' \le \alpha + \beta\) always holds.

The optimality theorem. Among all tests, sequential or fixed, with error probabilities no worse than \(\alpha\) and \(\beta\), the SPRT minimises both \(E_{\theta_0}(N)\) and \(E_{\theta_1}(N)\) — the Wald–Wolfowitz theorem. It is a remarkable result: the SPRT is simultaneously optimal at both hypotheses, which no fixed-sample test can be. Note what it does not say: nothing about \(\theta\) between the two, and section 3 shows that is exactly where the SPRT is at its worst.

2. Wald's Identity, and What It Is For

WALD'S EQUATION AND THE FUNDAMENTAL IDENTITY

Wald's equation. If \(N\) is a stopping time for the sequence \(z_1, z_2, \dots\) of independent identically distributed variables with finite mean, and \(E(N) < \infty\), then

\[ E\!\left(\sum_{i=1}^{N} z_i\right) = E(N)\,E(z). \]

The content is that \(N\) is allowed to depend on the data. For a fixed \(n\) the identity is trivial; for a random \(N\) it is not obvious at all, and it is false without the stopping-time condition — a rule that says “stop just before a large value” would break it. The proof writes \(\sum_{i\le N} z_i = \sum_{i\ge 1} z_i \mathbb{1}\{N \ge i\}\) and uses the fact that \(\{N \ge i\}\) depends only on \(z_1, \dots, z_{i-1}\), so it is independent of \(z_i\).

The fundamental identity extends this to the moment generating function: for \(h\) with \(E\!\left[e^{hz}\right]\) finite,

\[ E\!\left[e^{hZ_N}\left(E\,e^{hz}\right)^{-N}\right] = 1. \]

Setting \(E\!\left[e^{hz}\right] = 1\) kills the second factor, and that single choice of \(h\) is what makes the next section computable.

3. The Operating Characteristic and the Average Sample Number

THE TWO FUNCTIONS THAT DESCRIBE AN SPRT

The SPRT is built from two simple hypotheses, but it will be run on data from whatever \(\theta\) is true. Two functions of \(\theta\) describe what happens:

\[ L(\theta) = P_\theta(\text{accept } H_0), \qquad E_\theta(N) = \text{the average sample number}. \]

Wald's approximations. Let \(h(\theta)\) be the non-zero solution of

\[ E_\theta\!\left[\left(\frac{f_{\theta_1}(X)}{f_{\theta_0}(X)}\right)^{h}\right] = 1. \]

Then

\[ L(\theta) \approx \frac{A^{h(\theta)} - 1}{A^{h(\theta)} - B^{h(\theta)}}, \qquad E_\theta(N) \approx \frac{L(\theta)\ln B + \left[1 - L(\theta)\right]\ln A} {E_\theta(z)}. \]

Two traps in using these.

Two values are exact rather than approximate and serve as checks: \(h(\theta_0) = 1\) and \(h(\theta_1) = -1\), so \(L(\theta_0)\) and \(L(\theta_1)\) must come out as \(1-\alpha\) and \(\beta\).

EXAMPLE 4.1 — AN SPRT FOR A PROPORTION, WORKED THROUGH

Given. Bernoulli sampling, \(H_0: p = 0.1\) against \(H_1: p = 0.3\), with \(\alpha = \beta = 0.05\).

Step 1 — the boundaries.

\[ A = \frac{1 - 0.05}{0.05} = 19.0000, \qquad B = \frac{0.05}{1 - 0.05} = 0.052632. \]

Note \(B = 1/A\), which happens precisely because \(\alpha = \beta\).

Step 2 — the log-likelihood-ratio increments.

\[ z = \ln\frac{0.3}{0.1} = 1.098612 \ \text{ when } x = 1, \qquad z = \ln\frac{0.7}{0.9} = -0.251314 \ \text{ when } x = 0. \]

Step 3 — the boundaries as straight lines. After \(n\) trials with \(S_n = \sum x_i\) successes, \(Z_n = S_n(1.098612) + (n - S_n)(-0.251314)\). Writing \(B < \Lambda_n < A\) as \(\ln B < Z_n < \ln A\) and solving for \(S_n\):

\[ -2.181184 + 0.186169\,n \;<\; S_n \;<\; 2.181184 + 0.186169\,n. \]

The decision rule is two parallel lines on a plot of cumulative successes against trials, and sampling continues while the path stays between them. The slope \(0.186169\) is \(-z_{(0)}/(z_{(1)} - z_{(0)})\) and sits between \(p_0\) and \(p_1\), as it must.

Step 4 — the OC and ASN.

\(\theta\)\(h(\theta)\)\(L(\theta)\) \(E_\theta(N)\)
0.1000 \(=p_0\)1.00000.95000022.782
0.15000.37460.75080730.250
0.18620.00000.50000031.401
0.2500−0.58160.15284723.725
0.3000 \(=p_1\)−1.00000.05000017.245

The two checks hold exactly: \(L(p_0) = 0.950000 = 1 - \alpha\) and \(L(p_1) = 0.050000 = \beta\). At \(\theta = 0.186169\), where \(E_\theta(z) = 0\), the limiting forms give \(L = 0.500000\) — exactly one half, because \(B = 1/A\) makes \(\ln A/(\ln A - \ln B) = 1/2\).

Step 5 — against the fixed-sample test. A fixed-sample test of the same two hypotheses at the same \(\alpha\) and \(\beta\) needs

\[ n = \frac{\left[z_{\alpha}\sqrt{p_0q_0} + z_{\beta}\sqrt{p_1q_1}\right]^{2}} {(p_1 - p_0)^{2}} = 38.889 \;\longrightarrow\; 39. \]
UnderSPRT, \(E(N)\)Fixed sampleSaving
\(H_0\) true22.7823941.6%
\(H_1\) true17.2453955.8%
worst case31.6303918.9%

Interpretation. Under either hypothesis the SPRT uses roughly half the observations, which is what the Wald–Wolfowitz theorem promised. The third row is the honest one: between the two hypotheses the saving shrinks to \(18.9\%\), because there the data supports neither boundary and the path wanders. The peak is at \(\theta = 0.175740\), giving \(31.630\).

A detail worth noticing. That peak is not at \(\theta = 0.186169\) where \(E_\theta(z) = 0\), although it is close: there the ASN is \(31.401\). It is a common belief that the two coincide. Under Wald's approximation they do not, and the maximum sits a little below.

And the practical warning. \(N\) is unbounded — the path can in principle wander between the lines for ever. The probability of that is zero, but the probability of a very long run is not negligible, which is why real sequential designs truncate at some \(N_{\max}\) and accept the small distortion of the error probabilities that truncation causes.

4. Where Sequential Testing Is Actually Used

AND WHERE IT MUST NOT BE
SettingWhy sequential
Acceptance samplingeach item inspected is destroyed or costs money; stopping early on a clearly good or clearly bad lot is the whole point
Clinical trialscontinuing to randomise patients to a treatment already shown harmful is not a statistical question
Online experimentsthe cost of a bad variant accrues continuously
Fault detectionthe alternative is a process running out of control while data accumulates

The error that the whole theory exists to prevent. Running a fixed-sample test, looking at the \(p\) value, and continuing to collect data if it is not yet significant is not a sequential test. It is a fixed-sample test with its size destroyed: repeated looks at the same accumulating data will eventually cross any fixed threshold under the null with probability one. A sequential test is one whose stopping rule was fixed before the data arrived, and whose boundaries were derived from that rule. The distinction between the two is the difference between a valid \(5\%\) test and no test at all.

5. Elements of Decision Theory

TESTING AS A DECISION PROBLEM

The framework — action space, loss function \(L(\theta, a)\), risk \(R(\theta, \delta) = E_\theta L(\theta, \delta(X))\), admissibility, minimax and Bayes rules — is built in Estimation Theory, Unit 4 for the estimation problem, where the action is a number. Only two things change when the action is a choice between two hypotheses.

First, the action space has two points, so the loss function is four numbers rather than a curve. Taking the loss of a correct decision to be zero,

\[ L(\theta_0, a_1) = \ell_{\mathrm{I}}, \qquad L(\theta_1, a_0) = \ell_{\mathrm{II}}, \]

and the risk is \(\ell_{\mathrm{I}}\alpha\) at \(\theta_0\) and \(\ell_{\mathrm{II}}\beta\) at \(\theta_1\). The two error probabilities are the risk function; the whole of this course has been decision theory with the losses left unnamed.

Second, the Bayes rule is explicit. With prior probabilities \(\pi_0\) and \(\pi_1 = 1 - \pi_0\), the posterior expected loss of rejecting is smaller than that of accepting exactly when

\[ \frac{f_{\theta_1}(x)}{f_{\theta_0}(x)} \;>\; \frac{\pi_0\,\ell_{\mathrm{I}}} {\pi_1\,\ell_{\mathrm{II}}}. \]

That is the Neyman–Pearson test again, with the constant \(k\) of Unit 1 now identified: it is the ratio of prior-weighted losses. The lemma chose \(k\) to hit a size; the Bayes rule chooses it from what the errors cost. The two theories produce the same family of tests and disagree only about how to pick the member of it — which is a statement worth carrying away from this course as a whole.

And the minimax rule is the member of that family for which the two risks are equal, \(\ell_{\mathrm{I}}\alpha = \ell_{\mathrm{II}}\beta\) — the choice made when no prior is available and the worst case is what must be controlled.