Non-parametric (distribution-free) tests are statistical tests that do not assume any particular form for the underlying population distribution. They typically work with ranks or signs rather than the actual data values.
| Aspect | Parametric | Non-parametric |
|---|---|---|
| Distribution assumption | Specific (often normal) | None / minimal |
| Data type | Interval / ratio | Ordinal / nominal also OK |
| Test statistic based on | Means, variances | Ranks, signs, counts |
| Power (when assumptions hold) | Higher | Lower, by how much depends on the test: for normal data Mann–Whitney is about 95 % as efficient as the t test (\(3/\pi\)), the sign test only about 64 % (\(2/\pi\)) |
| Sensitivity to outliers | Sensitive | Robust |
The tests of Units 3 and 4 assume the form of the parent population is known (usually normal) and test its parameters — means, variances. They are parametric tests. A non-parametric (N.P.) test does not depend on the particular form of the frequency function from which the samples are drawn, and is not a statement about a parameter of that form. (The \(\chi^2\) test of goodness of fit in Unit 4 is already of this kind.)
Assumptions that remain:
Further advantages the textbook lists: the methods need little sampling theory; they apply to data given only as ranks, signs or grades (A, A+, B, …) and to data on a nominal scale, where no parametric technique applies; and socio-economic data, which are seldom normal, make them common in psychometry, sociology and educational statistics.
Further disadvantages: when a parametric test's assumptions do hold, the parametric test is more powerful; and there is no N.P. test for interactions in analysis of variance unless additivity of the model is assumed.
Source note. The textbook says N.P. tests “are designed to test statistical hypotheses only and not for estimating parameters”. That overstates it: the median has distribution-free confidence intervals built from the same binomial counts as the sign test (§3), and the Hodges–Lehmann estimate comes from the Wilcoxon statistic.
| Parametric tests | Non-parametric tests | |
|---|---|---|
| Observations | independent | independent |
| Variable | discrete or continuous | continuous (for exact tables) |
| Population | form assumed known, usually normal | no assumption about its form |
| Parameters | estimated and tested | the hypothesis is about the distribution or its median |
| Nominal or ordinal data; ranks and grades | cannot be used | can be used |
| Typical measure | mean and variance | median |
| Two independent samples | \(t\)-test | runs test (Wald–Wolfowitz), median test, Mann–Whitney U |
| Paired samples (one pair per unit) | paired \(t\)-test | sign test, Wilcoxon signed-rank test |
| Power when the parametric assumptions hold | higher | lower |
| Robustness to departures from the assumptions | sensitive to them | robust: valid whatever the continuous distribution |
Source note. The textbook's version of this table calls the parametric methods “robust (strong)” and the non-parametric ones “weak”. Robustness is the non-parametric tests' strength, as the table above and §1 say; what they lack is power when the parametric assumptions happen to hold. It also describes paired samples as “two samples of same size”. Pairing means each observation in one sample is matched to one unit in the other (the same student before and after); equal sizes follow from that, but two unrelated samples of equal size are not paired.
Tests whether a sequence of two types of items (or values above/below the median) is random.
Sequence: H T H T H T H T H T H T (12 tosses) — \(n_1 = n_2 = 6\), \(R = 12\) (alternating). \(E(R) = 7,\; \text{Var}(R) \approx 2.73,\; Z = 5/1.65 = 3.03\) ⇒ reject randomness.
Sequence: H H H H H H T T T T T T — \(R = 2,\; Z = (2-7)/1.65 = -3.03\) ⇒ also reject (clustering, not random).
A caution on both examples. With \(n_1 = n_2 = 6\) the samples are below the size at which the normal approximation is reliable, so \(Z\) is only a rough guide here. The exact answer agrees: of the \(\binom{12}{6} = 924\) equally likely arrangements, only 2 give 12 runs and only 2 give 2 runs, so each has probability \(2/924 = 0.002\).
A run is a maximal sequence of letters of one kind, bounded by letters of the other kind (or by the ends); the number of letters in it is its length. In
AA BBB A B AA B AA
there are 7 runs: 4 runs of A and 3 of B; the first has length 2.
Turning numbers into letters. Find the median \(M\) of the sample; write A for an observation \(\le M\) and B for one \(> M\), in the original order. (If \(n\) is odd, one observation equals \(M\); many texts drop it so that A and B are equally many.) Count the runs \(r\), the A's \(n_1\) and the B's \(n_2\).
Decision. For \(n_1, n_2 \le 20\) the runs tables give critical values \(r_1 < r_2\): accept randomness if \(r_1 < r < r_2\), reject if \(r \le r_1\) (too few runs: clustering, a trend) or \(r \ge r_2\) (too many: alternation). Beyond 20, use the normal form above.
Under \(H_0\) (random order) every arrangement of \(n_1\) A's and \(n_2\) B's is equally likely; there are \(\binom{n_1 + n_2}{n_1}\) of them.
Even number of runs, \(r = 2k\). Then there are \(k\) runs of A and \(k\) of B, alternating, starting with either letter (factor 2). Splitting \(n_1\) A's into \(k\) non-empty runs means choosing \(k - 1\) of the \(n_1 - 1\) gaps between them: \(\binom{n_1 - 1}{k - 1}\) ways; likewise \(\binom{n_2 - 1}{k - 1}\) for the B's. So
\[ P(r = 2k) = \frac{2\binom{n_1 - 1}{k - 1}\binom{n_2 - 1}{k - 1}}{\binom{n_1 + n_2}{n_1}} . \]Odd number, \(r = 2k + 1\). One letter has \(k + 1\) runs (and starts and ends the sequence), the other \(k\):
\[ P(r = 2k + 1) = \frac{\binom{n_1 - 1}{k}\binom{n_2 - 1}{k - 1} + \binom{n_1 - 1}{k - 1}\binom{n_2 - 1}{k}}{\binom{n_1 + n_2}{n_1}} . \]The mean. The number of runs is 1 plus the number of places where the letter changes. Each of the \(N - 1\) adjacent pairs (\(N = n_1 + n_2\)) is a change with probability \(2 \cdot \dfrac{n_1}{N} \cdot \dfrac{n_2}{N - 1}\), so
\[ E(r) = 1 + (N - 1)\cdot\frac{2 n_1 n_2}{N(N - 1)} = \frac{2 n_1 n_2}{n_1 + n_2} + 1 . \]The variance, by the same counting over pairs of positions, is the \(\text{Var}(R)\) above. The critical values in the runs tables are read off this exact distribution: \(r_1\) is the largest value with \(P(r \le r_1) \le 0.025\), and \(r_2\) the smallest with \(P(r \ge r_2) \le 0.025\).
Largest possible number of runs. Runs alternate, so the scarcer letter limits them: \(2\min(n_1, n_2)\) runs, plus one if \(n_1 \ne n_2\). It equals \(n_1 + n_2\) only when \(n_1 = n_2\).
Tests \(H_0\): median = \(M_0\) using only the signs of \((X_i - M_0)\). Or for paired samples, signs of \((X_i - Y_i)\).
15 measurements; 4 are below 50 (−), 11 above (+). \(n = 15, S = 11\). Two-tailed P-value from Binomial(15, 0.5): \(2 P(S \ge 11) = 2(0.0592) = 0.118\) ⇒ accept \(H_0\) at 5 %.
Before-after for 12 patients: +9 −2 (1 tie). \(n = 11, S_+ = 9\). \(P(S_+ \ge 9) = 0.0327\); two-tailed p = 0.065 ⇒ accept \(H_0\) at 5 %.
One sample. If \(M_0\) is the population median, then for a continuous variable \(P(X > M_0) = P(X < M_0) = \tfrac12\). The signs of \(x_i - M_0\) are independent, each “+” with probability \(\tfrac12\). Drop any zeros (they carry no sign) and let \(n\) be the number left. The number of + signs, \(S\), is then a count of successes in \(n\) independent trials with \(p = \tfrac12\):
\[ P(S = u) = \binom{n}{u}\left(\tfrac12\right)^u\left(\tfrac12\right)^{n-u} = \binom{n}{u}\left(\tfrac12\right)^n . \]Using the smaller count. The textbook works with \(u\) = the number of the less frequent sign, and the cumulative probability \(p = \sum_{i=0}^{u}\binom{n}{i}(\tfrac12)^n\). Since Binomial\((n, \tfrac12)\) is symmetric, a count as small as \(u\) is equally likely in either sign, so the two-sided p-value is \(2p\). Reject \(H_0\) if \(2p \le \alpha\).
Large \(n\) (the textbook's rule is \(n \ge 25\)): \(E(S) = np = \tfrac n2\) and \(V(S) = npq = \tfrac n4\), so \(z = \dfrac{S - n/2}{\sqrt{n/4}} \sim N(0,1)\) approximately.
Paired samples. With pairs \((x_i, y_i)\) take \(d_i = x_i - y_i\); under \(H_0\) (no difference between the two conditions), \(P(x_i > y_i) = P(x_i < y_i) = \tfrac12\) and the same binomial applies to \(u\) = the number of + signs. The critical values for a two-sided test at level \(\alpha\) are \(x_1\), the largest value with \(P(S \le x_1) \le \alpha/2\), and by symmetry \(x_2 = n - x_1\). Reject \(H_0\) if \(u \le x_1\) or \(u \ge x_2\); accept if \(x_1 < u < x_2\).
Source note. The textbook prints the binomial as \(\binom{n}{a}\) and its cumulative sum with \(\binom{n}{u}\) inside; both mean \(\binom{n}{i}\). For the paired test it states the decision the wrong way round (“if \(u \le x_1\) or \(u \ge x_2\), accept \(H_0\)”) and says to compare the cumulative \(P\) itself with \(\alpha\); for a two-sided test it is \(2P\), as in the one-sample case. Its \(H_0: f_1(x) = f_2(y)\) is more than the test checks: what the signs test is \(P(X > Y) = \tfrac12\), a zero median difference.
More powerful than the sign test because it uses the magnitudes of differences as well.
Differences \(d\): −2, +5, +1, −3, +6, +4. \(|d|\): 1, 2, 3, 4, 5, 6 ranked 1, 2, 3, 4, 5, 6. With signs: +1, −2, −3, +4, +5, +6 (matching order). \(W^+ = 1+4+5+6 = 16,\; W^- = 2+3 = 5\). \(W = 5\). For \(n = 6\), critical at 5 % two-tailed = 0; \(W = 5 > 0\) ⇒ accept \(H_0\).
10 paired differences with \(W^+ = 50, W^- = 5\). \(W = 5\). Critical Wilcoxon for \(n = 10\) at 5 % two-tailed = 8; \(W < 8\) ⇒ reject \(H_0\) — significant difference.
Tests whether two samples come from populations with the same median.
Sample A (10 obs): 4 above, 6 below grand median. Sample B (10 obs): 6 above, 4 below.
| Above M | Below M | |
|---|---|---|
| A | 4 | 6 |
| B | 6 | 4 |
Expected = 5 each. \(\chi^2 = 4(1)/5 = 0.80\). df = 1, \(\chi^2_{0.05} = 3.84\) ⇒ accept \(H_0\).
If Sample A has 9 above / 1 below and Sample B has 1 above / 9 below: \(\chi^2 = 4(4)^2/5 = 12.8\) ⇒ reject \(H_0\); medians differ.
Combine the two samples (sizes \(n_1, n_2\)) and find their common median \(M\). Count the observations \(\ge M\): \(m_1\) in the first sample, \(m_2\) in the second.
| Sample I | Sample II | Total | |
|---|---|---|---|
| \(\ge M\) | \(m_1\) | \(m_2\) | \(m_1 + m_2\) |
| \(< M\) | \(n_1 - m_1\) | \(n_2 - m_2\) | \(n_1 + n_2 - (m_1 + m_2)\) |
| Total | \(n_1\) | \(n_2\) | \(n_1 + n_2\) |
Where the hypergeometric comes from. Under \(H_0\) the two samples are one population, so given the margins, which \(m_1 + m_2\) of the \(n_1 + n_2\) observations land at or above \(M\) is a random choice. The chance that exactly \(m_1\) of them come from sample I is the hypergeometric probability
\[ P(m_1) = \frac{\binom{n_1}{m_1}\binom{n_2}{m_2}}{\binom{n_1 + n_2}{m_1 + m_2}} . \]The test uses a tail, not a single term. The p-value is the probability of a split at least as uneven as the one observed: \(P(M_1 \ge m_1)\), the sum of the terms from \(m_1\) upward (double it for a two-sided test). Reject \(H_0\) if it is at most \(\alpha\).
Larger samples (both \(n_1, n_2 > 10\)): use the \(2 \times 2\) \(\chi^2\) of Unit 4, \(\chi^2 = \dfrac{N(ad - bc)^2}{(a+b)(c+d)(a+c)(b+d)}\) with 1 d.f., with Yates' correction when a cell frequency is below 5.
Source note. The textbook compares the single probability \(P(m_1)\) with \(\alpha\). A single term can be small even when nothing unusual has happened (with large samples every exact outcome is improbable), so the comparison must be with the tail sum. It also prints the lower-left cell as \(n_1 - m_2\); it is \(n_1 - m_1\). And it introduces the test by saying the sign and Wilcoxon tests need equal sample sizes; what they need is paired observations.
The standard non-parametric test for comparing two independent samples: for normal data it is about 95 % as efficient as the two-sample t test, and for heavy-tailed data it can beat it. Tests whether the distributions are identical (or whether one tends to produce larger values).
Sample 1 (n₁=4): 12, 15, 9, 18. Sample 2 (n₂=5): 7, 10, 14, 11, 13.
Combined ranks: 7→1, 9→2, 10→3, 11→4, 12→5, 13→6, 14→7, 15→8, 18→9.
\(R_1 = 2 + 5 + 8 + 9 = 24,\; R_2 = 1+3+4+6+7 = 21\).
\(U_1 = 24 - 4 \cdot 5/2 = 14,\; U_2 = 21 - 5 \cdot 6/2 = 6\). \(U = 6\). Critical for (4,5) at 5 % two-tailed = 1; \(U = 6 > 1\) ⇒ accept \(H_0\).
Two methods scored over n₁ = 12, n₂ = 15 with \(R_1 = 250\). \(U_1 = 250 - 12(13)/2 = 172\). \(Z = (172 - 90)/\sqrt{12 \cdot 15 \cdot 28/12} = 82/\sqrt{420} = 82/20.49 = 4.00\) ⇒ reject \(H_0\); methods differ.
Tests whether two independent samples come from the same population. Uses runs in the combined ranked sample.
Sample A: 1,2,3,4,5; Sample B: 6,7,8,9,10. Combined ordered: A A A A A B B B B B → \(R = 2\). \(E(R) = 6, \text{Var}(R) = 2.22\). \(Z = (2-6)/1.49 = -2.68\) ⇒ reject \(H_0\); distributions differ.
Sample A: 1, 4, 6, 9; Sample B: 2, 3, 5, 7, 8, 10. Order: A B B A B A B B A B → \(R = 8\). \(n_1=4, n_2=6, E(R)=5.8, \text{Var}(R)=2.03\). \(Z = (8-5.8)/1.42 = 1.55\) ⇒ accept \(H_0\) at 5 %.
The two-sample runs test uses only the labels in the combined order, and the exact distribution of §2 (with \(n_1\) A's and \(n_2\) B's) applies unchanged. But if a value occurs in both samples, the order within that tie is arbitrary, and the number of runs depends on the choice. The textbook's saree data (Worked Problem 9) have five such ties; over every way of ordering them, \(r\) ranges from 18 to 26. The honest report gives that range (here every value leads to the same decision); a common alternative is to break ties at random.
Source note. The textbook gives the possible numbers of runs as \(2, 3, \ldots, n_1 + n_2\). The largest is \(2\min(n_1, n_2)\), plus one when \(n_1 \ne n_2\) (§2).
The Kruskal–Wallis test compares the locations of \(k \ge 3\) independent groups without assuming normality — the rank-based analogue of one-way ANOVA. Pool all \(N\) observations, rank them, and let \(R_j\) be the rank-sum of group \(j\) (size \(n_j\)).
Reject \(H_0\) (equal locations) when \(H\) exceeds \(\chi^2_{k-1}\) at level \(\alpha\).
Three groups — A: 12, 15, 18; B: 20, 22, 25; C: 8, 10, 14 (\(N = 9\)). Ranking all nine gives rank-sums \(R_A = 14,\ R_B = 24,\ R_C = 7\). Then
\(H = \dfrac{12}{9\cdot10}\!\left(\dfrac{14^2}{3}+\dfrac{24^2}{3}+\dfrac{7^2}{3}\right) - 3(10) = 6.49\). Since \(6.49 > \chi^2_{2,\,0.05} = 5.99\), the groups differ significantly.
The Kolmogorov–Smirnov test checks goodness of fit by comparing the empirical CDF \(F_n(x)\) with a fully specified theoretical CDF \(F_0(x)\) (one-sample), using the largest vertical gap:
Reject \(H_0\) (data follow \(F_0\)) if \(D\) exceeds the tabulated critical value \(D_{\alpha,n}\). Unlike the \(\chi^2\) goodness-of-fit test, K–S needs no class grouping and works well for small, continuous samples. (A two-sample version compares two empirical CDFs.)
Test whether 0.15, 0.32, 0.48, 0.61, 0.83 come from Uniform\((0,1)\). Comparing the step ECDF with \(F_0(x) = x\) gives \(D = 0.19\). Since \(0.19 < D_{0.05,\,5} = 0.563\), \(H_0\) is not rejected — the data are consistent with the uniform.
Eleven problems in the order the textbook sets them: two on the one-sample sign test, two on the paired sign test, two on the runs test for randomness, three on the two-sample runs test and two on the median test. Each states \(H_0\), turns the data into signs, runs or counts, and decides at the 5% level. Where a sample is small the decision is checked against the exact null distribution (binomial, runs or hypergeometric) as well as the textbook's tables.
Source note. Every count and statistic below was recomputed from the data, and every table value from the exact distribution. The counts all agree with the textbook. Its paired sign-test critical values are misprinted, one conclusion answers a different question from the one the test asks, and one probability is compared with \(\alpha\) in the wrong way; each is corrected in place with the printed version recorded.
Twenty students' marks: 93, 88, 107, 115, 82, 97, 103, 86, 113, 107, 112, 90, 98, 93, 99, 103, 100, 101, 96, 104. Test the hypothesis that the median mark in the school is 99.
(1) \(H_0: M = 99\). (2) \(H_1: M \ne 99\). Two-sided, 5%.
(3) The signs of \(x_i - 99\), in order:
− − + + − − + − + + + − − − 0 + + + − +
10 plus, 9 minus and one zero (the mark 99 itself), which is dropped: \(n = 19\). The rarer sign occurs \(u = 9\) times, and
\[ p = P(S \le 9) = \sum_{i=0}^{9}\binom{19}{i}\left(\tfrac12\right)^{19} = \frac12 , \]exactly one half, because with \(n = 19\) the values 0–9 and 10–19 are mirror images.
(4) \(2p = 1 > 0.05\), so \(H_0\) is accepted: the data are consistent with a median of 99. Ten above and nine below is as even a split as 19 signs allow.
Printing note. The textbook writes the sum with \(\binom{19}{9}\) inside; it is \(\binom{19}{i}\).
Of 100 students' weights, 30 are below 60 kg and 70 above. Is 60 kg the median weight (are the students split equally)?
(1) \(H_0: M = 60\). (2) \(H_1: M \ne 60\).
(3) \(n = 100\) is large, so the normal form applies, with \(u = 30\):
\[ z = \frac{u - n/2}{\sqrt{n/4}} = \frac{30 - 50}{\sqrt{25}} = \frac{-20}{5} = -4 . \](4) \(|z| = 4 > 1.96\), so \(H_0\) is rejected: 60 kg is not the median; more students are above it than below. (The exact binomial p-value, \(2P(S \le 30) = 0.00008\), agrees.)
Seventeen students were graded before and after a training course. Is there a difference?
| Student | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Test 1 \(x\) | 2 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 2 | 3 | 2 | 2 | 5 | 2 | 5 | 3 | 1 |
| Test 2 \(y\) | 4 | 4 | 5 | 5 | 3 | 2 | 5 | 3 | 1 | 5 | 5 | 5 | 4 | 5 | 5 | 5 | 5 |
| sign of \(x - y\) | − | − | − | − | 0 | + | − | 0 | + | − | − | − | + | − | 0 | − | − |
(1) \(H_0\): no difference between the two tests, \(P(x > y) = \tfrac12\). (2) \(H_1\): there is a difference. Two-sided.
(3) 3 plus, 11 minus, 3 zeros; drop the zeros: \(n = 14\), \(u = 3\) plus signs. For Binomial\((14, \tfrac12)\), \(P(S \le 2) = 0.0065\) and \(P(S \le 3) = 0.0287\), so the largest \(x_1\) with \(P(S \le x_1) \le 0.025\) is \(x_1 = 2\), and by symmetry \(x_2 = 14 - 2 = 12\).
(4) \(2 < u = 3 < 12\): \(H_0\) is accepted at 5%. The exact two-sided p-value is \(2P(S \le 3) = 0.057\): close, but not significant.
Correction note. The textbook reads the critical values as “2 and 11” and, a line later, as \(x_1 = 2\), \(x_2 = 10\). For \(n = 14\) the upper value is 12, the mirror image of 2; the decision is unchanged.
Thirty students took two tests:
| Student | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|
| Test I | 133 | 146 | 136 | 172 | 141 | 106 | 159 | 141 | 142 | 140 |
| Test II | 141 | 151 | 99 | 145 | 179 | 161 | 168 | 151 | 132 | 180 |
| sign | − | − | + | + | − | − | − | − | + | − |
| Student | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 |
| Test I | 174 | 156 | 82 | 138 | 154 | 163 | 175 | 150 | 134 | 134 |
| Test II | 163 | 133 | 158 | 180 | 204 | 157 | 215 | 145 | 167 | 137 |
| sign | + | + | − | − | − | + | − | + | − | − |
| Student | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 |
| Test I | 160 | 133 | 149 | 129 | 140 | 162 | 115 | 104 | 121 | 181 |
| Test II | 136 | 170 | 157 | 168 | 188 | 188 | 160 | 146 | 123 | 177 |
| sign | + | − | − | − | − | − | − | − | − | + |
Can the students' performance be taken as the same in both?
(1) \(H_0\): no difference. (2) \(H_1\): a difference.
(3) 9 plus signs, no zeros: \(u = 9\), \(n = 30 \ge 25\), so
\[ z = \frac{9 - 30/2}{\sqrt{30/4}} = \frac{-6}{2.739} = -2.191 . \](4) \(|z| = 2.191 > 1.96\), so \(H_0\) is rejected: the scores differ, Test II being higher for 21 of the 30 students. The exact binomial check, \(2P(S \le 9) = 0.043\), agrees.
Test the randomness of 109, 124, 173, 167, 148, 132, 168, 165, 118, 112, 114, 164, 180, 123, 180, 152.
(1) \(H_0\): the order is random. (2) \(H_1\): it is not.
(3) In ascending order the 8th and 9th values are 148 and 152, so \(M = (148 + 152)/2 = 150\). Writing A for \(\le 150\) and B for \(> 150\) in the original order:
AA BB AA BB AAA BB A BB
\(r = 8\) runs, \(n_1 = 8\) A's and \(n_2 = 8\) B's. From the runs tables (or the exact distribution of §2), \(r_1 = 4\) and \(r_2 = 14\).
(4) \(4 < 8 < 14\), so \(H_0\) is accepted: the sequence may be regarded as random. Eight is close to the expected \(2(8)(8)/16 + 1 = 9\).
A coin thrown 37 times gives 25 heads, in 13 runs. Is the sequence random, and is the coin unbiased?
Randomness. \(H_0\): the order of heads and tails is random. \(n_1 = 25\), \(n_2 = 12\), \(r = 13\); \(n_1 > 20\), so the normal form:
\[ E(r) = \frac{2(25)(12)}{37} + 1 = 17.216, \] \[ V(r) = \frac{2(25)(12)\,[2(25)(12) - 25 - 12]}{37^2 \times 36} = \frac{600 \times 563}{49284} = 6.854 , \] \[ z = \frac{13 - 17.216}{\sqrt{6.854}} = \frac{-4.216}{2.618} = -1.61 . \]\(|z| = 1.61 < 1.96\): the order may be regarded as random.
Bias is a different question. The runs test looks only at the order of the results, never at how many heads there are, so it cannot say whether the coin is fair. That is a test of \(p = \tfrac12\) on the count of heads (a sign test):
\[ z = \frac{25 - 37/2}{\sqrt{37/4}} = \frac{6.5}{3.041} = 2.14 . \]\(|z| = 2.14 > 1.96\) (exact two-sided p = 0.047): at 5% the coin is judged biased towards heads, although its sequence shows no departure from randomness.
Correction note. The textbook accepts \(H_0\) from the runs test and concludes “the coin is unbiased”; the runs test does not test that. Its working also prints the numerator without the minus sign, the variance with \(2(5)(12)(3 \times 25 \times 12 - 25 - 12)\) for \(2(25)(12)(2 \times 25 \times 12 - 25 - 12)\), and \(z = 1.61\) for \(-1.61\).
Scores of 6 Secretariat clerks: 40, 35, 52, 60, 46, 55; of 7 Directorate clerks: 47, 56, 42, 57, 50, 57, 62. Do the two sets of scores follow the same distribution?
(1) \(H_0\): the two populations have the same distribution. (2) \(H_1\): they do not.
(3) Combined in ascending order, A = Secretariat, B = Directorate:
| 35 | 40 | 42 | 46 | 47 | 50 | 52 | 55 | 56 | 57 | 57 | 60 | 62 |
| A | A | B | A | B | B | A | A | B | B | B | A | B |
AA B A BB AA BBB A B: \(r = 8\) runs, \(n_1 = 6\), \(n_2 = 7\). The critical values are \(r_1 = 3\), \(r_2 = 12\).
(4) \(3 < 8 < 12\): \(H_0\) is accepted; the two offices' scores may follow the same distribution.
A team's wins and losses in order: W L WW L W LLL W LL WW L WW L W L W L W LL W L W L. Is the order random?
Although the textbook sets this under the two-sample test, it is the runs test for randomness of §2: one sequence, two kinds of letter. The exact distribution is the same.
(1) \(H_0\): the order is random. (2) \(H_1\): it is not.
(3) \(n_1 = 14\) W's, \(n_2 = 15\) L's, \(r = 22\) runs. The exact critical values are \(r_1 = 9\), \(r_2 = 22\).
(4) \(r = 22 \ge r_2\), so \(H_0\) is rejected: the results alternate more often than chance allows (\(P(r \ge 22) = 0.011\), two-sided p = 0.021).
Quality scores of sarees made at two temperatures. Sample I (18): 235, 256, 315, 258, 220, 250, 225, 224, 247, 207, 248, 254, 206, 230, 251, 268, 245, 225. Sample II (23): 258, 228, 250, 225, 243, 249, 243, 229, 237, 239, 222, 206, 227, 211, 206, 220, 226, 225, 214, 232, 236, 205, 255. Is the quality the same? (Use the runs test.)
(1) \(H_0\): same distribution at both temperatures. (2) \(H_1\): not.
(3) Combined in ascending order, A = Sample I, B = Sample II, ties as the textbook orders them:
B A BB A BB A BB AAA BBBBBB A B A BBBBB AAA B A B AA B AA B AA
\(r = 22\), \(n_1 = 18\), \(n_2 = 23\). As \(n_2 > 20\), use the normal form:
\[ E(r) = \frac{2(18)(23)}{41} + 1 = 21.195, \qquad V(r) = \frac{828 \times 787}{41^2 \times 40} = 9.691, \] \[ z = \frac{22 - 21.195}{\sqrt{9.691}} = \frac{0.805}{3.113} = 0.26 . \](4) \(|z| = 0.26 < 1.96\): \(H_0\) is accepted; the quality may be taken as the same.
Ties. The values 206, 220, 225, 250 and 258 occur in both samples, and the order within each tie is arbitrary. Over every possible ordering, \(r\) runs from 18 to 26, and \(z\) from \(-1.03\) to \(1.54\) (§7). Every one of them accepts \(H_0\), so here the ties do not matter.
Daily numbers of defectives from machine A: 26, 27, 31, 26, 19, 21, 20, 25, 30; from machine B: 23, 28, 26, 24, 22, 19. Do the two samples come from the same population?
(1) \(H_0\): same population. (2) \(H_1\): not. Two-sided.
(3) Combined in order: 19A 19B 20A 21A 22B 23B 24B 25A 26A 26A 26B 27A 28B 30A 31A. With 15 values the median is the 8th, \(M = 25\).
| Machine A | Machine B | Total | |
|---|---|---|---|
| \(\ge 25\) | 6 | 2 | 8 |
| \(< 25\) | 3 | 4 | 7 |
| Total | 9 | 6 | 15 |
The samples are small, so the exact hypergeometric. The probability of this table alone is
\[ P(m_1 = 6) = \frac{\binom96\binom62}{\binom{15}{8}} = \frac{84 \times 15}{6435} = 0.1958, \]and the p-value adds the more extreme tables (\(m_1 = 7\): \(36 \times 6 = 216\); \(m_1 = 8\): 9):
\[ P(M_1 \ge 6) = \frac{1260 + 216 + 9}{6435} = \frac{1485}{6435} = \frac{3}{13} = 0.2308 . \](4) Doubled for a two-sided test, the p-value is \(0.4615 > 0.05\) (and even the one-sided 0.2308 is far above it): \(H_0\) is accepted; the samples may come from the same population.
Correction note. The textbook compares the single term 0.1958 with \(\alpha\) (§5). The verdict is the same here, but the p-value is the tail sum, 0.2308.
The data of Worked Problem 9. Do the two samples come from the same population?
(1) \(H_0\): same population. (2) \(H_1\): not.
(3) With 41 values the median is the 21st in order, \(M = 232\).
| Sample I | Sample II | Total | |
|---|---|---|---|
| \(\ge 232\) | 11 (\(a\)) | 10 (\(b\)) | 21 |
| \(< 232\) | 7 (\(c\)) | 13 (\(d\)) | 20 |
| Total | 18 | 23 | 41 |
Both samples exceed 10, so
\[ \chi^2 = \frac{N(ad - bc)^2}{(a+b)(c+d)(a+c)(b+d)} = \frac{41(11 \times 13 - 10 \times 7)^2}{21 \times 20 \times 18 \times 23} = \frac{41 \times 73^2}{173880} = 1.2566 . \](4) \(1.2566 < \chi^2_{0.05,1} = 3.84\): \(H_0\) is accepted, as by the runs test in Worked Problem 9. (With Yates' correction \(\chi^2 = 0.650\), the same verdict.)
| Question | Test | Parametric counterpart |
|---|---|---|
| Is a sequence random? | One-sample runs test | — |
| Test population median (one sample) | Sign / Wilcoxon signed-rank | One-sample t |
| Paired samples | Sign / Wilcoxon signed-rank | Paired t |
| Two independent medians | Median test | Two-sample t |
| Two independent samples (general) | Mann–Whitney U / Wald–Wolfowitz | Two-sample t |
Where this course has taken you. Five units, one narrowing question. Unit 1 estimated a parameter; Unit 2 turned estimation into a decision; Unit 3 made that decision when the sample was large enough for the CLT; Unit 4 when it was not, at the cost of assuming a normal population; and this unit when even normality is unavailable. Each step gave something up and bought robustness with it.
What is not here: the parametric way of comparing three or more groups, which is Analysis of Variance (its rank-based counterpart, Kruskal–Wallis, is section 8 above). To pick a test by what you are asking rather than by unit number, use the test chooser.