When the sample size \(n\) is large (rule of thumb: \(n > 30\)) and the population SD \(\sigma\) is known (or can be reliably estimated by \(s\)), the sampling distribution of any statistic that is a sum or average is approximately normal by the Central Limit Theorem.
Hence we use the standard normal \(Z\)-statistic for testing.
The whole theory of a test of significance turns on the sample size. In practice, if the sample size is greater than or equal to 30, the sample is called a large sample. The number is a working threshold, not a theorem.
What earns it is that for large \(n\) almost every distribution met in practice — Binomial, Poisson, Negative Binomial, Exponential — tends to the normal. So the normal distribution can be used to test the hypothesis whatever the parent population was, and the area property of the standard normal curve does the rest. If \(X \sim N(\mu, \sigma^2)\) then
\[ z = \frac{X - \mu}{\sigma} \sim N(0,1), \qquad\text{usually written}\qquad z = \frac{X - E(X)}{\sqrt{\operatorname{Var}(X)}} \sim N(0,1). \]The second form is the one that generalises: it is the shape of every statistic on this page.
From the standard normal tables,
\[ P(-3 < z < 3) = 0.9973, \qquad P(|z| \le 3) = 0.9973, \] \[ P(|z| > 3) = 1 - P(|z| \le 3) = 0.0027 . \]A standard normal variate is expected to lie between \(\pm 3\). So if \(|z| > 3\) the null hypothesis \(H_0\) is always rejected, at any level of significance in ordinary use; otherwise, if \(|z| \le 3\), \(H_0\) may be accepted. Several of the problems below finish on this rule alone, without consulting a table.
Correction. For the family-income population, the textbook offers “a group of families of low income group” as its example of a sample. That group is a subset of the population, but a sample drawn only from low-income families cannot tell us the average income of the locality: it is biased by construction. A sample that is to stand for the population must be drawn from all of it, at random.
A statistic is computed from observations that are themselves random (another sample would give other values), so a statistic is a random variable. Its probability distribution is its sampling distribution. For a random sample from \(N(\mu, \sigma^2)\), for example, \(\bar x \sim N(\mu, \sigma^2/n)\).
The standard error (S.E.) of a statistic is the standard deviation of its sampling distribution.
Note. The textbook gives the reason a statistic is random as “sample observations are independent”. Independence is what makes the formulae below simple, but a statistic is random because it is a function of random observations, independent or not.
Population \(\{2, 4, 6, 8\}\): \(\mu = 5\), \(\sigma^2 = \tfrac{9 + 1 + 1 + 9}{4} = 5\). Draw \(n = 2\) units with replacement: there are \(4 \times 4 = 16\) equally likely samples, and their means are
| \(\bar x\) | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|
| number of samples | 1 | 2 | 3 | 4 | 3 | 2 | 1 |
The mean of this sampling distribution is \(\tfrac{80}{16} = 5 = \mu\), and its variance is \(\tfrac{1}{16}\big[9 + 2(4) + 3(1) + 0 + 3(1) + 2(4) + 9\big] = \tfrac{40}{16} = \tfrac52 = \tfrac{\sigma^2}{n}\). So \(\text{S.E.}(\bar x) = \sqrt{5/2} = \sigma/\sqrt n\), exactly as the formula below says.
with \(Q = 1 - P\). The two-sample forms assume the samples are independent, so that the variances add. Sections 2–5 derive each one.
| Significance α | One-tailed | Two-tailed |
|---|---|---|
| 0.10 | 1.282 | 1.645 |
| 0.05 | 1.645 | 1.96 |
| 0.01 | 2.326 | 2.576 |
The critical value (or significant value, or tabulated value) is the value of a test statistic that separates the rejection region from the acceptance region. It depends on two things: the level of significance, and the alternative hypothesis — because it is \(H_1\) that decides whether the test is two-tailed or one-tailed. The table above is the answer; this is where it comes from.
Let the critical value at level \(\alpha\) be \(z_\alpha\). By the definition of \(\alpha\),
\[ P(|z| > z_\alpha) = \alpha \;\Longrightarrow\; P(z > z_\alpha) + P(z < -z_\alpha) = \alpha . \]The standard normal curve is symmetric, so the two tails are equal:
\[ P(z > z_\alpha) + P(z > z_\alpha) = \alpha \;\Longrightarrow\; 2P(z > z_\alpha) = \alpha \;\Longrightarrow\; P(z > z_\alpha) = \frac{\alpha}{2}, \]and likewise \(P(z < -z_\alpha) = \alpha/2\). Equivalently the middle carries the rest:
\[ P(-z_\alpha < z < z_\alpha) = 1 - \alpha . \]For a one-tailed test the whole of \(\alpha\) sits in a single tail. For a right-tailed test \(z_\alpha\) is fixed by \(P(z > z_\alpha) = \alpha\); for a left-tailed test \(-z_\alpha\) is fixed by \(P(z < -z_\alpha) = \alpha\). Both are read off the same table entry,
\[ P(z < z_\alpha) = 1 - \alpha . \]Two-tailed, 5%. From the standard normal tables,
\[ P(-1.96 < z < 1.96) = 0.95,\qquad P(|z| \le 1.96) = 0.95,\qquad P(|z| > 1.96) = 0.05 . \]So if \(|z| > 1.96\), \(H_0\) may be rejected at the 5% level; otherwise, if \(|z| \le 1.96\), \(H_0\) may be accepted.
Two-tailed, 1%.
\[ P(-2.58 < z < 2.58) = 0.99,\qquad P(|z| \le 2.58) = 0.99,\qquad P(|z| > 2.58) = 0.01 . \]So if \(|z| > 2.58\), \(H_0\) may be rejected at the 1% level.
One-tailed, 5%.
\[ P(z > 1.645) = 1 - P(-\infty < z < 1.645) = 1 - 0.95 = 0.05 . \]One-tailed, 1%.
\[ P(z > 2.33) = 1 - P(-\infty < z < 2.33) = 1 - 0.99 = 0.01 . \]Notice that at the same \(\alpha\) the one-tailed value is always the smaller: putting the whole rejection region in one tail makes it easier to reach.
Every one of the thirty-one worked problems below runs through these same five steps, in this order. Learning the shape once is worth more than learning nine formulae.
Step 4 is the only one that changes from test to test, and all it ever needs is \(E(t)\) and \(\text{S.E.}(t)\) for the statistic in hand. The subsections that follow derive exactly those two quantities, nine times over.
\(H_0: \mu = \mu_0\). Test statistic:
\[ Z \;=\; \dfrac{\bar X - \mu_0}{\sigma/\sqrt n} \]If \(\sigma\) is unknown but \(n\) large, replace \(\sigma\) by sample SD \(s\).
Under \(H_0\), \(Z \sim N(0, 1)\).
Let \(x_1, x_2, \ldots, x_n\) be a random sample of size \(n\) drawn from a normal population with mean \(\mu\) and variance \(\sigma^{2}\), with \(\sigma^{2}\) known.
Hypotheses. \(H_0: \mu = \mu_0\) — the sample has been drawn from that population, or equivalently there is no significant difference between the sample mean and the population mean. Against it, \(H_1: \mu \ne \mu_0\), or \(\mu > \mu_0\), or \(\mu < \mu_0\).
Statistic. Since the \(x_i\) are normal, the sample mean \(\bar x\) is a normal variate too, and
\[ \bar x \sim N\!\left(\mu,\; \frac{\sigma^{2}}{n}\right) \;\Longrightarrow\; E(\bar x) = \mu, \qquad \text{S.E.}(\bar x) = \frac{\sigma}{\sqrt n}. \]Feeding those two into the standard shape \(z = \big(t - E(t)\big)/\text{S.E.}(t)\),
\[ z = \frac{\bar x - E(\bar x)}{\text{S.E.}(\bar x)} = \frac{\bar x - \mu}{\sigma/\sqrt n} \sim N(0,1), \qquad\text{and under } H_0,\qquad z = \frac{\bar x - \mu_0}{\sigma/\sqrt n} \sim N(0,1). \]If \(\sigma\) is not known we estimate it by the sample standard deviation, \(\hat\sigma = s\), which for a large sample costs nothing:
\[ z = \frac{\bar x - \mu_0}{\hat\sigma/\sqrt n} = \frac{\bar x - \mu_0}{s/\sqrt n} \sim N(0,1). \]That substitution is exactly what fails for a small sample, and is why Unit 4 needs Student's \(t\).
A drug company claims their tablets weigh on average 500 mg with σ = 6 mg. A sample of 64 tablets gives \(\bar X = 498.5\) mg. Test at 5 %.
\(Z = (498.5 - 500)/(6/8) = -2.0\). \(|Z| = 2.0 > 1.96\) ⇒ reject \(H_0\). The mean weight differs significantly.
A school principal claims average IQ = 100 (σ = 15). Sample of 100 students has \(\bar X = 104\). Test if IQ is higher (one-tailed).
\(Z = (104 - 100)/(15/10) = 2.67 > 1.645\) ⇒ reject \(H_0\) at 5 %. Evidence that mean IQ > 100.
\(H_0: \mu_1 = \mu_2\). Independent samples of sizes \(n_1, n_2\) (both large) with means \(\bar X_1, \bar X_2\) and SDs \(\sigma_1, \sigma_2\). Test statistic:
\[ Z \;=\; \dfrac{\bar X_1 - \bar X_2}{\sqrt{\sigma_1^2/n_1 + \sigma_2^2/n_2}} \;\sim\; N(0, 1). \]Let \(x_1, \ldots, x_{n_1}\) be a random sample of size \(n_1\) from a normal population with mean \(\mu_1\) and variance \(\sigma_1^{2}\), and let \(y_1, \ldots, y_{n_2}\) be another sample from a second normal population with mean \(\mu_2\) and variance \(\sigma_2^{2}\).
For large \(n_1\) and \(n_2\) both sample means are normal:
\[ \bar x \sim N\!\left(\mu_1, \frac{\sigma_1^{2}}{n_1}\right), \qquad \bar y \sim N\!\left(\mu_2, \frac{\sigma_2^{2}}{n_2}\right). \]Take the standard normal variate for the difference \(\bar x - \bar y\). Its mean is the difference of the means, and — because the two samples are independent — its variance is the sum of the variances, not the difference:
\[ E(\bar x - \bar y) = E(\bar x) - E(\bar y) = \mu_1 - \mu_2, \] \[ \operatorname{Var}(\bar x - \bar y) = \operatorname{Var}(\bar x) + \operatorname{Var}(\bar y) = \frac{\sigma_1^{2}}{n_1} + \frac{\sigma_2^{2}}{n_2}, \qquad \text{S.E.}(\bar x - \bar y) = \sqrt{\frac{\sigma_1^{2}}{n_1} + \frac{\sigma_2^{2}}{n_2}}. \]Hence
\[ z = \frac{\bar x - \bar y - (\mu_1 - \mu_2)} {\sqrt{\dfrac{\sigma_1^{2}}{n_1} + \dfrac{\sigma_2^{2}}{n_2}}} \sim N(0,1), \qquad\text{and under } H_0: \mu_1 = \mu_2,\qquad z = \frac{\bar x - \bar y} {\sqrt{\dfrac{\sigma_1^{2}}{n_1} + \dfrac{\sigma_2^{2}}{n_2}}} \sim N(0,1). \]Three cases then arise, and the worked problems below use all three.
Case 1 — a common, known \(\sigma\). If \(\sigma_1^{2} = \sigma_2^{2} = \sigma^{2}\) is known, it comes outside:
\[ z = \frac{\bar x - \bar y}{\sigma\sqrt{\dfrac{1}{n_1} + \dfrac{1}{n_2}}} \sim N(0,1). \]Case 2 — neither \(\sigma\) known. Estimate each from its own sample, \(\hat\sigma_1 = s_1\) and \(\hat\sigma_2 = s_2\):
\[ z = \frac{\bar x - \bar y} {\sqrt{\dfrac{s_1^{2}}{n_1} + \dfrac{s_2^{2}}{n_2}}} \sim N(0,1). \]Case 3 — a common but unknown \(\sigma\). Pool the two samples into a single estimate, weighting each by its own size, and use Case 1 with it:
\[ \hat\sigma^{2} = \frac{n_1 s_1^{2} + n_2 s_2^{2}}{n_1 + n_2}, \qquad z = \frac{\bar x - \bar y}{\hat\sigma\sqrt{\dfrac{1}{n_1} + \dfrac{1}{n_2}}} \sim N(0,1). \]Boys: \(n_1 = 100,\; \bar X_1 = 165\) cm, \(\sigma_1 = 5\). Girls: \(n_2 = 80,\; \bar X_2 = 158\) cm, \(\sigma_2 = 4\).
\(Z = (165 - 158)/\sqrt{25/100 + 16/80} = 7/\sqrt{0.45} = 10.43\) ⇒ highly significant.
Two methods of teaching: \(n_1 = 50,\bar X_1 = 72,\sigma_1 = 8\); \(n_2 = 60,\bar X_2 = 75,\sigma_2 = 7\). Test at 5 %.
\(Z = (72 - 75)/\sqrt{64/50 + 49/60} = -3/\sqrt{2.097} = -2.07\). \(|Z| > 1.96\) ⇒ reject \(H_0\).
\(H_0: p = p_0\). Sample proportion \(\hat p = X/n\). Under \(H_0\):
\[ Z \;=\; \dfrac{\hat p - p_0}{\sqrt{p_0(1-p_0)/n}} \;\sim\; N(0, 1). \]Let \(A\) be an attribute observed on \(n\) persons; count the presence of the attribute as a success and its absence as a failure. Let \(x\) be the number of successes in \(n\) independent trials with probability \(P\) of success at each. Then \(x\) is binomial, so
\[ E(x) = nP, \qquad \operatorname{Var}(x) = nPQ, \qquad Q = 1 - P, \]and for large \(n\) the binomial tends to the normal, which is what licenses a \(z\) test at all.
Hypotheses. \(H_0: P = P_0\) — the sample proportion is coming from the population proportion. Against it \(H_1: P \ne P_0\), or \(P > P_0\), or \(P < P_0\).
Statistic. Let the sample proportion be \(p = x/n\). Dividing a random variable by the constant \(n\) divides its mean by \(n\) and its variance by \(n^{2}\):
\[ E(p) = E\!\left(\frac{x}{n}\right) = \frac{E(x)}{n} = \frac{nP}{n} = P, \] \[ \operatorname{Var}(p) = \operatorname{Var}\!\left(\frac{x}{n}\right) = \frac{1}{n^{2}}\operatorname{Var}(x) = \frac{1}{n^{2}}\,nPQ = \frac{PQ}{n}, \qquad \text{S.E.}(p) = \sqrt{\frac{PQ}{n}} . \]So the sample proportion is an unbiased estimator of \(P\), and
\[ z = \frac{p - E(p)}{\text{S.E.}(p)} = \frac{p - P}{\sqrt{\dfrac{PQ}{n}}} \sim N(0,1), \qquad\text{and under } H_0: P = P_0,\qquad z = \frac{p - P_0}{\sqrt{\dfrac{P_0 Q_0}{n}}} \sim N(0,1). \]Inference. Compute \(|z|\), read the tabulated \(z_\alpha\) at level \(\alpha\) from the standard normal tables. If \(|z| > z_\alpha\) we may reject \(H_0\); otherwise, if \(|z| \le z_\alpha\), \(H_0\) may be accepted.
Out of 1000 customers, 540 prefer brand A. Test \(H_0: p = 0.5\) at 5 %.
\(\hat p = 0.54\). \(Z = (0.54 - 0.5)/\sqrt{0.5(0.5)/1000} = 0.04/0.01581 = 2.53\) ⇒ reject \(H_0\). Significant majority for A.
A coin is tossed 400 times yielding 230 heads. Test fairness.
\(\hat p = 0.575\). \(Z = (0.575 - 0.5)/\sqrt{0.25/400} = 0.075/0.025 = 3.0\) ⇒ reject \(H_0\). Coin is biased.
\(H_0: p_1 = p_2\). Pooled proportion \(\hat p = (X_1 + X_2)/(n_1 + n_2)\). Test statistic:
\[ Z \;=\; \dfrac{\hat p_1 - \hat p_2}{\sqrt{\hat p(1-\hat p)\left(\dfrac{1}{n_1} + \dfrac{1}{n_2}\right)}}. \]Here two different populations are compared with respect to one attribute. Let \(x_1\) and \(x_2\) be the numbers of persons having the attribute in samples of sizes \(n_1\) and \(n_2\) drawn from the two populations. Both counts are binomial, with parameters \((n_1, P_1)\) and \((n_2, P_2)\), where \(P_1\) and \(P_2\) are the population proportions. With sample proportions \(p_1 = x_1/n_1\) and \(p_2 = x_2/n_2\), the single-proportion result gives
\[ E(p_1) = P_1, \quad \operatorname{Var}(p_1) = \frac{P_1 Q_1}{n_1}; \qquad E(p_2) = P_2, \quad \operatorname{Var}(p_2) = \frac{P_2 Q_2}{n_2}. \]Hypotheses. \(H_0: P_1 = P_2\) — there is no significant difference between the two sample proportions, or the population proportions are equal. Against it \(H_1: P_1 \ne P_2\), or \(P_1 > P_2\), or \(P_1 < P_2\).
Statistic. Take the standard normal variate for \(p_1 - p_2\). As with two means, the variances add because the samples are independent:
\[ E(p_1 - p_2) = E(p_1) - E(p_2) = P_1 - P_2, \] \[ \operatorname{Var}(p_1 - p_2) = \operatorname{Var}(p_1) + \operatorname{Var}(p_2) = \frac{P_1 Q_1}{n_1} + \frac{P_2 Q_2}{n_2}, \qquad \text{S.E.}(p_1 - p_2) = \sqrt{\frac{P_1 Q_1}{n_1} + \frac{P_2 Q_2}{n_2}}, \] \[ z = \frac{p_1 - p_2 - (P_1 - P_2)} {\sqrt{\dfrac{P_1 Q_1}{n_1} + \dfrac{P_2 Q_2}{n_2}}} \sim N(0,1). \]Under \(H_0\) the two proportions are one and the same, say \(P_1 = P_2 = P\) and therefore \(Q_1 = Q_2 = Q\). The numerator's \((P - P)\) vanishes and the common \(PQ\) factors out of the denominator:
\[ z = \frac{p_1 - p_2 - (P - P)}{\sqrt{\dfrac{PQ}{n_1} + \dfrac{PQ}{n_2}}} = \frac{p_1 - p_2}{\sqrt{PQ\left(\dfrac{1}{n_1} + \dfrac{1}{n_2}\right)}} \sim N(0,1). \]When \(P\) is not known. If nothing is known about the population proportion, estimate it by pooling both samples — every observed success over every observed trial:
\[ \hat P = \frac{n_1 p_1 + n_2 p_2}{n_1 + n_2} = \frac{x_1 + x_2}{n_1 + n_2}, \qquad \hat Q = 1 - \hat P, \] \[ z = \frac{p_1 - p_2}{\sqrt{\hat P \hat Q\left(\dfrac{1}{n_1} + \dfrac{1}{n_2}\right)}} \sim N(0,1). \]The two forms of \(\hat P\) are the same number written two ways, and the problems below use whichever is more convenient from the data given.
Conclusion. If \(|z| > z_\alpha\) we may reject \(H_0\) at level \(\alpha\); otherwise, if \(|z| \le z_\alpha\), we may accept \(H_0\).
Printing note. The source writes \(p_1 = x_1/n_2\) where it means \(x_1/n_1\), and carries a stray subscript in \(PQ_1/n_1\) where the pooled form needs \(PQ/n_1\). Both are corrected above; the surrounding derivation is unaffected.
Town A: 60 of 200 favoured proposal. Town B: 90 of 250. Test \(H_0: p_1 = p_2\).
\(\hat p_1 = 0.30,\; \hat p_2 = 0.36;\; \hat p = 150/450 = 1/3\). SE = \(\sqrt{(1/3)(2/3)(1/200 + 1/250)} = 0.0447\).
\(Z = -0.06/0.0447 = -1.34 < 1.96\) ⇒ do not reject; no significant difference.
Defective rates of two factories: F1: 50 defective in 500; F2: 30 in 400. Test \(p_1 = p_2\).
\(\hat p_1 = 0.10, \hat p_2 = 0.075, \hat p = 80/900 = 0.0889\). SE = \(\sqrt{0.0889 \cdot 0.9111(1/500 + 1/400)} = 0.0190\).
\(Z = 0.025/0.019 = 1.32\) ⇒ not significant at 5 %.
For large \(n\), the sampling distribution of sample SD \(s\) is approximately \(N(\sigma,\; \sigma^2/2n)\).
Sections 6.1 and 6.2 test a standard deviation. The same data can be asked about in terms of variance instead, and for large \(n\) that has its own statistic.
Let \(x_1, \ldots, x_n\) be a random sample from a normal population with mean \(\mu\) and variance \(\sigma^{2}\). Test \(H_0: \sigma^{2} = \sigma_0^{2}\) — the sample variance is drawn from the population variance — against \(H_1: \sigma^{2} \ne \sigma_0^{2}\).
If \(s^{2}\) is the sample variance, then from the chi-square distribution
\[ \frac{n s^{2}}{\sigma^{2}} \sim \chi^{2}_{(n-1)} . \]For large \(n\) the chi-square distribution itself tends to normal, and its mean and variance are known outright:
\[ E\!\left(\frac{n s^{2}}{\sigma^{2}}\right) = n - 1 \simeq n, \qquad \operatorname{Var}\!\left(\frac{n s^{2}}{\sigma^{2}}\right) = 2(n-1) \simeq 2n . \]Standardising \(ns^{2}/\sigma^{2}\) with those two,
\[ z = \frac{\dfrac{n s^{2}}{\sigma^{2}} - n}{\sqrt{2n}} \sim N(0,1), \qquad\text{and under } H_0,\qquad z = \frac{\dfrac{n s^{2}}{\sigma_0^{2}} - n}{\sqrt{2n}} \sim N(0,1). \]If \(|z| > z_\alpha\) we reject \(H_0\) at level \(\alpha\); otherwise, if \(|z| \le z_\alpha\), we accept it. For a small sample the normal approximation is unavailable and the chi-square statistic is used directly, which is Unit 4's business.
A single standard deviation. For large \(n\), the sample standard deviation \(s\) itself follows a normal distribution with mean \(\sigma\) and variance \(\sigma^{2}/2n\):
\[ s \sim N\!\left(\sigma, \frac{\sigma^{2}}{2n}\right) \;\Longrightarrow\; E(s) = \sigma, \qquad \operatorname{Var}(s) = \frac{\sigma^{2}}{2n}, \qquad \text{S.E.}(s) = \frac{\sigma}{\sqrt{2n}} . \]Hence, testing \(H_0: \sigma = \sigma_0\) against \(H_1: \sigma \ne \sigma_0\),
\[ z = \frac{s - E(s)}{\text{S.E.}(s)} = \frac{s - \sigma}{\sigma/\sqrt{2n}} \sim N(0,1), \qquad\text{under } H_0,\qquad z = \frac{s - \sigma_0}{\sigma_0/\sqrt{2n}} \sim N(0,1). \]Two standard deviations. Both sample standard deviations are normal for large \(n_1, n_2\), so take the standard normal variate for \(s_1 - s_2\); once again the variances add, because \(s_1\) and \(s_2\) are independent:
\[ s_1 \sim N\!\left(\sigma_1, \frac{\sigma_1^{2}}{2n_1}\right), \qquad s_2 \sim N\!\left(\sigma_2, \frac{\sigma_2^{2}}{2n_2}\right), \] \[ E(s_1 - s_2) = \sigma_1 - \sigma_2, \qquad \operatorname{Var}(s_1 - s_2) = \frac{\sigma_1^{2}}{2n_1} + \frac{\sigma_2^{2}}{2n_2}, \qquad \text{S.E.}(s_1 - s_2) = \sqrt{\frac{\sigma_1^{2}}{2n_1} + \frac{\sigma_2^{2}}{2n_2}}, \] \[ z = \frac{s_1 - s_2 - (\sigma_1 - \sigma_2)} {\sqrt{\dfrac{\sigma_1^{2}}{2n_1} + \dfrac{\sigma_2^{2}}{2n_2}}} \sim N(0,1), \qquad\text{under } H_0: \sigma_1 = \sigma_2,\qquad z = \frac{s_1 - s_2} {\sqrt{\dfrac{\sigma_1^{2}}{2n_1} + \dfrac{\sigma_2^{2}}{2n_2}}} \sim N(0,1). \]The same three cases arise as for two means. Case 1: a common known \(\sigma\), which comes outside as \(\sigma\sqrt{1/(2n_1) + 1/(2n_2)}\). Case 2: neither known, so estimate each from its own sample, \(\hat\sigma_1 = s_1\) and \(\hat\sigma_2 = s_2\), giving \(\sqrt{s_1^{2}/(2n_1) + s_2^{2}/(2n_2)}\). Case 3: a common but unknown \(\sigma\), pooled exactly as before,
\[ \hat\sigma^{2} = \frac{n_1 s_1^{2} + n_2 s_2^{2}}{n_1 + n_2}, \qquad z = \frac{s_1 - s_2}{\hat\sigma\sqrt{\dfrac{1}{2n_1} + \dfrac{1}{2n_2}}} \sim N(0,1). \]From a sample of 200, \(s = 14.5\); test \(H_0: \sigma = 15\).
\(Z = (14.5 - 15)/(15/\sqrt{400}) = -0.5/0.75 = -0.67\) ⇒ not significant at 5 %.
\(n_1 = 100,\; s_1 = 8;\; n_2 = 150,\; s_2 = 6\). \(Z = (8-6)/\sqrt{64/200 + 36/300} = 2/\sqrt{0.44} = 3.02\). Reject \(H_0\).
To test \(H_0: \rho = 0\) (no correlation in the population), use:
(Strictly a t-statistic with \(n-2\) df, but for large \(n\) treated as Z.)
Let \((x_1,y_1), (x_2,y_2), \ldots, (x_n,y_n)\) be a bivariate random sample of size \(n\) from a normal population. Write \(r\) for the sample correlation coefficient and \(\rho\) for the population one. Test \(H_0: \rho = \rho_0\) against \(H_1: \rho \ne \rho_0\). Which statistic to use depends on how large \(\rho\) is, and the split is sharp.
Case 1 — \(\rho\) small, tending to zero. For large \(n\) the sampling distribution of \(r\) is normal with mean \(\rho\) and variance \((1-\rho^{2})^{2}/n\), so
\[ z = \frac{r - E(r)}{\text{S.E.}(r)} = \frac{r - \rho}{\sqrt{\dfrac{(1-\rho^{2})^{2}}{n}}} = \frac{r - \rho}{(1 - \rho^{2})/\sqrt n} \sim N(0,1), \]and under \(H_0\),
\[ z = \frac{r - \rho_0}{(1 - \rho_0^{2})/\sqrt n} \sim N(0,1). \]At \(\rho_0 = 0\) this collapses to the memorable \(z = r\sqrt n\). That is a shade more conservative than the \(t\)-form \(r\sqrt{n-2}/\sqrt{1-r^{2}}\) given above, which uses the sample \(r\) in the denominator rather than the hypothesised \(\rho_0\); for large \(n\) the two agree closely, and either is acceptable. Use whichever form the question's own notation points at.
Case 2 — \(\rho\) tending to 1 (in practice \(|\rho| > 0.7\)). The variance \((1-\rho^{2})^{2}/n\) collapses towards zero and the normal approximation fails. Fisher's transformation, below, is what rescues it.
Test statistic: \(Z = (Z' - \zeta_0)\sqrt{n - 3}\) where \(\zeta_0 = \frac{1}{2}\ln((1+\rho_0)/(1-\rho_0))\).
\(r = 0.4,\; n = 100\). Test \(H_0: \rho = 0\). \(Z = 0.4\sqrt{98}/\sqrt{0.84} = 4.32\) ⇒ highly significant.
\(r = 0.7,\; n = 28\). Test \(H_0: \rho = 0.5\) using Fisher's Z.
\(Z' = 0.5\ln(1.7/0.3) = 0.867\). \(\zeta_0 = 0.5\ln(1.5/0.5) = 0.549\). \(Z = (0.867 - 0.549)\sqrt{25} = 1.59\) ⇒ not significant at 5 %.
Let \((x_1,y_1), \ldots, (x_{n_1}, y_{n_1})\) be a bivariate random sample of size \(n_1\) from one normal population, and \((x_1,y_1), \ldots, (x_{n_2}, y_{n_2})\) another of size \(n_2\) from a second. Write \(r_1, r_2\) for the sample correlation coefficients and \(\rho_1, \rho_2\) for the population ones.
Hypotheses. \(H_0: \rho_1 = \rho_2\) — the two population correlation coefficients are equal, or there is no significant difference between the sample correlation coefficients. Against it \(H_1: \rho_1 \ne \rho_2\).
Statistic. Correlation coefficients cannot be compared directly, because the variance of \(r\) depends on \(\rho\) itself. Fisher's transformation removes that dependence: apply it to each sample,
\[ z_1 = \frac{1}{2}\log_e\frac{1+r_1}{1-r_1}, \qquad z_2 = \frac{1}{2}\log_e\frac{1+r_2}{1-r_2}, \]whose means and variances are
\[ E(z_1) = \frac{1}{2}\log_e\frac{1+\rho_1}{1-\rho_1}, \qquad E(z_2) = \frac{1}{2}\log_e\frac{1+\rho_2}{1-\rho_2}, \] \[ \operatorname{Var}(z_1) = \frac{1}{n_1 - 3}, \qquad \operatorname{Var}(z_2) = \frac{1}{n_2 - 3}. \]The variances no longer involve \(\rho\), which is the whole point of the transformation. Now take the standard normal variate for the difference \(z_1 - z_2\); the variances add, because the two samples are independent:
\[ E(z_1 - z_2) = \frac{1}{2}\log_e\frac{1+\rho_1}{1-\rho_1} - \frac{1}{2}\log_e\frac{1+\rho_2}{1-\rho_2}, \] \[ \operatorname{Var}(z_1 - z_2) = \frac{1}{n_1 - 3} + \frac{1}{n_2 - 3}, \qquad \text{S.E.}(z_1 - z_2) = \sqrt{\frac{1}{n_1 - 3} + \frac{1}{n_2 - 3}}, \] \[ V = \frac{z_1 - z_2 - E(z_1 - z_2)}{\text{S.E.}(z_1 - z_2)} \sim N(0,1). \]Under \(H_0: \rho_1 = \rho_2\) the two transformed means are equal, so \(E(z_1 - z_2) = 0\) and the numerator reduces to the observed difference alone:
\[ V = \frac{\dfrac{1}{2}\log_e\dfrac{1+r_1}{1-r_1} - \dfrac{1}{2}\log_e\dfrac{1+r_2}{1-r_2}} {\sqrt{\dfrac{1}{n_1 - 3} + \dfrac{1}{n_2 - 3}}} \sim N(0,1). \]Inference. If \(|V| > V_\alpha\) we reject \(H_0\) at level \(\alpha\); otherwise, if \(|V| \le V_\alpha\), we may accept \(H_0\). The critical values \(V_\alpha\) are the ordinary standard normal ones, because \(V\) is a standard normal variate.
Thirty-two problems in the order the textbook sets them, grouped by the test they use: five on a single proportion, seven on two proportions, six on a single mean, six on two means, four on variance and standard deviation, and four on correlation coefficients, all worked in full. Every one runs the same five steps from §1: null hypothesis, alternative hypothesis, level of significance, test statistic, conclusion. Read as a run rather than picked over, they show that nine different-looking formulae are one formula with nine different standard errors.
Source note. Every answer below was recomputed independently before it was written down. Where the book's printed figure differs, the reason is given in place: two are genuine misprints and are corrected, six are rounding artefacts where the printed value follows exactly from the book's own rounded intermediates, and those are left standing with a note.
A dice is thrown 900 times and a face of 3 or 5 is observed 335 times. Test whether the dice is unbiased.
Given \(n = 900\), \(x = 335\). For a fair dice,
\[ P = P(\text{3 or 5}) = \frac{2}{6} = \frac13, \qquad Q = 1 - \frac13 = \frac23 . \](1) Null hypothesis. \(H_0: P = \tfrac13\), i.e. the dice is unbiased.
(2) Alternative hypothesis. \(H_1: P \ne \tfrac13\), i.e. the dice is not unbiased. Two-tailed test.
(3) Test statistic under \(H_0\). The sample proportion is
\[ p = \frac{x}{n} = \frac{335}{900} = 0.3722, \] \[ z = \frac{p - P_0}{\sqrt{\dfrac{P_0 Q_0}{n}}} = \frac{\dfrac{335}{900} - \dfrac13}{\sqrt{\dfrac{\frac13 \cdot \frac23}{900}}} = \frac{0.03889}{0.01571} = 2.475 . \](4) Conclusion. \(|z| = 2.475\). The tabulated value at the 5% level for a two-tailed test is \(z_\alpha = 1.96\). Since \(|z| > z_\alpha\), \(H_0\) is rejected: the dice is not unbiased.
Printing note. The source prints \(z = 2.326\) here. That is not what the data give: \(35/900 = 0.03889\) divided by \(\sqrt{(2/9)/900} = 0.01571\) is \(2.475\), and no rounding of the intermediates reaches 2.326. The figure 2.326 is the 1% one-tailed critical value, so it appears to have been copied into the statistic's slot by mistake. The conclusion is unaffected, since both exceed 1.96.
A coin was thrown 400 times and head resulted 240 times. Test whether the coin is unbiased at the 1% level of significance.
Given \(n = 400\), \(x = 240\), and for a fair coin \(P = \tfrac12\), \(Q = \tfrac12\), so \(p = 240/400 = 0.6\).
(1) \(H_0: P = \tfrac12\), the coin is unbiased. (2) \(H_1: P \ne \tfrac12\), the coin is biased. Two-tailed.
(3)
\[ z = \frac{p - P_0}{\sqrt{\dfrac{P_0 Q_0}{n}}} = \frac{\dfrac{240}{400} - \dfrac12}{\sqrt{\dfrac{\frac12 \cdot \frac12}{400}}} = \frac{0.1}{0.025} = 4 . \](4) \(|z| = 4\), and \(z_\alpha = 2.58\) at the 1% level for a two-tailed test. \(|z| > z_\alpha\), so \(H_0\) is rejected. Indeed \(|z| > 3\), so \(H_0\) is rejected outright whatever the level. The coin is not unbiased.
In a sample of 1000 people in Maharashtra, 540 are rice eaters and the rest are wheat eaters. Can we assume that both rice and wheat are equally popular in Maharashtra at the 1% level of significance?
Given \(n = 1000\), \(x = 540\), so \(p = 540/1000 = 0.54\). Equal popularity means \(P = \tfrac12 = 0.5\) and \(Q = 1 - P = 0.5\).
(1) \(H_0: P = 0.5\), both are equally popular. (2) \(H_1: P \ne 0.5\). Two-tailed.
(3)
\[ z = \frac{0.54 - 0.5}{\sqrt{\dfrac{(0.5)(0.5)}{1000}}} = \frac{0.04}{0.01581} = 2.53 . \](4) \(|z| = 2.53\) and \(z_\alpha = 2.58\) at the 1% level, two-tailed. \(|z| < z_\alpha\), so \(H_0\) is accepted: both rice and wheat are equally popular in Maharashtra. Note how narrowly — at the 5% level, where \(z_\alpha = 1.96\), the same data would have rejected \(H_0\).
Twenty people were attacked by a disease and only 18 survived. Will you reject the hypothesis that the survival rate, if attacked by this disease, is 85% in favour of the hypothesis that it is more, at the 5% level?
Given \(n = 20\), \(x = 18\), so \(p = 18/20 = 0.9\); and \(P = 85\% = 0.85\), \(Q = 0.15\).
(1) \(H_0: P = 0.85\), the survival rate is 85%. (2) \(H_1: P > 0.85\), the survival rate is more than 85%. The wording “in favour of the hypothesis that it is more” is what makes this a one-tailed test.
(3)
\[ z = \frac{0.9 - 0.85}{\sqrt{\dfrac{(0.85)(0.15)}{20}}} = \frac{0.05}{0.0798} = 0.633 . \](4) \(|z| = 0.633\) and \(z_\alpha = 1.645\) at the 5% level for a one-tailed test. \(|z| < z_\alpha\), so \(H_0\) is accepted: the survival rate is 85%, and the evidence does not show it to be more.
Sample-size note. Here \(nQ = 20 \times 0.15 = 3\), below the usual guide that \(nP\) and \(nQ\) should both be at least 5, so the normal approximation is rough. The exact binomial test agrees: \(P(X \ge 18) = 0.2293 + 0.1368 + 0.0388 = 0.405\) when \(P = 0.85\), far above 0.05, so \(H_0\) is not rejected either way.
Rounding note. Carried to full precision the statistic is \(0.6262\); the printed \(0.633\) follows from rounding the standard error to \(0.079\) before dividing. Both accept \(H_0\) comfortably. Figure note. The source labels this one-tailed test's critical values \(z_{\alpha/2} = \pm 1.645\); for a one-tailed test they are \(z_\alpha\), and there is only one of them.
A manufacturer claims that 2% of the product is defective. In one day's production of 200 items, only 8 are defective. Test his claim that production of defective items is 2% or more at the 5% level. Find 95% confidence limits for the proportion of defective items.
Given \(n = 200\), \(x = 8\), so \(p = 8/200 = 0.04\); and \(P = 2\% = 0.02\), \(Q = 1 - P = 0.98\).
(1) \(H_0: P = 0.02\), the proportion of defective items is 2%. (2) \(H_1: P > 0.02\), the proportion is more than 2%. One-tailed.
(3)
\[ z = \frac{0.04 - 0.02}{\sqrt{\dfrac{(0.02)(0.98)}{200}}} = \frac{0.02}{0.0099} = 2.02 . \](4) \(|z| = 2.02\) and \(z_\alpha = 1.645\) at the 5% level, one-tailed. \(|z| > z_\alpha\), so \(H_0\) is rejected: the manufacturer's claim is not correct, and the proportion of defective items is more than 2%.
Confidence limits. The hypothesised \(P\) has been rejected, so the limits are built from the estimated proportion instead: \(\hat P = p = 0.04\), \(\hat Q = 1 - p = 0.96\).
\[ \left(p - 1.96\sqrt{\frac{\hat P \hat Q}{n}},\;\; p + 1.96\sqrt{\frac{\hat P \hat Q}{n}}\right) = \left(0.04 - 1.96\sqrt{\frac{(0.04)(0.96)}{200}},\;\; 0.04 + 1.96\sqrt{\frac{(0.04)(0.96)}{200}}\right) \] \[ = (0.0128,\; 0.0672). \]So the 95% confidence limits for the proportion of defective items are \((1.28,\; 6.72)\) per cent. Note that this interval contains 2%, although the test rejected \(P = 0.02\). The two are answering different questions: the test is one-tailed and measures the gap in units of the standard error at the claimed \(P = 0.02\), while the two-sided interval is built on the larger standard error at \(p = 0.04\). A one-tailed rejection and a two-sided interval need not agree.
A survey was conducted on T.B. patients in India. The data revealed that 1% of the population are suffering from T.B. in the country. A sample was collected in two colleges. In college A there are 5 T.B. patients out of 400 students, and in college B there are 10 T.B. patients out of 1200 students. Test the significance of the difference between the proportion of T.B. patients in the two colleges.
Given \(n_1 = 400\), \(n_2 = 1200\), \(x_1 = 5\), \(x_2 = 10\):
\[ p_1 = \frac{5}{400} = 0.0125, \qquad p_2 = \frac{10}{1200} = 0.0083, \]and here \(P\) is known from the national survey: \(P = 1\% = 0.01\), \(Q = 1 - P = 0.99\).
(1) \(H_0: P_1 = P_2 = P = 0.01\), no significant difference between the two colleges. (2) \(H_1: P_1 \ne P_2\). Two-tailed.
(3) Because \(P\) is known, the pooled estimate is not needed:
\[ z = \frac{p_1 - p_2}{\sqrt{PQ\left(\dfrac{1}{n_1} + \dfrac{1}{n_2}\right)}} = \frac{0.0125 - 0.0083} {\sqrt{(0.01)(0.99)\left(\dfrac{1}{400} + \dfrac{1}{1200}\right)}} = 0.7311 . \](4) \(|z| = 0.7311\) and \(z_\alpha = 1.96\) at the 5% level, two-tailed. \(|z| < z_\alpha\), so \(H_0\) is accepted: there is no significant difference between the proportion of T.B. patients in the two colleges.
Rounding note. With \(p_2\) kept exact at \(1/120 = 0.008333\) the statistic is \(0.7253\); the printed \(0.7311\) is exactly what \(p_2 = 0.0083\) gives, so it is the book's own rounding carried forward, not an error.
Random samples of 400 men and 600 women were asked whether they would like to have a flyover near their residence. 200 men and 325 women were in favour of the proposal. Test the hypothesis that the proportions of men and women in favour of the proposal are the same or not, at the 5% level of significance.
Given \(n_1 = 400\), \(n_2 = 600\), \(x_1 = 200\), \(x_2 = 325\):
\[ p_1 = \frac{200}{400} = 0.5, \qquad p_2 = \frac{325}{600} = 0.5417 . \]\(P\) is not known, so pool the two samples:
\[ \hat P = \frac{x_1 + x_2}{n_1 + n_2} = \frac{200 + 325}{400 + 600} = \frac{525}{1000} = 0.525, \qquad \hat Q = 1 - 0.525 = 0.475 . \](1) \(H_0: P_1 = P_2 = P\), no significant difference between the opinion of men and women about the flyover. (2) \(H_1: P_1 \ne P_2\). Two-tailed.
(3)
\[ z = \frac{p_1 - p_2}{\sqrt{\hat P \hat Q\left(\dfrac{1}{n_1} + \dfrac{1}{n_2}\right)}} = \frac{0.5 - 0.5417} {\sqrt{(0.525)(0.475)\left(\dfrac{1}{400} + \dfrac{1}{600}\right)}} = -1.294 . \](4) \(|z| = 1.294\) and \(z_\alpha = 1.96\) at the 5% level, two-tailed. \(|z| < z_\alpha\), so \(H_0\) is accepted: the proportions of men and women in favour of the flyover are the same.
Printing note. The source prints \(z = -1.269\). The data give \(-1.294\): the numerator is \(-0.0417\) and the standard error \(0.03223\), and no rounding of either reaches 1.269, so this looks like transposed digits. The conclusion is unaffected.
A survey was conducted on the eyesight of the students of Andhra Pradesh. Two random samples of sizes 900 and 1600 students were selected from the cities Visakhapatnam and Tirupathi, out of these 20% and 15.5% have the eyesight defect respectively. Test the significance that the eyesight defect in Visakhapatnam is more than that of Tirupathi at the 1% level of significance.
Given \(n_1 = 900\), \(n_2 = 1600\), \(p_1 = 20\% = 0.2\), \(p_2 = 15.5\% = 0.155\). \(P\) is not known, so pool — and here the counts are not given directly, so use the weighted form:
\[ \hat P = \frac{n_1 p_1 + n_2 p_2}{n_1 + n_2} = \frac{900(0.2) + 1600(0.155)}{900 + 1600} = 0.1712, \qquad \hat Q = 1 - 0.1712 = 0.8288 . \](1) \(H_0: P_1 = P_2 = P\), no significant difference between the eyesight of students in Visakhapatnam and in Tirupathi. (2) \(H_1: P_1 > P_2\), the eyesight defect in Visakhapatnam is more. One-tailed.
(3)
\[ z = \frac{0.2 - 0.155} {\sqrt{(0.1712)(0.8288)\left(\dfrac{1}{900} + \dfrac{1}{1600}\right)}} = 2.87 . \](4) \(|z| = 2.87\) and \(z_\alpha = 2.33\) at the 1% level, one-tailed. \(|z| > z_\alpha\), so \(H_0\) is rejected: the eyesight defect among students in Visakhapatnam is more than among students in Tirupathi.
Printing note. The source's problem statement gives the second sample as 1200 students with 15%, but every line of its working — the pooled \(\hat P = 0.1712\), the term \(1/1600\) inside the standard error, and the answer 2.87 — uses 1600 and 15.5%. The statement has been brought into line with the working above, since that is the answer the printed solution reaches. Worked instead from the printed statement, \(n_2 = 1200\) and \(p_2 = 0.15\) give \(\hat P = 0.1786\) and \(z = 3.01\), which rejects \(H_0\) more emphatically but by the same reasoning.
Before an increase in excise duty on tea, 800 persons out of a sample of 1000 persons were found to be tea drinkers. After an increase in excise duty, 800 people were tea drinkers in a sample of 1200 people. Test whether there is a significant decrease in the consumption of tea after the increase in excise duty.
Given \(n_1 = 1000\), \(n_2 = 1200\), \(x_1 = 800\), \(x_2 = 800\):
\[ p_1 = \frac{800}{1000} = 0.8, \qquad p_2 = \frac{800}{1200} = 0.67 . \]\(P\) is not known:
\[ \hat P = \frac{800 + 800}{1000 + 1200} = \frac{1600}{2200} = 0.7273, \qquad \hat Q = 1 - 0.7273 = 0.2727 . \](1) \(H_0: P_1 = P_2 = P\), no significant difference between the consumption of tea before and after the increase in excise duty. (2) \(H_1: P_1 > P_2\), there is a significant decrease after the increase. One-tailed.
(3)
\[ z = \frac{0.80 - 0.67} {\sqrt{(0.7273)(0.2727)\left(\dfrac{1}{1000} + \dfrac{1}{1200}\right)}} = 6.842 . \](4) \(|z| = 6.842\), and since \(|z| > 3\) we reject \(H_0\) at any level of significance without consulting a table: there is a significant decrease in the consumption of tea after the increase in excise duty.
Rounding note. Keeping \(p_2 = 2/3\) exact gives \(6.817\); the printed \(6.842\) follows from the rounded \(p_2 = 0.67\) and a standard error rounded to \(0.019\). Both are far beyond 3.
In a random sample of 500 men from a particular district of U.P., 300 are found to be smokers. In one of 1,000 men from another district, 550 are smokers. Do the data indicate that the two districts are significantly different with respect to the prevalence of smoking among men?
Given \(n_1 = 500\), \(n_2 = 1000\), \(x_1 = 300\), \(x_2 = 550\):
\[ p_1 = \frac{300}{500} = 0.6, \qquad p_2 = \frac{550}{1000} = 0.55, \] \[ \hat P = \frac{300 + 550}{500 + 1000} = \frac{850}{1500} = 0.57, \qquad \hat Q = 1 - 0.57 = 0.43 . \](1) \(H_0: P_1 = P_2 = P\), no significant difference between the smokers of the two districts. (2) \(H_1: P_1 \ne P_2\). Two-tailed.
(3)
\[ z = \frac{0.6 - 0.55} {\sqrt{(0.57)(0.43)\left(\dfrac{1}{500} + \dfrac{1}{1000}\right)}} = 1.84 . \](4) \(|z| = 1.84\) and \(z_\alpha = 1.96\) at the 5% level, two-tailed. \(|z| < z_\alpha\), so \(H_0\) is accepted: there is no significant difference between the smokers of the two districts in U.P.
1000 apples kept under one type of storage were found to show rotting to the extent of 4%. 1500 apples kept under another kind of storage showed 3% rotting. Can it be reasonably concluded that the second type of storage is superior to the first?
Given \(n_1 = 1000\), \(n_2 = 1500\), \(p_1 = 4\% = 0.04\), \(p_2 = 3\% = 0.03\). \(P\) is not known:
\[ \hat P = \frac{1000(0.04) + 1500(0.03)}{1000 + 1500} = 0.034, \qquad \hat Q = 1 - 0.034 = 0.966 . \](1) \(H_0: P_1 = P_2 = P\), no significant difference between the two storages. (2) \(H_1: P_1 > P_2\), the second type of storage is superior to the first — superior meaning less rotting in the second, so the first storage's proportion is the larger. One-tailed, right tail, which is why the statistic below is \(p_1 - p_2\).
(3)
\[ z = \frac{0.04 - 0.03} {\sqrt{(0.034)(0.966)\left(\dfrac{1}{1000} + \dfrac{1}{1500}\right)}} = 1.35 . \](4) \(|z| = 1.35\) and \(z_\alpha = 1.645\) at the 5% level, one-tailed. \(|z| < z_\alpha\), so \(H_0\) is accepted: the second type of storage is not superior to the first, and both storages are equal in their preference for keeping apples.
A machine produces 16 defective bolts in a batch of 500 bolts. After the machine is overhauled, it produces 3 defective bolts in a batch of 100 bolts. Has the machine improved?
Given \(n_1 = 500\), \(n_2 = 100\), \(x_1 = 16\), \(x_2 = 3\):
\[ p_1 = \frac{16}{500} = 0.032, \qquad p_2 = \frac{3}{100} = 0.03, \] \[ \hat P = \frac{16 + 3}{500 + 100} = \frac{19}{600} = 0.0317, \qquad \hat Q = 1 - 0.0317 = 0.9683 . \](1) \(H_0: P_1 = P_2 = P\), the machine has not improved — the condition is the same. (2) \(H_1: P_1 > P_2\), the machine has improved, i.e. it now produces a smaller proportion of defectives. One-tailed.
(3)
\[ z = \frac{0.032 - 0.03} {\sqrt{(0.0317)(0.9683)\left(\dfrac{1}{500} + \dfrac{1}{100}\right)}} = 0.1042 . \](4) \(|z| = 0.1042\) and \(z_\alpha = 1.645\) at the 5% level, one-tailed. \(|z| < z_\alpha\), so \(H_0\) is accepted: the machine has not improved. A drop from 3.2% to 3.0% on 100 bolts is well inside what chance alone produces.
A random sample of 1600 students has a mean score 99. Test whether the sample has been drawn from a population with mean score 100 and S.D. 15.
Given \(n = 1600\), \(\bar x = 99\), \(\mu = 100\), \(\sigma = 15\). Here \(\sigma\) is given, so no estimate is needed.
(1) \(H_0: \mu = 100\), the sample has been drawn from that population. (2) \(H_1: \mu \ne 100\). Two-tailed.
(3)
\[ z = \frac{\bar x - \mu_0}{\sigma/\sqrt n} = \frac{99 - 100}{15/\sqrt{1600}} = \frac{-1}{0.375} = -2.67 . \](4) \(|z| = 2.67\) and \(z_\alpha = 1.96\) at the 5% level, two-tailed. \(|z| > z_\alpha\), so \(H_0\) is rejected: the sample has not been drawn from that population.
A sample of 100 students is taken from a large population. The mean height of the students in the sample is 160 cms. Can it be reasonably regarded that, in the population, the mean height is 165 cm and the S.D. is 10 cm?
Given \(n = 100\), \(\bar x = 160\), \(\mu = 165\), \(\sigma = 10\).
(1) \(H_0: \mu = 165\), the sample has been drawn from the population with mean height 165 cms. (2) \(H_1: \mu \ne 165\). Two-tailed.
(3)
\[ z = \frac{160 - 165}{\dfrac{10}{\sqrt{100}}} = \frac{-5}{1} = -5 . \](4) \(|z| = 5\). Since \(|z| > 3\) we always reject \(H_0\): the given sample has not been drawn from the population with mean height 165 cm.
A random sample of 100 items, drawn from a universe with mean value 64 and S.D. 3, has a mean value 63.5. Is the difference in the means significant? What will be your inference if the sample had 200 items?
Given \(n = 100\), \(\bar x = 63.5\), and the universe's \(\mu = 64\) and \(\sigma = 3\). Here \(\sigma\) is given, so no estimate is needed.
(1) \(H_0: \mu = 64\), the difference in the means is not significant. (2) \(H_1: \mu \ne 64\). Two-tailed.
(3)
\[ z = \frac{\bar x - \mu_0}{\sigma/\sqrt n} = \frac{63.5 - 64}{3/\sqrt{100}} = \frac{-0.5}{0.3} = -1.67 . \](4) \(|z| = 1.67\) and \(z_\alpha = 1.96\) at the 5% level, two-tailed. \(|z| < z_\alpha\), so \(H_0\) is accepted: the difference between the means is not significant.
If instead \(n = 200\), with everything else unchanged,
\[ z = \frac{63.5 - 64}{3/\sqrt{200}} = -2.36, \]and now \(|z| = 2.36 > z_\alpha = 1.96\), so \(H_0\) is rejected. In this case the difference between the means is significant. The same half-unit gap in the means becomes significant purely by doubling the sample: the standard error shrinks like \(1/\sqrt n\), which is worth pausing on, because it is what makes “significant” and “large” different questions.
A sample of 400 individuals is found to have a mean height of 67.47 inches, S.D. 1.3 inches. Can it be reasonably regarded as a sample from a large population with mean height more than 67.39 inches at the 1% level of significance?
Given \(n = 400\), \(\bar x = 67.47\), \(s = 1.3\), \(\mu = 67.39\); \(\sigma\) is not known.
(1) \(H_0: \mu = 67.39\), the sample has been drawn from the population with mean height 67.39 inches. (2) \(H_1: \mu > 67.39\), the mean height of the population is more than 67.39 inches. One-tailed.
(3)
\[ z = \frac{67.47 - 67.39}{1.3/\sqrt{400}} = \frac{0.08}{0.065} = 1.23 . \](4) \(|z| = 1.23\) and \(z_\alpha = 2.33\) at the 1% level, one-tailed. \(|z| < z_\alpha\), so \(H_0\) is accepted: the sample has been drawn from the population, and the average height is not more than 67.39 inches.
A random sample of 400 students has a mean 4.5 cms and S.D. 1.5 cms. Can we assume that the sample is drawn from the population with mean 4.3 cms? Also find 95% confidence limits and 99% confidence limits for the population mean if the population S.D. is 2.
Given \(n = 400\), \(\bar x = 4.5\), \(s = 1.5\), \(\mu = 4.3\). For the test, \(\sigma\) is not known, so \(\hat\sigma = s = 1.5\).
(1) \(H_0: \mu = 4.3\), the sample has been drawn from the population. (2) \(H_1: \mu \ne 4.3\). Two-tailed.
(3)
\[ z = \frac{4.5 - 4.3}{1.5/\sqrt{400}} = \frac{0.2}{0.075} = 2.67 . \](4) \(|z| = 2.67 > z_\alpha = 1.96\) at the 5% level, two-tailed, so \(H_0\) is rejected: the sample has not been drawn from that population.
Confidence limits. For these the population S.D. is given, \(\sigma = 2\), so use it rather than \(s\).
\[ \left(\bar x - 1.96\frac{\sigma}{\sqrt n},\; \bar x + 1.96\frac{\sigma}{\sqrt n}\right) = \left(4.5 - 1.96\frac{2}{\sqrt{400}},\; 4.5 + 1.96\frac{2}{\sqrt{400}}\right) = (4.304,\; 4.696), \] \[ \left(\bar x - 2.58\frac{\sigma}{\sqrt n},\; \bar x + 2.58\frac{\sigma}{\sqrt n}\right) = \left(4.5 - 2.58\frac{2}{\sqrt{400}},\; 4.5 + 2.58\frac{2}{\sqrt{400}}\right) = (4.242,\; 4.758). \]The 95% interval just excludes 4.3 and the 99% interval (4.242, 4.758) contains it; and the 99% interval is the wider of the two, as more confidence always costs precision. Neither is a re-run of the test above: the test used \(s = 1.5\), the intervals use the stated \(\sigma = 2\).
The mean and S.D. of 60 students was found to be 145 and 40. Find 95% and 98% confidence limits for the population mean.
Given \(n = 60\), \(\bar x = 145\), \(s = 40\). \(\sigma\) is not known and is estimated by \(s\), which for a large sample is free: \(\hat\sigma = s = 40\).
\[ \left(145 - 1.96\frac{40}{\sqrt{60}},\; 145 + 1.96\frac{40}{\sqrt{60}}\right) = (134.88,\; 155.12), \] \[ \left(145 - 2.33\frac{40}{\sqrt{60}},\; 145 + 2.33\frac{40}{\sqrt{60}}\right) = (132.97,\; 157.03). \]So the 95% limits are \((134.88, 155.12)\) and the 98% limits \((132.97, 157.03)\). Note the multiplier for 98% two-tailed is 2.33, the same number that serves as the 1% one-tailed critical value — because 1% in one tail and 2% split between two tails put the same 1% in the upper tail.
In a random sample of size 500, the mean is found to be 30. In another random sample of size 400, the mean is 35. Could the samples have been drawn from the same population with S.D. 8?
Given \(n_1 = 500\), \(n_2 = 400\), \(\bar x = 30\), \(\bar y = 35\), and \(\sigma_1 = \sigma_2 = \sigma = 8\) — Case 1, a common known \(\sigma\).
(1) \(H_0: \mu_1 = \mu_2\), the two samples have been drawn from the same population. (2) \(H_1: \mu_1 \ne \mu_2\). Two-tailed.
(3)
\[ z = \frac{\bar x - \bar y}{\sigma\sqrt{\dfrac{1}{n_1} + \dfrac{1}{n_2}}} = \frac{30 - 35}{8\sqrt{\dfrac{1}{500} + \dfrac{1}{400}}} = \frac{-5}{0.5367} = -9.32 . \](4) \(|z| = 9.32 > 3\), so we always reject \(H_0\): the two samples were not drawn from the same population.
Figure note. The source's diagram for this problem is labelled \(|z| = 1.32\); its own working and conclusion both give 9.32, so the label is a slip.
The means of two large samples of 1000 and 2000 members are 67.5″ and 68.0″ respectively. Can the samples be regarded as drawn from the same population with standard deviation 2.5″?
Given \(n_1 = 1000\), \(n_2 = 2000\), \(\bar x = 67.5\), \(\bar y = 68\), \(\sigma_1 = \sigma_2 = \sigma = 2.5\) — Case 1 again.
(1) \(H_0: \mu_1 = \mu_2\). (2) \(H_1: \mu_1 \ne \mu_2\). Two-tailed.
(3)
\[ z = \frac{67.5 - 68}{2.5\sqrt{\dfrac{1}{1000} + \dfrac{1}{2000}}} = \frac{-0.5}{0.09682} = -5.16 . \](4) \(|z| = 5.16 > 3\), so \(H_0\) is always rejected: the two samples have not come from the same population. Half an inch is a small difference in heights, but on samples this size it is far beyond chance — again the \(1/\sqrt n\) effect.
A random sample of heights of 6400 men of the state Andhra Pradesh has a mean 170 cm and S.D. 6.4 cm, while a random sample of heights of 1600 men of the state Punjab has a mean 172 cm and S.D. 6.3 cm. Do the data conclude that Punjab men are on an average taller than the Andhra Pradesh men?
Given \(n_1 = 6400\), \(n_2 = 1600\), \(\bar x = 170\), \(\bar y = 172\), \(s_1 = 6.4\), \(s_2 = 6.3\). The \(\sigma\)'s are not known, so estimate each from its own sample — Case 2: \(\hat\sigma_1 = s_1 = 6.4\), \(\hat\sigma_2 = s_2 = 6.3\).
(1) \(H_0: \mu_1 = \mu_2\), the average heights are equal. (2) \(H_1: \mu_1 < \mu_2\), Punjab men are taller. One-tailed.
(3)
\[ z = \frac{\bar x - \bar y}{\sqrt{\dfrac{s_1^{2}}{n_1} + \dfrac{s_2^{2}}{n_2}}} = \frac{170 - 172}{\sqrt{\dfrac{(6.4)^{2}}{6400} + \dfrac{(6.3)^{2}}{1600}}} = -11.32 . \](4) \(|z| = 11.32 > 3\), so \(H_0\) is always rejected: Punjab men are on an average taller than Andhra Pradesh men.
The mean and S.D. from a sample of size 400 are 250 and 40 respectively. Those of another sample of size 400 are 220 and 55. Test at the 1% level of significance whether the means of the two populations from which the samples have been drawn are equal.
Given \(n_1 = n_2 = 400\), \(\bar x = 250\), \(\bar y = 220\), \(s_1 = 40\), \(s_2 = 55\); the \(\sigma\)'s are not known, so Case 2 applies.
(1) \(H_0: \mu_1 = \mu_2\), the two samples have been drawn from the same population. (2) \(H_1: \mu_1 \ne \mu_2\). Two-tailed.
(3)
\[ z = \frac{250 - 220}{\sqrt{\dfrac{(40)^{2}}{400} + \dfrac{(55)^{2}}{400}}} = \frac{30}{3.4004} = 8.82 . \](4) \(|z| = 8.82 > 3\), so \(H_0\) is always rejected: the two samples were not drawn from the same population.
The data of life time of two types of electrical bulbs is given below.
| Type of bulb | Size | Average life time (hrs.) | S.D. (hrs.) |
|---|---|---|---|
| A | 50 | 1980 | 80 |
| B | 70 | 2010 | 60 |
According to this data, can we assume that the life time of type A is superior to type B at the 1% level?
Given \(n_1 = 50\), \(n_2 = 70\), \(\bar x = 1980\), \(\bar y = 2010\), \(s_1 = 80\), \(s_2 = 60\); the \(\sigma\)'s are not known, so Case 2.
(1) \(H_0: \mu_1 = \mu_2\), the life times of the two types are equal. (2) \(H_1: \mu_1 > \mu_2\), type A is superior to type B. One-tailed.
(3)
\[ z = \frac{1980 - 2010}{\sqrt{\dfrac{(80)^{2}}{50} + \dfrac{(60)^{2}}{70}}} = \frac{-30}{13.395} = -2.24 . \](4) \(|z| = 2.24\) and \(z_\alpha = 2.33\) at the 1% level, one-tailed. \(|z| < z_\alpha\), so \(H_0\) is accepted: the life times of bulbs of types A and B are equal. In fact no level could reject here: \(H_1\) is right-tailed and \(z\) is negative, so the statistic points away from the alternative. Type A's mean is the lower of the two, and comparing \(|z|\) with a one-tailed value is only valid when \(z\) has the sign \(H_1\) predicts.
Intelligence tests were given to two groups of boys and girls of the same age group chosen from the same college and the following results were obtained.
| Size | Mean | S.D. | |
|---|---|---|---|
| Boys | 100 | 73 | 10 |
| Girls | 60 | 75 | 8 |
Examine whether the difference between the means is significant or not.
Given \(n_1 = 100\), \(n_2 = 60\), \(\bar x = 73\), \(\bar y = 75\), \(s_1 = 10\), \(s_2 = 8\); Case 2 again.
(1) \(H_0: \mu_1 = \mu_2\), no significant difference between the sample means. (2) \(H_1: \mu_1 \ne \mu_2\). Two-tailed.
(3)
\[ z = \frac{73 - 75}{\sqrt{\dfrac{(10)^{2}}{100} + \dfrac{(8)^{2}}{60}}} = \frac{-2}{1.4376} = -1.39 . \](4) \(|z| = 1.39\) and \(z_\alpha = 1.96\) at the 5% level, two-tailed. \(|z| < z_\alpha\), so \(H_0\) is accepted: there is no significant difference between the sample means.
A random sample of size 72 has variance 144 drawn from a population. Can we assume that the sample has been drawn from the population with variance 100 at the 1% level of significance?
Given \(n = 72\), \(s^{2} = 144\), \(\sigma^{2} = 100\).
(1) \(H_0: \sigma^{2} = 100\), the sample has been drawn from the population. (2) \(H_1: \sigma^{2} \ne 100\). Two-tailed.
(3)
\[ z = \frac{\dfrac{n s^{2}}{\sigma_0^{2}} - n}{\sqrt{2n}} = \frac{\dfrac{72 \times 144}{100} - 72}{\sqrt{2 \times 72}} = \frac{103.68 - 72}{12} = \frac{31.68}{12} = 2.64 . \](4) \(|z| = 2.64\) and \(z_\alpha = 2.58\) at the 1% level, two-tailed. \(|z| > z_\alpha\), so \(H_0\) is rejected: the sample has not been drawn from the population. It is a near thing — 2.64 against 2.58 — and a slightly smaller sample would have gone the other way.
A random sample of size 50 has S.D. 11.8 drawn from a normal population. Can we assume that the sample has been drawn from the population with S.D. 10?
Given \(n = 50\), \(s = 11.8\), \(\sigma = 10\).
(1) \(H_0: \sigma = 10\), the sample has been drawn from the population. (2) \(H_1: \sigma \ne 10\). Two-tailed.
(3)
\[ z = \frac{s - \sigma_0}{\sigma_0/\sqrt{2n}} = \frac{11.8 - 10}{10/\sqrt{2 \times 50}} = \frac{1.8}{1} = 1.8 . \](4) \(|z| = 1.8\) and \(z_\alpha = 1.96\) at the 5% level, two-tailed. \(|z| < z_\alpha\), so \(H_0\) is accepted: the sample has been drawn from the population.
Worth comparing with Problem 25. There a variance 44% above the hypothesised value was rejected on a sample of 72; here a standard deviation 18% above it is accepted on a sample of 50. The two statistics are different large-sample approximations, and on the same data they need not agree: Problem 26's figures put into the variance form of Problem 25 give \(z = (50 \times 139.24/100 - 50)/10 = 1.96\), right on the boundary, where the S.D. form gave 1.8. With a sample this small the exact \(\chi^2\) test of Unit 4 is the one to trust.
Two random samples of sizes 100 each have been drawn from two populations with the standard deviations 2.823 and 1.548. Test the significance difference between the sample standard deviations if the population standard deviation is 2.
This is Case 1 of §6: \(n_1 = n_2 = 100\), \(s_1 = 2.823\), \(s_2 = 1.548\), and a common known \(\sigma_1 = \sigma_2 = 2\), so
\[ z = \frac{s_1 - s_2}{\sigma\sqrt{\dfrac{1}{2n_1} + \dfrac{1}{2n_2}}} . \](1) \(H_0: \sigma_1 = \sigma_2\), the sample standard deviations do not differ significantly. (2) \(H_1: \sigma_1 \ne \sigma_2\). Two-tailed.
(3)
\[ z = \frac{2.823 - 1.548}{2\sqrt{\dfrac{1}{200} + \dfrac{1}{200}}} = \frac{1.275}{2 \times 0.1} = 6.38 . \](4) \(|z| = 6.38 > 1.96\) (and beyond 3), so \(H_0\) is rejected: the difference between the sample standard deviations is significant.
Two large samples of heights, of sizes 1000 and 1200, have means 67.42″ and 67.25″ and standard deviations 2.58″ and 2.50″. Test (i) whether the means differ significantly, and (ii) whether the standard deviations differ significantly.
Statement note. The wording of this question is reconstructed from the data its working uses; the figures themselves are the book's.
Given \(n_1 = 1000\), \(n_2 = 1200\), \(\bar x = 67.42\), \(\bar y = 67.25\), \(s_1 = 2.58\), \(s_2 = 2.50\).
(i) Test for means.
(1) \(H_0: \mu_1 = \mu_2\), no significant difference between the sample means. (2) \(H_1: \mu_1 \ne \mu_2\). Two-tailed.
(3) The \(\sigma\)'s are not known, so Case 2:
\[ z = \frac{\bar x - \bar y}{\sqrt{\dfrac{s_1^{2}}{n_1} + \dfrac{s_2^{2}}{n_2}}} = \frac{67.42 - 67.25}{\sqrt{\dfrac{(2.58)^{2}}{1000} + \dfrac{(2.5)^{2}}{1200}}} = 1.56 . \](4) \(|z| = 1.56 < z_\alpha = 1.96\) at the 5% level, two-tailed, so \(H_0\) is accepted: there is no significant difference between the sample means.
(ii) Test for standard deviations. The same two samples, now asked about their spread rather than their centre.
(1) \(H_0: \sigma_1 = \sigma_2\), no significant difference between the sample standard deviations. (2) \(H_1: \sigma_1 \ne \sigma_2\). Two-tailed.
(3) Case 2 again, with the \(2n\) in the denominators that marks a spread test:
\[ z = \frac{s_1 - s_2}{\sqrt{\dfrac{s_1^{2}}{2n_1} + \dfrac{s_2^{2}}{2n_2}}} = \frac{2.58 - 2.50}{\sqrt{\dfrac{(2.58)^{2}}{2 \times 1000} + \dfrac{(2.50)^{2}}{2 \times 1200}}} = 1.03 . \](4) \(|z| = 1.03 < z_\alpha = 1.96\) at the 5% level, two-tailed, so \(H_0\) is accepted: there is no significant difference between the sample standard deviations either. The two populations look alike in both respects.
Rounding note. Carried to full precision part (ii) gives \(1.039\); the printed \(1.03\) follows from rounding the standard error before dividing.
A bivariate random sample of size 1600 has the correlation coefficient 0.5. Test whether this sample has been drawn from the population with correlation coefficient 0.6.
Given \(n = 1600\), \(r = 0.5\), \(\rho = 0.6\).
(1) \(H_0: \rho = 0.6\), the sample has been drawn from the population. (2) \(H_1: \rho \ne 0.6\). Two-tailed.
(3) \(\rho\) is small enough for Case 1:
\[ z = \frac{r - \rho_0}{(1 - \rho_0^{2})/\sqrt n} = \frac{0.5 - 0.6}{\big(1 - (0.6)^{2}\big)/\sqrt{1600}} = \frac{-0.1}{0.016} = -6.25 . \](4) \(|z| = 6.25 > 3\), so \(H_0\) is always rejected: the sample has not been drawn from the population.
A bivariate random sample of heights and weights of 50 students has the correlation coefficient 0.2. Test the significance of the correlation between heights and weights of the students.
Given \(n = 50\), \(r = 0.2\). “Test the significance of the correlation” with no population value named means testing against zero.
(1) \(H_0: \rho = 0\), there is no correlation between the heights and weights of the students. (2) \(H_1: \rho \ne 0\). Two-tailed.
(3) With \(\rho_0 = 0\) the Case 1 statistic reduces to \(r\sqrt n\):
\[ z = \frac{r - \rho_0}{(1 - \rho_0^{2})/\sqrt n} = \frac{0.2 - 0}{(1 - 0)/\sqrt{50}} = 0.2\sqrt{50} = 1.41 . \](4) \(|z| = 1.41\) and \(z_\alpha = 1.96\) at the 5% level, two-tailed. \(|z| < z_\alpha\), so \(H_0\) is accepted: there is no correlation between the heights and weights of the students. A sample \(r\) of 0.2 on fifty students is not enough to establish a relationship.
A bivariate random sample of size 100 has the correlation coefficient 0.92. Can we assume that the sample be regarded as from the population with correlation coefficient 0.85?
Given \(n = 100\), \(r = 0.92\), \(\rho = 0.85\).
(1) \(H_0: \rho = 0.85\), the given sample has been drawn from the population. (2) \(H_1: \rho \ne 0.85\). Two-tailed.
(3) Here \(\rho\) tends to 1, well past the 0.7 mark, so Case 1 is unavailable and Fisher's transformation is used:
\[ V = \frac{\dfrac{1}{2}\log_e\dfrac{1+r}{1-r} - \dfrac{1}{2}\log_e\dfrac{1+\rho_0}{1-\rho_0}} {\sqrt{\dfrac{1}{n-3}}} = \frac{\dfrac{1}{2}\log_e\dfrac{1.92}{0.08} - \dfrac{1}{2}\log_e\dfrac{1.85}{0.15}} {\sqrt{\dfrac{1}{97}}} = \frac{1.59 - 1.26}{0.102} = 3.24 . \](4) \(|V| = 3.24 > 3\), so \(H_0\) is always rejected: the sample cannot be regarded as drawn from the population.
Rounding note. The two transformed values are \(1.5890\) and \(1.2562\) and the standard error \(0.10153\), giving \(V = 3.278\) exactly. The printed \(3.24\) is what the rounded \(1.59\), \(1.26\) and \(0.102\) give, so it is the book's own rounding carried through, not a misprint. Both are past 3, so the conclusion is the same either way.
Two bivariate random samples of sizes 420 and 225 have the correlation coefficients 0.72 and 0.65 respectively. Test whether there is any significance difference between the sample correlation coefficients at the 1% level.
Given \(n_1 = 420\), \(n_2 = 225\), \(r_1 = 0.72\), \(r_2 = 0.65\).
(1) \(H_0: \rho_1 = \rho_2\), no significant difference between the two sample correlation coefficients. (2) \(H_1: \rho_1 \ne \rho_2\). Two-tailed.
(3) Transform each, then compare:
\[ V = \frac{\dfrac{1}{2}\log_e\dfrac{1+r_1}{1-r_1} - \dfrac{1}{2}\log_e\dfrac{1+r_2}{1-r_2}} {\sqrt{\dfrac{1}{n_1 - 3} + \dfrac{1}{n_2 - 3}}} = \frac{\dfrac{1}{2}\log_e\dfrac{1.72}{0.28} - \dfrac{1}{2}\log_e\dfrac{1.65}{0.35}} {\sqrt{\dfrac{1}{420 - 3} + \dfrac{1}{225 - 3}}} = \frac{0.908 - 0.775}{0.0831} = 1.601 . \](4) \(|V| = 1.601\) and \(V_\alpha = 2.58\) at the 1% level, two-tailed. \(|V| \le V_\alpha\), so \(H_0\) is accepted: there is no significant difference between the two sample correlation coefficients.
Rounding note. Exactly, the transforms are \(0.9076\) and \(0.7753\) and \(V = 1.593\); the printed \(1.601\) follows from the rounded \(0.908\) and \(0.775\).
Try it. Samples of sizes 64 and 39 have correlation coefficients 0.58 and 0.76. Test whether they differ significantly. (Answer: the transforms are \(0.662\) and \(0.996\), the standard error is \(\sqrt{1/61 + 1/36} = 0.210\), and \(V = 1.59 < 1.96\): not significant.)
| Hypothesis | Test Statistic |
|---|---|
| \(\mu = \mu_0\) | \((\bar X - \mu_0)/(\sigma/\sqrt n)\) |
| \(\mu_1 = \mu_2\) | \((\bar X_1 - \bar X_2)/\sqrt{\sigma_1^2/n_1 + \sigma_2^2/n_2}\) |
| \(p = p_0\) | \((\hat p - p_0)/\sqrt{p_0(1-p_0)/n}\) |
| \(p_1 = p_2\) | \((\hat p_1 - \hat p_2)/\sqrt{\hat p \hat q (1/n_1 + 1/n_2)}\) |
| \(\sigma = \sigma_0\) | \((s - \sigma_0)/(\sigma_0/\sqrt{2n})\) |
| \(\rho = 0\) | \(r\sqrt{n-2}/\sqrt{1 - r^2}\) |
| \(\sigma^2 = \sigma_0^2\) | \(\left(ns^2/\sigma_0^2 - n\right)/\sqrt{2n}\) |
| \(\sigma_1 = \sigma_2\) | \((s_1 - s_2)/\sqrt{\sigma_1^2/(2n_1) + \sigma_2^2/(2n_2)}\) |
| \(\rho = \rho_0\) (small \(\rho\)) | \((r - \rho_0)/\big((1-\rho_0^2)/\sqrt n\big)\) |
| \(\rho_1 = \rho_2\) | \((z_1 - z_2)/\sqrt{1/(n_1-3) + 1/(n_2-3)}\), \(z_i\) Fisher-transformed |
Nine rows, one shape. Every statistic is \(\big(\text{estimate} - \text{hypothesised value}\big)/\text{standard error}\); only the standard error changes, and deriving it is the whole of the work.
Where this goes next. Every test above borrowed its normality from the Central Limit Theorem, and the CLT needed \(n > 30\). Below that the prop is gone twice over: the population must now be assumed normal, and estimating \(\sigma\) by \(s\) is no longer free — the extra uncertainty thickens the tails. That is the whole reason Unit 4 exists, and why \(Z\) gives way to Student's \(t\), with \(\chi^2\) and \(F\) for questions about variance rather than about means.