Skip to the content

Topics Covered

t-test (single mean) t-test (difference) Paired t-test t-test (correlation) χ² Goodness of Fit χ² Independence χ² for Variance F-test
On this page
  1. 1. Why Small Sample Tests?
  2. 2. t-test for Single Mean
  3. 3. t-test for Difference of Two Means (Independent Samples)
  4. 4. Paired t-test (Dependent Samples)
  5. t-test for a Correlation Coefficient
  6. 5. χ² Test for Goodness of Fit
  7. 6. χ² Test for Independence (Contingency Tables)
  8. 7. χ² Test for Single Variance
  9. 8. F-test for Equality of Two Variances
  10. 9. Choosing the Right Small-sample Test
  11. Key Take-aways
  12. Extra Practical Problems
  13. Worked Problems on Small Sample Tests

1. Why Small Sample Tests?

When \(n\) is small (\(n \le 30\)) and the population is normal but \(\sigma\) is unknown, the Z-test does not apply. Instead we use Student's t, χ² and F distributions developed precisely for small samples.

Assumption: the parent population is normal (or approximately so).

Student's t vs. standard normal -4 -3 -2 -1 0 1 2 3 4 Normal (df = ∞) t, df = 5 t, df = 1 (heavy tails) heavier tails
Fig 4.1 — Why a special distribution for small samples: Student's \(t\) has the same bell shape as the standard normal but a lower peak and heavier tails. With few degrees of freedom (\(df = 1\)) the extra tail probability is large, so critical values are bigger than \(1.96\); as \(df\) grows the \(t\) curve tightens onto the normal (\(df = \infty\)).

When is a Sample Small?

The working rule used in the chapter: a sample of size less than 30 is treated as small, and the tests built for such samples are the small sample tests. The page above says \(n \le 30\); the two statements differ only at the single value \(n = 30\), which is where the rule of thumb sits, not a sharp boundary.

What makes a test "small-sample" is not the size itself but what the test relies on. A large sample test leans on the central limit theorem: whatever the parent population, the statistic is approximately normal. A small sample cannot lean on that, so these tests use the exact sampling distributions of their statistics — \(\chi^2\), \(t\) and \(F\) (and \(z\) when the population standard deviation is known) — which are exact only when the parent population is normal.

Assumptions of Student's t-test

  1. The parent population from which the sample is drawn is normal.
  2. The sample observations are independent, that is, the sample is random.
  3. The population standard deviation \(\sigma\) is unknown. (If it is known, the \(z\) statistic of Unit 3 applies, even to a small sample — Worked Problem 1 below is exactly that case.)

Two Ways of Writing the Sample Variance

Textbooks write the \(t\) statistic in two forms that look different and are the same number. The chapter mostly uses the sample variance with divisor \(n\); the formula boxes of §§2–4 below use the one with divisor \(n-1\), and call it \(s\). From here on, \(s^2\) means divisor \(n\) and \(S^2\) divisor \(n-1\):

\[ s^2 = \frac{1}{n}\sum_{i=1}^{n}(x_i - \bar x)^2, \qquad S^2 = \frac{1}{n-1}\sum_{i=1}^{n}(x_i - \bar x)^2 . \]

Statement. \(\dfrac{s}{\sqrt{n-1}} = \dfrac{S}{\sqrt n}\), so \(\dfrac{\bar x - \mu_0}{s/\sqrt{n-1}}\) and \(\dfrac{\bar x - \mu_0}{S/\sqrt n}\) are the same statistic.

Proof. Both variances divide the same sum of squares, so

\[ n s^2 = \sum_{i=1}^{n}(x_i - \bar x)^2 = (n-1) S^2 . \]

Divide both sides by \(n(n-1)\):

\[ \frac{s^2}{n-1} = \frac{S^2}{n} . \]

Both sides are positive, so taking square roots keeps the equality: \(\dfrac{s}{\sqrt{n-1}} = \dfrac{S}{\sqrt n}\). □

The computing form of \(s^2\) used in the problems comes from expanding the square: \(\sum(x_i - \bar x)^2 = \sum x_i^2 - 2\bar x\sum x_i + n\bar x^2 = \sum x_i^2 - n\bar x^2\), because \(\sum x_i = n\bar x\). Dividing by \(n\),

\[ s^2 = \frac{1}{n}\sum x_i^2 - \bar x^2 . \]

2. t-test for Single Mean

\(H_0: \mu = \mu_0\). With sample mean \(\bar X\) and sample SD \(s\) (with \(n-1\) divisor):

\[ t \;=\; \dfrac{\bar X - \mu_0}{s/\sqrt n} \;\sim\; t_{n-1}. \]

Where the t Statistic Comes From

Let \(x_1, x_2, \ldots, x_n\) be a random sample from \(N(\mu, \sigma^2)\). Test \(H_0: \mu = \mu_0\) (the sample has been drawn from a population with mean \(\mu_0\)) against \(H_1: \mu \ne \mu_0\), \(\mu > \mu_0\) or \(\mu < \mu_0\).

Case 1: \(\sigma\) known. The sample mean of a normal sample is itself normal, \(\bar x \sim N(\mu, \sigma^2/n)\), whatever \(n\) is. Standardising it,

\[ z = \frac{\bar x - \mu_0}{\sigma/\sqrt n} \sim N(0, 1) \quad\text{under } H_0, \]

so the normal test applies to a small sample too.

Case 2: \(\sigma\) unknown. Now \(\sigma\) must be estimated, and the estimate brings its own randomness. Two facts about a normal sample are used.

  1. \(\bar x \sim N(\mu, \sigma^2/n)\), so \(Z = \dfrac{\bar x - \mu}{\sigma/\sqrt n} \sim N(0,1)\).
  2. \(\dfrac{n s^2}{\sigma^2} \sim \chi^2_{n-1}\), and it is independent of \(\bar x\).

By the definition of Student's \(t\): a standard normal variate divided by the square root of an independent \(\chi^2\) variate over its degrees of freedom has the \(t\) distribution with those degrees of freedom. So

\[ t = \frac{Z}{\sqrt{\dfrac{n s^2/\sigma^2}{n-1}}} = \frac{\dfrac{\bar x - \mu}{\sigma/\sqrt n}}{\dfrac{s}{\sigma}\sqrt{\dfrac{n}{n-1}}} \sim t_{n-1} . \]

The unknown \(\sigma\) appears once in the numerator's denominator and once in the denominator, so it cancels:

\[ t = \frac{\bar x - \mu}{\sigma/\sqrt n}\cdot\frac{\sigma}{s}\sqrt{\frac{n-1}{n}} = \frac{(\bar x - \mu)\sqrt{n-1}}{s} = \frac{\bar x - \mu}{s/\sqrt{n-1}} . \]

That cancellation is the whole point: the statistic can be computed without knowing \(\sigma\), and its distribution, \(t_{n-1}\), does not depend on \(\sigma\) either. Under \(H_0\),

\[ t = \frac{\bar x - \mu_0}{s/\sqrt{n-1}} = \frac{\bar x - \mu_0}{S/\sqrt n} \sim t_{n-1} , \]

the two forms being equal by §1. Inference: if \(|t|\) exceeds the table value \(t_{\alpha, n-1}\) for the test in hand (two-tailed or one-tailed), reject \(H_0\); otherwise accept it.

Confidence Limits for the Mean

Turning the test around gives an interval. Since \(P\!\left(-t_{\alpha/2} \le \dfrac{\bar x - \mu}{s/\sqrt{n-1}} \le t_{\alpha/2}\right) = 1 - \alpha\), rearranging the inequalities for \(\mu\) gives the \(100(1-\alpha)\%\) confidence limits

\[ \bar x \pm t_{\alpha/2,\,n-1}\,\frac{s}{\sqrt{n-1}} . \]

An interval has two ends, so it uses the two-tailed point: for 95% limits \(t_{0.025}\), for 99% limits \(t_{0.005}\). Worked Problem 5 below shows what goes wrong when a one-tailed table value is used instead.

EXAMPLE 1

A sample of 10 observations: \(\bar X = 48,\; s = 5\). Test \(H_0: \mu = 50\) at 5 %.

\(t = (48 - 50)/(5/\sqrt{10}) = -2/1.581 = -1.265\). df = 9. \(t_{0.025, 9} = 2.262\). \(|t| < 2.262\) ⇒ accept \(H_0\).

EXAMPLE 2

10 students score: 65, 70, 68, 72, 75, 73, 71, 69, 74, 70 — \(\bar X = 70.7,\; s = 2.98\). Test \(H_0: \mu = 68\).

\(t = (70.7 - 68)/(2.98/\sqrt{10}) = 2.7/0.943 = 2.86\). df = 9. \(t_{0.025, 9} = 2.262\) ⇒ reject \(H_0\). Mean exceeds 68.

3. t-test for Difference of Two Means (Independent Samples)

\(H_0: \mu_1 = \mu_2\). With pooled variance:

\[ s_p^2 \;=\; \dfrac{(n_1-1)s_1^2 + (n_2-1)s_2^2}{n_1 + n_2 - 2}, \] \[ t \;=\; \dfrac{\bar X_1 - \bar X_2}{s_p\sqrt{\dfrac{1}{n_1} + \dfrac{1}{n_2}}} \;\sim\; t_{n_1+n_2-2}. \]

Assumption: \(\sigma_1 = \sigma_2\) (equal variances). Use Welch's correction otherwise.

Where the Two-sample t Statistic Comes From

Let \(x_1, \ldots, x_{n_1}\) be a random sample from \(N(\mu_1, \sigma_1^2)\) and \(y_1, \ldots, y_{n_2}\) an independent one from \(N(\mu_2, \sigma_2^2)\). Test \(H_0: \mu_1 = \mu_2\).

Step 1. Each mean is normal: \(\bar x \sim N(\mu_1, \sigma_1^2/n_1)\) and \(\bar y \sim N(\mu_2, \sigma_2^2/n_2)\). The samples are independent, so the variances of the two means add, and the difference is normal:

\[ \bar x - \bar y \sim N\!\left(\mu_1 - \mu_2,\; \frac{\sigma_1^2}{n_1} + \frac{\sigma_2^2}{n_2}\right). \]

Case 3 (the chapter's order): \(\sigma_1^2\) and \(\sigma_2^2\) both known. Standardise directly:

\[ z = \frac{\bar x - \bar y}{\sqrt{\dfrac{\sigma_1^2}{n_1} + \dfrac{\sigma_2^2}{n_2}}} \sim N(0,1) \quad\text{under } H_0 . \]

Case 2: \(\sigma_1 = \sigma_2 = \sigma\), known. The denominator becomes \(\sqrt{\sigma^2/n_1 + \sigma^2/n_2} = \sigma\sqrt{1/n_1 + 1/n_2}\), so

\[ z = \frac{\bar x - \bar y}{\sigma\sqrt{\dfrac{1}{n_1} + \dfrac{1}{n_2}}} \sim N(0,1). \]

Case 1: \(\sigma_1 = \sigma_2 = \sigma\), unknown. This is the \(t\)-test. Start from the standard normal variate of Case 2, \(\xi = \dfrac{\bar x - \bar y - (\mu_1 - \mu_2)}{\sigma\sqrt{1/n_1 + 1/n_2}} \sim N(0,1)\). With \(s_1^2, s_2^2\) the sample variances (divisor \(n\)),

\[ \frac{n_1 s_1^2}{\sigma^2} \sim \chi^2_{n_1 - 1}, \qquad \frac{n_2 s_2^2}{\sigma^2} \sim \chi^2_{n_2 - 1}, \]

independently, and by the additive property of \(\chi^2\) their sum is \(\chi^2 = \dfrac{n_1 s_1^2 + n_2 s_2^2}{\sigma^2} \sim \chi^2_{n_1 + n_2 - 2}\). Divide \(\xi\) by \(\sqrt{\chi^2/(n_1+n_2-2)}\), as the definition of \(t\) requires; \(\sigma\) cancels exactly as in §2:

\[ t = \frac{\bar x - \bar y - (\mu_1 - \mu_2)}{S\sqrt{\dfrac{1}{n_1} + \dfrac{1}{n_2}}} \sim t_{n_1 + n_2 - 2}, \qquad S^2 = \frac{n_1 s_1^2 + n_2 s_2^2}{n_1 + n_2 - 2} . \]

Under \(H_0\) the term \(\mu_1 - \mu_2\) is zero. Since \(n_1 s_1^2 = (n_1 - 1)S_1^2\), this \(S^2\) is the pooled \(s_p^2\) written above with the other divisor — the same number.

Why the F-test comes first. Case 1 assumes the two population variances are equal. So before applying it, it is theoretically desirable to test \(\sigma_1^2 = \sigma_2^2\) with the F-test of §8; Worked Problem 34 below does exactly that.

EXAMPLE 1

Method A: \(n_1 = 8,\; \bar X_1 = 75,\; s_1 = 5\). Method B: \(n_2 = 10,\; \bar X_2 = 70,\; s_2 = 6\). Test at 5 %.

\(s_p^2 = [7(25) + 9(36)]/16 = (175 + 324)/16 = 31.19;\;s_p = 5.59\).

\(t = 5/(5.59 \sqrt{1/8 + 1/10}) = 5/(5.59 \cdot 0.474) = 1.89\). df = 16. \(t_{0.025, 16} = 2.12\). \(|t| < 2.12\) ⇒ no significant difference.

EXAMPLE 2

Drug A: \(n_1 = 12,\bar X_1 = 25,\; s_1 = 2\). Drug B: \(n_2 = 14,\bar X_2 = 28,\; s_2 = 2.5\). Test \(H_0: \mu_A = \mu_B\) at 5 %.

\(s_p^2 = (11 \cdot 4 + 13 \cdot 6.25)/24 = 5.22\); \(s_p = 2.28\).

\(t = -3/(2.28\sqrt{1/12 + 1/14}) = -3/0.898 = -3.34\). df = 24, \(t_{0.025} = 2.06\) ⇒ reject \(H_0\).

4. Paired t-test (Dependent Samples)

Used when the two samples are paired (before–after, matched pairs, twin studies).

Let \(d_i = X_i - Y_i\) (paired difference) with mean \(\bar d\) and SD \(s_d\). Then

\[ t \;=\; \dfrac{\bar d - 0}{s_d/\sqrt n} \;\sim\; t_{n-1}. \]

Where the Paired Statistic Comes From

In paired data \((x_1, y_1), \ldots, (x_n, y_n)\) the two values in a pair belong to the same sample unit — the same patient before and after, the same cow on two foods. They are not independent, so the two-sample test of §3, which assumed independent samples, does not apply.

The way out is to work with one number per unit, the difference \(d_i = x_i - y_i\). The \(d_i\) are a single random sample, and \(H_0: \mu_1 = \mu_2\) becomes \(H_0: \mu_d = 0\). The single-mean test of §2 applied to the \(d_i\) is then the paired test:

\[ \bar d = \frac{1}{n}\sum d_i, \qquad s^2 = \frac{1}{n}\sum (d_i - \bar d)^2 = \frac{1}{n}\sum d_i^2 - \bar d^2, \] \[ t = \frac{\bar d}{s/\sqrt{n-1}} = \frac{\bar d}{S/\sqrt n} \sim t_{n-1}, \qquad S^2 = \frac{1}{n-1}\left[\sum d_i^2 - n\bar d^2\right]. \]

Keep the direction straight. With \(d = x - y\), "the second reading is higher" means \(\mu_d < 0\), so an increase from \(x\) to \(y\) is a lower-tailed test: reject \(H_0\) when \(t < -t_\alpha\). Two problems in the source get this backwards (Worked Problems 12 and 14 below).

EXAMPLE 1 (Before–after)

Weights of 8 patients before and after a diet:

Before7278698085768274
After7076677883747972
d22222232

\(\bar d = 17/8 = 2.125;\; s_d = 0.354\). \(t = 2.125/(0.354/\sqrt 8) = 2.125/0.125 = 17.0\). df = 7, highly significant ⇒ reject \(H_0\): weights fell. (That the diet caused the fall needs a control group; a before–after design cannot show it.)

EXAMPLE 2

10 students' scores before and after a coaching: \(\bar d = 4.5,\; s_d = 5.2\). Test \(H_0: \mu_d = 0\).

\(t = 4.5/(5.2/\sqrt{10}) = 4.5/1.644 = 2.74\). df = 9. \(t_{0.025} = 2.26\) ⇒ reject \(H_0\); scores rose significantly after the coaching.

t-test for a Correlation Coefficient

Let \((x_1, y_1), \ldots, (x_n, y_n)\) be a random sample of size \(n\) from a bivariate normal population, with sample correlation coefficient \(r\) and population correlation \(\rho\). Test

\(H_0: \rho = 0\) (the variables are uncorrelated) against \(H_1: \rho \ne 0\).

Under \(H_0\),

\[ t = \frac{r}{\sqrt{\dfrac{1 - r^2}{n - 2}}} = \frac{r\sqrt{n-2}}{\sqrt{1 - r^2}} \sim t_{n-2} . \]

Two parameters, the two means, are estimated before \(r\) can be computed, which is where the \(n - 2\) degrees of freedom come from. If \(|t| > t_{\alpha, n-2}\), reject \(H_0\).

Only for \(\rho = 0\). The source states the statistic as \((r - \rho)/\sqrt{(1-r^2)/(n-2)} \sim t_{n-2}\) for general \(\rho\). That is true only when \(\rho = 0\): for any other value the distribution of \(r\) is skewed and the statistic is not \(t\). A hypothesis \(\rho = \rho_0 \ne 0\) is tested with Fisher's \(Z\)-transformation, in Unit 3.

Confidence limits. For a large sample the source gives \(r \pm z_{\alpha/2}\,\dfrac{1 - r^2}{\sqrt n}\), from the large-sample standard error of \(r\); Worked Problem 18 uses it and compares it with Fisher's interval.

5. χ² Test for Goodness of Fit

Tests whether observed frequencies fit a hypothesized distribution.

\[ \chi^2 \;=\; \sum_{i=1}^{k}\dfrac{(O_i - E_i)^2}{E_i} \;\sim\; \chi^2_{k - 1 - r}, \]

where \(k\) = number of categories, \(r\) = number of parameters estimated from data, \(O_i\) = observed, \(E_i\) = expected frequency. Reject \(H_0\) if \(\chi^2 > \chi^2_{\alpha, df}\).

Conditions: All \(E_i \ge 5\); if any are smaller, combine adjacent categories.

Conditions for the Validity of the χ² Test

For the \(\chi^2\) test of goodness of fit, and the test of independence of §6, to be valid:

  1. The sample observations should be independent.
  2. The constraints on the cell frequencies should be linear, for example \(\sum_i (A_i) = \sum_j (B_j) = N\) in a test of independence.
  3. The total frequency \(N\) should be reasonably large, say greater than 50.
  4. No theoretical (expected) cell frequency should be less than 5.

Pooling. If an expected frequency is less than 5, it is pooled with the preceding or succeeding frequency until the pooled frequency exceeds 5, and the degrees of freedom are reduced by the cells lost: if three frequencies are pooled into one, two degrees of freedom are subtracted. Estimated parameters. If parameters of the fitted distribution are estimated from the data, that number is subtracted as well. So with \(k\) classes, \(K\) degrees of freedom lost to pooling and \(l\) parameters estimated,

\[ \text{d.f.} = k - 1 - K - l . \]

Karl Pearson's Statistic and its Short Form

Karl Pearson's test asks whether the deviation of experiment from theory is just chance, or is due to the inadequacy of the theory to fit the observed data. With \(O_i\) the observed and \(E_i\) the expected frequencies, \(H_0\): the fitted distribution is a good fit,

\[ \chi^2 = \sum_{i=1}^{k}\frac{(O_i - E_i)^2}{E_i} \sim \chi^2_{k-1} . \]

The short form. Expand the square and use \(\sum O_i = \sum E_i = N\):

\[ \sum\frac{(O_i - E_i)^2}{E_i} = \sum\frac{O_i^2}{E_i} - 2\sum O_i + \sum E_i = \sum\frac{O_i^2}{E_i} - 2N + N = \sum\frac{O_i^2}{E_i} - N . \]

The short form needs only one division per class, which is why it is convenient by hand.

EXAMPLE 1 (Fair die)

Throw a die 60 times. Observed frequencies: 8, 11, 9, 12, 10, 10. Test fairness.

\(E_i = 60/6 = 10\) for each. \(\chi^2 = [(8-10)^2 + (11-10)^2 + (9-10)^2 + (12-10)^2 + 0 + 0]/10 = (4+1+1+4)/10 = 1.0\). df = 5. \(\chi^2_{0.05,5} = 11.07\) ⇒ do not reject \(H_0\); the data are consistent with a fair die.

EXAMPLE 2 (Mendelian ratio 9:3:3:1)

Of 800 plants: 460 round-yellow, 140 round-green, 130 wrinkled-yellow, 70 wrinkled-green. Expected: 450, 150, 150, 50.

\(\chi^2 = (10)^2/450 + (10)^2/150 + (20)^2/150 + (20)^2/50 = 0.222 + 0.667 + 2.667 + 8.0 = 11.56\). df = 3, \(\chi^2_{0.05,3} = 7.81\) ⇒ reject; data deviates from Mendelian ratio.

6. χ² Test for Independence (Contingency Tables)

Tests whether two categorical variables are independent.

For an \(r \times c\) contingency table with observed frequencies \(O_{ij}\):

\[ E_{ij} \;=\; \dfrac{(\text{row total}_i)(\text{col total}_j)}{N}, \] \[ \chi^2 \;=\; \sum_{i,j}\dfrac{(O_{ij} - E_{ij})^2}{E_{ij}} \;\sim\; \chi^2_{(r-1)(c-1)}. \]

Expected Frequencies Under Independence

Attribute A has \(r\) classes \(A_1, \ldots, A_r\) and B has \(s\) classes \(B_1, \ldots, B_s\). In the \(r \times s\) contingency table, \((A_iB_j)\) is the number of units with both \(A_i\) and \(B_j\), \((A_i)\) and \((B_j)\) are the margins, and \(\sum_i (A_i) = \sum_j (B_j) = N\).

Estimate the class probabilities by their proportions, \(P[A_i] = (A_i)/N\) and \(P[B_j] = (B_j)/N\). Under \(H_0\) the attributes are independent, so the probability of a cell is the product:

\[ P[A_iB_j] = P[A_i]\,P[B_j] = \frac{(A_i)}{N}\cdot\frac{(B_j)}{N} . \]

Multiplying by \(N\) gives the expected frequency,

\[ E_{ij} = N\,P[A_iB_j] = \frac{(A_i)(B_j)}{N}, \qquad \chi^2 = \sum_{i=1}^{r}\sum_{j=1}^{s}\frac{(O_{ij} - E_{ij})^2}{E_{ij}} \sim \chi^2_{(r-1)(s-1)} . \]

The degrees of freedom are \((r-1)(s-1)\): once the margins are fixed, only that many cells can be filled freely; the rest follow by subtraction.

The 2 × 2 Shortcut, Proved

Statement. For the table with cells \(a, b\) (first row) and \(c, d\) (second row), \(N = a + b + c + d\),

\[ \chi^2 = \frac{N(ad - bc)^2}{(a+b)(c+d)(a+c)(b+d)} . \]

Proof. The margins are \(a+b\), \(c+d\) (rows) and \(a+c\), \(b+d\) (columns), so the expected frequencies are

\[ E_{11} = \frac{(a+b)(a+c)}{N},\quad E_{12} = \frac{(a+b)(b+d)}{N},\quad E_{21} = \frac{(c+d)(a+c)}{N},\quad E_{22} = \frac{(c+d)(b+d)}{N} . \]

Step 1: every deviation has the same size. Put \(a\) over the common denominator \(N\) and expand:

\[ a - E_{11} = \frac{a(a+b+c+d) - (a+b)(a+c)}{N} = \frac{a^2 + ab + ac + ad - a^2 - ac - ab - bc}{N} = \frac{ad - bc}{N} . \]

The same expansion gives \(b - E_{12} = -\dfrac{ad - bc}{N}\), \(c - E_{21} = -\dfrac{ad - bc}{N}\) and \(d - E_{22} = \dfrac{ad - bc}{N}\). Squared, all four are \(\dfrac{(ad - bc)^2}{N^2}\).

Step 2: take the common factor out.

\[ \chi^2 = \frac{(ad - bc)^2}{N^2}\left[\frac{1}{E_{11}} + \frac{1}{E_{12}} + \frac{1}{E_{21}} + \frac{1}{E_{22}}\right]. \]

Each \(E\) has \(N\) in its denominator, so each \(1/E\) has \(N\) on top, and one \(N\) cancels against the \(N^2\):

\[ \chi^2 = \frac{(ad - bc)^2}{N}\left[\frac{1}{(a+b)(a+c)} + \frac{1}{(a+b)(b+d)} + \frac{1}{(c+d)(a+c)} + \frac{1}{(c+d)(b+d)}\right]. \]

Step 3: add the four fractions. Pair the first two, which share \((a+b)\): \(\dfrac{1}{a+b}\left[\dfrac{1}{a+c} + \dfrac{1}{b+d}\right] = \dfrac{1}{a+b}\cdot\dfrac{N}{(a+c)(b+d)}\), because \((b+d) + (a+c) = N\). The last two give \(\dfrac{1}{c+d}\cdot\dfrac{N}{(a+c)(b+d)}\) the same way. Adding,

\[ \frac{N}{(a+c)(b+d)}\left[\frac{1}{a+b} + \frac{1}{c+d}\right] = \frac{N}{(a+c)(b+d)}\cdot\frac{N}{(a+b)(c+d)} = \frac{N^2}{(a+b)(c+d)(a+c)(b+d)} . \]

Step 4. Substitute into Step 2: \(\chi^2 = \dfrac{(ad - bc)^2}{N}\cdot\dfrac{N^2}{(a+b)(c+d)(a+c)(b+d)} = \dfrac{N(ad - bc)^2}{(a+b)(c+d)(a+c)(b+d)}\). □

Source note. The chapter's proof reaches the last line through \(\dfrac{(ad + bc)^2}{N}\cdots\); the plus is a misprint for the minus it then uses. Steps 1 and 3 are the two pieces of algebra it leaves to the reader.

Yates' Correction

In a \(2 \times 2\) table, when a cell frequency is small (an expected frequency below 5), the continuity correction is

\[ \chi^2 = \frac{N\left(|ad - bc| - \tfrac{N}{2}\right)^2}{(a+b)(c+d)(a+c)(b+d)} . \]

The \(\chi^2\) distribution is continuous and the cell counts are whole numbers; subtracting \(N/2\) moves each deviation half a unit towards zero, which makes the approximation better for small counts. The source recommends using the correction even when no expected frequency is below 5 — a common textbook practice, and the reason Extra Practical Problem 7 below applies it.

EXAMPLE 1 (2 × 2 table)

Smoking vs cancer in 200 people:

CancerNoTotal
Smoker6040100
Non-smoker2080100
Total80120200

Expected: \(E_{11} = 100 \cdot 80/200 = 40\); rest by subtraction: \(E_{12} = 60,\; E_{21} = 40,\; E_{22} = 60\).

\(\chi^2 = (60-40)^2/40 + (40-60)^2/60 + (20-40)^2/40 + (80-60)^2/60 = 10 + 6.67 + 10 + 6.67 = 33.34\). df = 1, \(\chi^2_{0.05,1} = 3.84\) ⇒ reject; smoking and cancer are not independent.

EXAMPLE 2 (3 × 2 table)

Voting preference of 200 voters in three regions, between two candidates:

Candidate ACandidate BTotal
Region 1251540
Region 2273360
Region 36832100
Total12080200

Expected counts, row total × column total / 200: Region 1: \(40 \cdot 120/200 = 24\) and 16; Region 2: 36 and 24; Region 3: 60 and 40.

\(\chi^2 = \dfrac{(25-24)^2}{24} + \dfrac{(15-16)^2}{16} + \dfrac{(27-36)^2}{36} + \dfrac{(33-24)^2}{24} + \dfrac{(68-60)^2}{60} + \dfrac{(32-40)^2}{40}\) \(= 0.042 + 0.063 + 2.250 + 3.375 + 1.067 + 1.600 = 8.40\).

df = (3−1)(2−1) = 2, \(\chi^2_{0.05,2} = 5.99\). Since \(8.40 > 5.99\), reject (p ≈ 0.015): preference depends on region. Region 2, where B is ahead, contributes most of the statistic.

7. χ² Test for Single Variance

\(H_0: \sigma^2 = \sigma_0^2\). Test statistic:

\[ \chi^2 \;=\; \dfrac{(n - 1) s^2}{\sigma_0^2} \;\sim\; \chi^2_{n-1}. \]

Where the Single-variance Statistic Comes From

For a random sample from \(N(\mu, \sigma^2)\) with sample variance \(s^2\) (divisor \(n\)), \(\dfrac{n s^2}{\sigma^2} = \dfrac{\sum (x_i - \bar x)^2}{\sigma^2} \sim \chi^2_{n-1}\). Under \(H_0: \sigma^2 = \sigma_0^2\),

\[ \chi_0^2 = \frac{n s^2}{\sigma_0^2} = \frac{(n-1)S^2}{\sigma_0^2} \sim \chi^2_{n-1}, \]

the second form being the one above, since \(n s^2 = (n-1)S^2\) (§1). The source states the rule as "reject if \(\chi_0^2 > \chi^2_{\alpha, n-1}\)", which is the test against \(\sigma^2 > \sigma_0^2\). Against \(\sigma^2 < \sigma_0^2\) the rejection region is the lower tail, and against \(\sigma^2 \ne \sigma_0^2\) it is both tails, as Example 1 below uses.

EXAMPLE 1

Sample of 25 with \(s^2 = 18\). Test \(H_0: \sigma^2 = 16\) at 5 %.

\(\chi^2 = 24 \cdot 18/16 = 27\). df = 24. \(\chi^2_{0.025, 24} = 39.36;\; \chi^2_{0.975, 24} = 12.40\). Since \(12.4 < 27 < 39.4\) ⇒ accept \(H_0\).

EXAMPLE 2

\(n = 21,\; s^2 = 6\). Test \(H_0: \sigma^2 = 4\) (one-tailed: \(\sigma^2 > 4\)).

\(\chi^2 = 20 \cdot 6/4 = 30\). df = 20. \(\chi^2_{0.05, 20} = 31.41 > 30\) ⇒ accept \(H_0\) at 5 %.

8. F-test for Equality of Two Variances

\(H_0: \sigma_1^2 = \sigma_2^2\). Take the ratio of larger to smaller variance:

\[ F \;=\; \dfrac{s_1^2}{s_2^2} \;\sim\; F_{n_1 - 1,\; n_2 - 1}\;\;(s_1^2 \ge s_2^2). \]

Reject \(H_0\) if \(F > F_{\alpha, n_1-1, n_2-1}\) (one-tailed) or beyond two-tailed bounds.

Where the F Statistic Comes From

Let a sample of size \(n_1\) come from \(N(\mu_1, \sigma_1^2)\) and an independent one of size \(n_2\) from \(N(\mu_2, \sigma_2^2)\), with sample variances \(s_1^2, s_2^2\). Then

\[ \frac{n_1 s_1^2}{\sigma_1^2} \sim \chi^2_{n_1 - 1}, \qquad \frac{n_2 s_2^2}{\sigma_2^2} \sim \chi^2_{n_2 - 1}, \]

independently. By definition, the ratio of two independent \(\chi^2\) variates, each divided by its degrees of freedom, is an \(F\) variate:

\[ F = \frac{\dfrac{n_1 s_1^2}{\sigma_1^2}\Big/(n_1 - 1)}{\dfrac{n_2 s_2^2}{\sigma_2^2}\Big/(n_2 - 1)} \sim F_{(n_1 - 1,\; n_2 - 1)} . \]

Under \(H_0: \sigma_1^2 = \sigma_2^2\) the population variances cancel, leaving the ratio of the unbiased estimates \(S_1^2 = \dfrac{n_1 s_1^2}{n_1 - 1}\) and \(S_2^2 = \dfrac{n_2 s_2^2}{n_2 - 1}\):

\[ F = \frac{S_1^2}{S_2^2} \sim F_{(n_1 - 1,\; n_2 - 1)} . \]

Which Variance Goes on Top

F tables give only upper points, so the larger estimate is put in the numerator and the degrees of freedom follow it:

Reject \(H_0\) if \(F\) exceeds the table value. One caution the source passes over: putting the larger estimate on top makes a two-sided alternative \(\sigma_1^2 \ne \sigma_2^2\) a two-tailed test in disguise, so at an exact 5% level the upper 2.5% point is the right comparison. Textbooks, this one included, conventionally read the 5% table; the problems below say where the difference could matter.

An F-test also states its question in terms of the unbiased estimates \(S^2\). If a problem already gives unbiased estimates, they are used as they stand — converting them again by \(n/(n-1)\) is the slip in Worked Problem 33.

EXAMPLE 1

\(n_1 = 11,\; s_1^2 = 25;\;\; n_2 = 16,\; s_2^2 = 16\). Test at 5 %.

\(F = 25/16 = 1.5625\). df = (10, 15). \(F_{0.05, 10, 15} = 2.54\) ⇒ accept \(H_0\); variances equal.

EXAMPLE 2

Two production lines: \(s_1^2 = 36\) (n₁=21), \(s_2^2 = 16\) (n₂=16). \(F = 36/16 = 2.25\). df = (20, 15). \(F_{0.05} = 2.33\) ⇒ accept \(H_0\) at 5 %.

9. Choosing the Right Small-sample Test

QuestionTest
Single mean (σ known, any n)z, standard normal
Single mean (σ unknown)t with n−1 df
Difference of two means (independent)Pooled t with n₁+n₂−2 df
Difference of paired meansPaired t with n−1 df
Correlation coefficient (ρ = 0)t with n−2 df
Population varianceχ² with n−1 df
Equality of two variancesF with (n₁−1, n₂−1) df
Goodness of fit / independenceχ² with appropriate df

Key Take-aways

Extra Practical Problems

PRACTICE

Additional worked problems with step-by-step procedures to support self-study, matching this unit's topics.

STEP-BY-STEP PROCEDURE (general test of significance)
  1. State the null hypothesis \(H_0\) and alternative hypothesis \(H_1\) (one- or two-sided).
  2. Choose the level of significance \(\alpha\) (usually 5% or 1%).
  3. Compute the appropriate test statistic from the sample.
  4. Find the degrees of freedom and read the table (critical) value at \(\alpha\).
  5. Decide: if \(|\text{calculated}| > \text{tabulated}\), reject \(H_0\); otherwise accept \(H_0\). State the conclusion in terms of \(H_0\).

Test statistics used in this unit

Problem 1 — t-test for a Single Mean

DATA

A new greengram variety is expected to yield 12 q/ha. Tested on 10 fields: 14.3, 12.6, 13.7, 10.9, 13.7, 12.0, 11.4, 12.0, 12.6, 13.1.

\(H_0: \mu = 12\) vs \(H_1: \mu \ne 12\). Sample mean \(\bar x = 12.63\), \(s = \sqrt{\frac{\sum(x-\bar x)^2}{n-1}} = \sqrt{\frac{10.60}{9}} = 1.085\).

\(t = \dfrac{\bar x - \mu}{s/\sqrt n} = \dfrac{12.63 - 12}{1.085/\sqrt{10}} = 1.836\), df = 9. Table \(t_{0.05,9} = 2.262\). Since \(1.836 < 2.262\), accept \(H_0\) — the variety yields about 12 q/ha.

Problem 2 — t-test for Two Independent Sample Means

DATA

Yields (q) under two manures: Manure I (n=8): 14, 20, 34, 48, 32, 42, 30, 44; Manure II (n=7): 31, 18, 22, 28, 40, 26, 45.

\(\bar x = 33,\ \bar y = 30\); \(\sum(x-\bar x)^2 = 968,\ \sum(y-\bar y)^2 = 554\).

\(s^2 = \dfrac{968 + 554}{8+7-2} = 117.07\), \(s = 10.82\).

\(t = \dfrac{\bar x - \bar y}{s\sqrt{\frac1{n_1}+\frac1{n_2}}} = \dfrac{33-30}{10.82\sqrt{\frac18+\frac17}} = 0.54\), df = 13. Table \(t_{0.05,13} = 2.16\). Since \(0.54 < 2.16\), accept \(H_0\) — no significant difference between the manures.

Problem 3 — Paired t-test

DATA

Body-weight increase (oz) of animals given treatments A and B from six litters:

Litter123456
A283229362934
B252427303029

Differences \(d = A - B\): 3, 8, 2, 6, −1, 5; \(\bar d = 23/6 = 3.83\), \(s^2 = \frac{50.83}{5} = 10.17\), \(s = 3.19\).

\(t = \dfrac{\bar d}{s/\sqrt n} = \dfrac{3.83}{3.19/\sqrt 6} = 2.94\), df = 5. Table \(t_{0.05,5} = 2.571\). Since \(2.94 > 2.571\), reject \(H_0\) — treatments A and B differ significantly.

Problem 4 — F-test (Variance-Ratio Test)

DATA

Sample I (n=10): 20, 16, 26, 27, 23, 22, 18, 24, 25, 19. Sample II (n=12): 17, 23, 32, 25, 22, 24, 28, 18, 31, 33, 20, 27.

\(\bar x = 22,\ \bar y = 25\); \(s_1^2 = \frac{120}{9} = 13.33\), \(s_2^2 = \frac{314}{11} = 28.55\).

\(F = \dfrac{\text{larger } s^2}{\text{smaller } s^2} = \dfrac{28.55}{13.33} = 2.14\), df = (11, 9). Table \(F_{0.05}(11, 9) = 3.10\). Since \(2.14 < 3.10\), accept \(H_0\) — equal variances.

Problem 5 — Chi-square Goodness of Fit

DATA

200 random digits with observed frequencies of 0–9: 22, 21, 16, 20, 23, 15, 18, 21, 19, 25. Expected frequency \(= 200/10 = 20\) each.

\(\chi^2 = \sum\dfrac{(O_i - E_i)^2}{E_i} = \dfrac{86}{20} = 4.3\), df = 9. Table \(\chi^2_{0.05,9} = 16.91\). Since \(4.3 < 16.91\), accept \(H_0\) — digits are equally distributed.

Problem 6 — Chi-square for a 2×2 Table

DATA

Cholera epidemic data:

AttackedNot attackedTotal
Inoculated31469500
Not inoculated18513151500
Total21617842000

Expected: E(31)=54, E(469)=446, E(185)=162, E(1315)=1338. \(\chi^2 = \sum\dfrac{(O-E)^2}{E} = 14.64\), df = (2−1)(2−1) = 1. Table \(\chi^2_{0.05,1} = 3.841\). Since \(14.64 > 3.841\), reject \(H_0\) — inoculation is effective.

Problem 7 — Chi-square with Yates' Correction (small cell frequency)

DATA

50 small shops:

In TownsIn VillagesTotal
Run by men171835
Run by women31215
Total203050

The counts are small (one observed cell is 3; the smallest expected count is \(15 \times 20/50 = 6\)), and many texts apply Yates' correction to every 2×2 table, so it is applied here: \(\chi^2 = \dfrac{N\left(|ad-bc| - \frac N2\right)^2}{(a+b)(c+d)(a+c)(b+d)} = \dfrac{50\left(|17\cdot12 - 18\cdot3| - 25\right)^2}{35\cdot15\cdot20\cdot30} = 2.48\), df = 1. Since \(2.48 < 3.841\), accept \(H_0\) — no evidence of relatively more women owners in villages. (Without the correction \(\chi^2 = 3.57\), which also does not reach 3.841.)

Unsolved Exercises

PRACTICE
  1. Marks of 6 boys 63, 63, 64, 66, 60, 68; test whether mean = 66. (Ans: \(t_{cal} = -1.78\))
  2. Onion yield: Method I (n=12, mean 25.25, SS 186.25) vs Method II (n=12, mean 28.83, SS 737.67). (Ans: \(t_{cal} = -1.35\))
  3. Blood-pressure changes d = 5, 2, 8, −1, 3, 0, −2, 1, 5, 0, 4; paired test for an increase. (Ans: \(\bar d = 2.27,\ s_d = 3.04,\ t_{cal} = 2.48\), df = 10)
  4. Output A: 40,30,38,41,38,35; B: 39,38,41,23,32,39,40,34. Is B more stable (F-test)? (Ans: \(s_1^2 = 16,\ s_2^2 = 35.93\) (both with \(n-1\)); \(F = 35.93/16 = 2.25\), df = (7, 5) — A is the more stable line)
  5. Anesthetic mortality: Local 511 alive/24 dead, General 147/18. \(\chi^2\) test. (Ans: \(\chi^2_{cal} = 9.22\))
  6. Serum on 22 animals: Inoculated 7 rec/3 died, Uninoculated 3/9. \(\chi^2\) test of association. (Ans: \(\chi^2_{cal} = 2.82\))

Worked Problems on Small Sample Tests

Thirty-four problems in the order the textbook sets them, grouped by the test they use: five on a single mean, six on two independent means, four on paired samples, three on a correlation coefficient, six on goodness of fit, five on independence of attributes and five on the F-test. They are numbered 1–34 separately from the seven Extra Practical Problems above, and the rest of this page calls them “Worked Problem n”. Every one runs the same steps: null hypothesis, alternative hypothesis, test statistic under \(H_0\), conclusion at the stated level (5% where none is stated). The chapter's \(\chi^2\) test for a single variance (§7) is theory only; it sets no problem.

As the chapter does, these solutions write \(s^2\) for the sample variance with divisor \(n\) and \(S^2\) for the one with divisor \(n-1\); the two forms of each statistic give the same number (§1).

Source note. Every answer below was recomputed from the data before it was written down. Where the book's printed figure differs, the reason is given in place. Misprints in data, labels and intermediate values are corrected, with the printed figure kept in the note. Rounding artefacts, where the printed value follows from the book's own rounded intermediates, are left standing with a note (Worked Problems 12–15, 21, 22, 24 and 32). In five problems (5, 12, 14, 29 and 33) the book's working or conclusion is wrong; the correct one is worked here and the book's version is recorded.

A. t-test for a Single Mean

WORKED PROBLEM 1 — steel rods, σ known

The specified mean breaking strength of steel rods is 18.5 thousand lb, with standard deviation 1.955 thousand lb. A sample of 14 rods gave a mean breaking strength of 17.85 thousand lb. Is the difference significant?

Given \(\mu_0 = 18.5\), \(\sigma = 1.955\), \(n = 14\), \(\bar x = 17.85\).

(1) Null hypothesis. \(H_0: \mu = 18.5\): the rods meet the specification; the difference is not significant.

(2) Alternative hypothesis. \(H_1: \mu \ne 18.5\). Two-tailed test.

(3) Test statistic under \(H_0\). The population standard deviation is known, so by Case 1 of §2 the normal test applies even though the sample is small:

\[ z = \frac{\bar x - \mu_0}{\sigma/\sqrt n} = \frac{17.85 - 18.5}{1.955/\sqrt{14}} = \frac{-0.65}{1.955/3.7417} = \frac{-0.65}{0.5225} = -1.244 . \]

(4) Conclusion. \(|z| = 1.244\), and the two-tailed 5% value is \(z_\alpha = 1.96\). Since \(|z| < z_\alpha\), \(H_0\) is accepted: the difference is not significant.

Printing note. The source's conclusion line reads “\(|z| < z_\alpha \Rightarrow\) we reject \(H_0\)” and then “i.e. not significant”. “Reject” is a slip for “accept”: both the inequality and the stated meaning say accept.

WORKED PROBLEM 2 — bolts from a factory

The average diameter of the bolts made in a factory is 21 mm. A sample of 25 bolts has mean diameter 22.6 mm and standard deviation 3 mm. Can the sample be regarded as drawn from this population at the 5% level?

Given \(\mu_0 = 21\), \(n = 25\), \(\bar x = 22.6\), \(s = 3\); \(\sigma\) is unknown.

(1) \(H_0: \mu = 21\): the sample is drawn from the population. (2) \(H_1: \mu \ne 21\). Two-tailed.

(3)

\[ t = \frac{\bar x - \mu_0}{s/\sqrt{n-1}} = \frac{22.6 - 21}{3/\sqrt{24}} = \frac{1.6}{3/4.8990} = \frac{1.6}{0.6124} = 2.613 \sim t_{24} . \]

(4) The two-tailed 5% value for 24 d.f. is \(t_{0.025,24} = 2.064\). Since \(|t| = 2.613 > 2.064\), \(H_0\) is rejected: the sample is not drawn from a population with mean 21 mm.

−2.064 2.064 t = 2.613 t = 0 μ = 21 in the rejection region → reject H₀ ±t = ±2.064, 24 d.f.
Fig 4.2 — Worked Problem 2. The statistic 2.613 is past 2.064, so the bolts are judged not to come from a population with mean 21 mm. (The source's figure labels the left critical value −2.64; it is −2.064.)
WORKED PROBLEM 3 — axle diameters, one-tailed at 1%

The specified diameter of an engine axle is 1.75 mm. A sample of 10 parts has mean diameter 1.85 mm and standard deviation 0.1 mm. Test at the 1% level whether the mean diameter is more than 1.75 mm.

Given \(\mu_0 = 1.75\), \(n = 10\), \(\bar x = 1.85\), \(s = 0.1\).

(1) \(H_0: \mu = 1.75\). (2) \(H_1: \mu > 1.75\). One-tailed (right).

(3)

\[ t = \frac{\bar x - \mu_0}{s/\sqrt{n-1}} = \frac{1.85 - 1.75}{0.1/\sqrt 9} = \frac{0.10}{0.0333} = 3.00 \sim t_9 . \]

(4) The one-tailed 1% value for 9 d.f. is \(t_{0.01,9} = 2.821\). Since \(t = 3.00 > 2.821\), \(H_0\) is rejected: the mean diameter is more than 1.75 mm.

2.821 t = 3.00 t = 0 μ = 1.75 in the rejection region → reject H₀ t = 2.821, 9 d.f.
Fig 4.3 — Worked Problem 3. A right-tailed test at 1%: only the upper 1% is shaded, and t = 3 lies in it. (The source's figure labels the sample mean 1.8; the data give 1.85.)
WORKED PROBLEM 4 — IQ of ten boys

The IQs of 10 boys are 70, 120, 110, 101, 88, 83, 95, 98, 107, 100. Do these data support the assumption that the population mean IQ is 100?

\(x\)7012011010188839598107100\(\sum x = 972\)
\(x^2\)490014400121001020177446889902596041144910000\(\sum x^2 = 96312\)
\[ \bar x = \frac{972}{10} = 97.2, \qquad s^2 = \frac{\sum x^2}{n} - \bar x^2 = 9631.2 - 9447.84 = 183.36, \qquad s = 13.54 . \]

(1) \(H_0: \mu = 100\). (2) \(H_1: \mu \ne 100\). Two-tailed.

(3)

\[ t = \frac{\bar x - \mu_0}{s/\sqrt{n-1}} = \frac{97.2 - 100}{13.54/\sqrt 9} = \frac{-2.8}{4.514} = -0.62 \sim t_9 . \]

(4) \(|t| = 0.62 < t_{0.025,9} = 2.262\), so \(H_0\) is accepted: the data support a population mean IQ of 100.

Printing note. The source's table prints \(100^2\) as 1000; it is 10000, and the printed total 96312 already uses 10000.

WORKED PROBLEM 5 — weights below 66 kg? with confidence limits

The weights (kg) of 10 males are 62, 64, 67, 71, 69, 68, 70, 71, 72, 66. Test whether the average weight is below 66 kg, and find the 95% and 99% confidence limits for the population mean.

\(x\)62646771696870717266\(\sum x = 680\)
\(x^2\)3844409644895041476146244900504151844356\(\sum x^2 = 46336\)
\[ \bar x = 68, \qquad s^2 = 4633.6 - 68^2 = 4633.6 - 4624 = 9.6, \qquad s = 3.10 . \]

(1) \(H_0: \mu = 66\). (2) \(H_1: \mu < 66\). One-tailed (left): \(H_0\) is rejected only if \(t < -t_{0.05,9} = -1.833\).

(3)

\[ t = \frac{\bar x - \mu_0}{s/\sqrt{n-1}} = \frac{68 - 66}{3.10/\sqrt 9} = \frac{2}{1.033} = 1.94 \sim t_9 . \]

(4) The statistic is positive. A left-tailed test can reject only a large negative \(t\); \(t = +1.94\) is in the opposite tail, so \(H_0\) is accepted: the data give no evidence that the average weight is below 66 kg. They could not: the sample mean, 68, is above 66.

−1.833 t = 1.94 t = 0 μ = 66 outside the rejection region → accept H₀ t = −1.833, 9 d.f.
Fig 4.4 — Worked Problem 5. The alternative μ < 66 puts the whole rejection region in the left tail. The statistic is +1.94, on the other side of the centre, so H₀ stands. The source compared |t| with 1.833 and rejected.

Confidence limits. An interval has two ends, so it uses the two-tailed points (§2). With \(s/\sqrt{n-1} = 3.10/3 = 1.033\):

\[ \text{95\%:}\quad 68 \pm t_{0.025,9}(1.033) = 68 \pm 2.262(1.033) = 68 \pm 2.34 \;\Rightarrow\; (65.66,\; 70.34), \] \[ \text{99\%:}\quad 68 \pm t_{0.005,9}(1.033) = 68 \pm 3.250(1.033) = 68 \pm 3.36 \;\Rightarrow\; (64.64,\; 71.36). \]

Correction note. The source compares \(|t| = 1.94\) with 1.833, rejects \(H_0\) and concludes that the average weight is below 66 kg. That ignores the direction of the alternative: a statistic in the wrong tail can never support it. (Had the question asked whether the average is above 66, \(H_1: \mu > 66\), the same \(t = 1.94 > 1.833\) would reject \(H_0\) at 5%.) For the limits, the source uses the one-tailed points 1.833 and 2.821 and gets (66.11, 69.89) and (65.09, 70.92); those are 90% and 98% intervals, not 95% and 99%.

B. t-test for the Difference of Two Means

WORKED PROBLEM 6 — population variances known

Two samples of sizes 10 and 8 have means 29 and 32. The population variances are 7.52 and 6.84. Test whether the difference of means is significant.

Given \(n_1 = 10\), \(n_2 = 8\), \(\bar x = 29\), \(\bar y = 32\), \(\sigma_1^2 = 7.52\), \(\sigma_2^2 = 6.84\).

(1) \(H_0: \mu_1 = \mu_2\). (2) \(H_1: \mu_1 \ne \mu_2\). Two-tailed.

(3) Both population variances are known, so Case 3 of §3 applies, small samples included:

\[ z = \frac{\bar x - \bar y}{\sqrt{\dfrac{\sigma_1^2}{n_1} + \dfrac{\sigma_2^2}{n_2}}} = \frac{29 - 32}{\sqrt{\dfrac{7.52}{10} + \dfrac{6.84}{8}}} = \frac{-3}{\sqrt{0.752 + 0.855}} = \frac{-3}{1.2677} = -2.37 . \]

(4) \(|z| = 2.37 > 1.96\), so \(H_0\) is rejected: the samples are not drawn from populations with the same mean.

−1.96 1.96 z = −2.37 z = 0 μ₁ = μ₂ in the rejection region → reject H₀ ±zα = ±1.96
Fig 4.5 — Worked Problem 6. The population variances are known, so the normal curve applies to these small samples; z = −2.37 is beyond −1.96.
WORKED PROBLEM 7 — same population with SD 5? at 1%

Samples of sizes 10 and 12 have means 24 and 30. Test at the 1% level whether they can be regarded as drawn from the same population with standard deviation 5.

(1) \(H_0: \mu_1 = \mu_2\). (2) \(H_1: \mu_1 \ne \mu_2\). Two-tailed.

(3) A common, known \(\sigma = 5\): Case 2 of §3.

\[ z = \frac{\bar x - \bar y}{\sigma\sqrt{\dfrac{1}{n_1} + \dfrac{1}{n_2}}} = \frac{24 - 30}{5\sqrt{\dfrac{1}{10} + \dfrac{1}{12}}} = \frac{-6}{5(0.4282)} = \frac{-6}{2.141} = -2.80 . \]

(4) At 1%, two-tailed, \(z_\alpha = 2.58\). \(|z| = 2.80 > 2.58\), so \(H_0\) is rejected: the samples are not from the same population.

−2.58 2.58 z = −2.80 z = 0 μ₁ = μ₂ in the rejection region → reject H₀ ±zα = ±2.58
Fig 4.6 — Worked Problem 7. At 1% the critical values move out to ±2.58; z = −2.80 still clears them.
WORKED PROBLEM 8 — salaries in two groups

Group A: 12 employees, mean salary 1050, standard deviation 68. Group B: 10 employees, mean 980, standard deviation 74. Test whether the mean salaries differ.

The population standard deviations are unknown, so the \(t\)-test (Case 1 of §3).

\[ S^2 = \frac{n_1 s_1^2 + n_2 s_2^2}{n_1 + n_2 - 2} = \frac{12(68)^2 + 10(74)^2}{20} = \frac{55488 + 54760}{20} = 5512.4, \qquad S = 74.25 . \]

(1) \(H_0: \mu_1 = \mu_2\). (2) \(H_1: \mu_1 \ne \mu_2\). Two-tailed.

(3)

\[ t = \frac{\bar x - \bar y}{S\sqrt{\dfrac{1}{n_1} + \dfrac{1}{n_2}}} = \frac{1050 - 980}{74.25\sqrt{\dfrac{1}{12} + \dfrac{1}{10}}} = \frac{70}{74.25(0.4282)} = \frac{70}{31.79} = 2.20 \sim t_{20} . \]

(4) \(|t| = 2.20 > t_{0.025,20} = 2.086\), so \(H_0\) is rejected: the mean salaries differ significantly.

−2.086 2.086 t = 2.20 t = 0 μ₁ = μ₂ in the rejection region → reject H₀ ±t = ±2.086, 20 d.f.
Fig 4.7 — Worked Problem 8. x̄ − ȳ = +70, so the statistic is +2.20, just inside the right rejection region. (The source's figure labels it t = −2.2.)

Printing note. The source writes the statistic as \(|z| = 2.20\); it is \(t\), as its own table value from the \(t\) table shows.

WORKED PROBLEM 9 — two types of bulb, one-tailed at 1%

Type I bulbs: 8 tested, mean life 1234 hours, standard deviation 36. Type II: 7 tested, mean 1036, standard deviation 40. Is Type I superior at the 1% level?

\[ S^2 = \frac{8(36)^2 + 7(40)^2}{8 + 7 - 2} = \frac{10368 + 11200}{13} = 1659.08, \qquad S = 40.73 . \]

(1) \(H_0: \mu_1 = \mu_2\). (2) \(H_1: \mu_1 > \mu_2\). One-tailed (right).

(3)

\[ t = \frac{1234 - 1036}{40.73\sqrt{\dfrac18 + \dfrac17}} = \frac{198}{40.73(0.5175)} = \frac{198}{21.08} = 9.39 \sim t_{13} . \]

(4) \(t = 9.39 > t_{0.01,13} = 2.650\), so \(H_0\) is rejected: Type I is superior.

2.650 t = 9.39 → off the scale t = 0 μ₁ = μ₂ in the rejection region → reject H₀ t = 2.650, 13 d.f.
Fig 4.8 — Worked Problem 9. A right-tailed test at 1%. The statistic 9.39 lies far beyond the drawable range, so the marker sits at the edge.
WORKED PROBLEM 10 — two small samples from the raw data

Two independent samples of sizes 8 and 7 gave the values below. Is the difference of means significant?

\(x\)1917152116181614\(\sum x = 136\)
\(x^2\)361289225441256324256196\(\sum x^2 = 2348\)
\(y\)15141519151816\(\sum y = 112\)
\(y^2\)225196225361225324256\(\sum y^2 = 1812\)
\[ \bar x = \frac{136}{8} = 17, \quad \bar y = \frac{112}{7} = 16, \quad s_1^2 = \frac{2348}{8} - 17^2 = 4.5, \quad s_2^2 = \frac{1812}{7} - 16^2 = 2.86 . \] \[ S^2 = \frac{8(4.5) + 7(2.857)}{13} = \frac{36 + 20}{13} = 4.308, \qquad S = 2.08 . \]

(1) \(H_0: \mu_1 = \mu_2\). (2) \(H_1: \mu_1 \ne \mu_2\). Two-tailed.

(3)

\[ t = \frac{17 - 16}{2.08\sqrt{\dfrac18 + \dfrac17}} = \frac{1}{2.08(0.5175)} = \frac{1}{1.074} = 0.93 \sim t_{13} . \]

(4) \(|t| = 0.93 < t_{0.025,13} = 2.160\), so \(H_0\) is accepted: the difference of means is not significant.

−2.160 2.160 t = 0.93 t = 0 μ₁ = μ₂ outside the rejection region → accept H₀ ±t = ±2.160, 13 d.f.
Fig 4.9 — Worked Problem 10. The statistic 0.93 sits well inside the acceptance region.
WORKED PROBLEM 11 — are sailors taller than soldiers?

Heights (inches) of 6 sailors: 63, 65, 68, 69, 71, 72. Heights of 10 soldiers: 61, 62, 65, 66, 69, 69, 70, 71, 72, 73. Test at 5% whether sailors are on average taller than soldiers.

Sailors \(x\)636568697172\(\sum x = 408,\ \sum x^2 = 27804\)
Soldiers \(y\)61, 62, 65, 66, 69, 69, 70, 71, 72, 73\(\sum y = 678,\ \sum y^2 = 46122\)
\[ \bar x = \frac{408}{6} = 68, \qquad s_1^2 = \frac{27804}{6} - 68^2 = 4634 - 4624 = 10, \] \[ \bar y = \frac{678}{10} = 67.8, \qquad s_2^2 = \frac{46122}{10} - 67.8^2 = 4612.2 - 4596.84 = 15.36 . \] \[ S^2 = \frac{6(10) + 10(15.36)}{14} = \frac{213.6}{14} = 15.26, \qquad S = 3.91 . \]

(1) \(H_0: \mu_1 = \mu_2\). (2) \(H_1: \mu_1 > \mu_2\). One-tailed (right).

(3)

\[ t = \frac{68 - 67.8}{3.91\sqrt{\dfrac16 + \dfrac1{10}}} = \frac{0.2}{3.91(0.5164)} = \frac{0.2}{2.017} = 0.099 \sim t_{14} . \]

(4) \(t = 0.099 < t_{0.05,14} = 1.761\), so \(H_0\) is accepted: the data do not show that sailors are taller on average.

C. Paired t-test

WORKED PROBLEM 12 — haemoglobin before and after treatment

A treatment was given to 10 patients. Their HB percentages before (\(x\)) and after (\(y\)) were as below. Has the treatment increased HB?

Before \(x\)11.9411.9911.9812.0312.0311.9611.9511.9611.9212.00
After \(y\)12.0011.9911.9512.0712.0311.9812.0312.0212.0111.99
\(d = x - y\)−0.0600.03−0.040−0.02−0.08−0.06−0.090.01

The same patients are measured twice, so the samples are paired.

\[ \sum d = -0.31, \qquad \sum d^2 = 0.0247, \qquad \bar d = \frac{-0.31}{10} = -0.031, \] \[ s^2 = \frac{0.0247}{10} - (0.031)^2 = 0.002470 - 0.000961 = 0.001509, \qquad s = 0.03885 . \]

(1) \(H_0: \mu_1 = \mu_2\): the treatment does not change HB.

(2) \(H_1: \mu_1 < \mu_2\): HB is higher after treatment. With \(d = x - y\), that means \(\mu_d < 0\): one-tailed, left.

(3)

\[ t = \frac{\bar d}{s/\sqrt{n-1}} = \frac{-0.031}{0.03885/3} = \frac{-0.031}{0.01295} = -2.39 \sim t_9 . \]

(4) The left-tailed 5% point is \(-1.833\). \(t = -2.39 < -1.833\), so \(H_0\) is rejected: HB has increased after the treatment.

−1.833 t = −2.39 t = 0 μ₁ = μ₂ in the rejection region → reject H₀ t = −1.833, 9 d.f.
Fig 4.10 — Worked Problem 12. With d = before − after, an increase in HB makes d̄ negative, so the test is left-tailed. t = −2.39 lies in that tail: HB increased.

Correction note. The source takes \(H_1: \mu_1 > \mu_2\), the wrong direction for “increased” when \(x\) is the reading before; it then finds \(|t| > 1.833\), rejects \(H_0\), and concludes “HB percent was not increased”. Rejecting \(H_0\) against the correct alternative means the opposite: HB increased. Rounding note. The source prints \(t = -2.40\), from \(s\) rounded to 0.0388; the unrounded value is \(-2.39\).

WORKED PROBLEM 13 — blood pressure after a stimulus

A stimulus given to 12 patients produced these increases in blood pressure: 5, 2, 8, −1, 3, 0, −2, 1, 5, 0, 4, 6. Can it be concluded that the stimulus in general increases blood pressure?

The data are already differences: \(d\) = after − before.

\[ \sum d = 31, \qquad \sum d^2 = 185, \qquad \bar d = \frac{31}{12} = 2.583, \] \[ s^2 = \frac{185}{12} - 2.583^2 = 15.417 - 6.674 = 8.743, \qquad s = 2.957 . \]

(1) \(H_0\): the stimulus does not change blood pressure, \(\mu_d = 0\). (2) \(H_1\): it increases it, \(\mu_d > 0\). One-tailed (right).

(3)

\[ t = \frac{\bar d}{s/\sqrt{n-1}} = \frac{2.583}{2.957/\sqrt{11}} = \frac{2.583}{0.8915} = 2.90 \sim t_{11} . \]

(4) \(t = 2.90 > t_{0.05,11} = 1.796\), so \(H_0\) is rejected: the stimulus increases blood pressure.

Rounding note. The source rounds \(\bar d\) to 2.58 before squaring, giving \(s^2 = 8.76\), \(s = 2.96\) and \(t = 2.89\); unrounded, \(s^2 = 8.743\) and \(t = 2.90\).

WORKED PROBLEM 14 — marks before and after coaching

Marks of 11 students before (\(x\)) and after (\(y\)) a course of coaching:

Before \(x\)1923162417182018211920
After \(y\)1724202420222020182219
\(d = x - y\)2−1−40−3−40−23−31

Did the students benefit from the coaching? The same students are tested twice: paired.

\[ \sum d = -11, \quad \sum d^2 = 69, \quad \bar d = -1, \quad s^2 = \frac{69}{11} - 1 = 5.273, \quad s = 2.296 . \]

(1) \(H_0: \mu_1 = \mu_2\): the coaching is of no benefit. (2) \(H_1: \mu_1 < \mu_2\): marks are higher after coaching (\(\mu_d < 0\)). One-tailed (left).

(3)

\[ t = \frac{\bar d}{s/\sqrt{n-1}} = \frac{-1}{2.296/\sqrt{10}} = \frac{-1}{0.7261} = -1.38 \sim t_{10} . \]

(4) The left-tailed 5% point is \(-1.812\). \(t = -1.38\) is not below it, so \(H_0\) is accepted: the data do not show that the coaching benefited the students.

Correction note. The source accepts \(H_0\) too, but words the conclusion as “coaching was a benefit to the students”. Accepting \(H_0\) means the opposite: no significant benefit was shown. Rounding note. The source's \(t = -1.37\) comes from \(s\) rounded to 2.30; unrounded, \(-1.38\).

WORKED PROBLEM 15 — two foods for cows, unpaired and paired

Two foods A and B were tried on 8 cows; the increases in weight were:

Food A \(x\)4953515247505253\(\sum x = 407,\ \sum x^2 = 20737\)
Food B \(y\)5255525350545453\(\sum y = 423,\ \sum y^2 = 22383\)

Is food B better than food A (i) if the two samples are independent, (ii) if the same 8 cows received both foods?

In both parts \(H_0: \mu_1 = \mu_2\) (no difference) and \(H_1: \mu_1 < \mu_2\) (B gives a larger increase). One-tailed (left).

(i) Independent samples.

\[ \bar x = 50.875, \quad \bar y = 52.875, \quad s_1^2 = \frac{20737}{8} - 50.875^2 = 3.859, \quad s_2^2 = \frac{22383}{8} - 52.875^2 = 2.109 , \] \[ S^2 = \frac{8(3.859) + 8(2.109)}{14} = \frac{47.75}{14} = 3.411, \qquad S = 1.847 , \] \[ t = \frac{50.875 - 52.875}{1.847\sqrt{\dfrac18 + \dfrac18}} = \frac{-2}{1.847(0.5)} = \frac{-2}{0.9234} = -2.17 \sim t_{14} . \]

The left-tailed 5% point for 14 d.f. is \(-1.761\); \(t = -2.17 < -1.761\), so \(H_0\) is rejected: food B is better.

(ii) The same cows. Now each cow gives a pair, so the paired test applies, with \(d = x - y\): −3, −2, −1, −1, −3, −4, −2, 0.

\[ \sum d = -16, \quad \sum d^2 = 44, \quad \bar d = -2, \quad s^2 = \frac{44}{8} - 4 = 1.5, \quad s = 1.225 , \] \[ t = \frac{\bar d}{s/\sqrt{n-1}} = \frac{-2}{1.225/\sqrt 7} = \frac{-2}{0.4629} = -4.32 \sim t_7 . \]

The left-tailed 5% point for 7 d.f. is \(-1.895\); \(t = -4.32 < -1.895\), so \(H_0\) is rejected: food B is better.

Why (ii) is so much stronger. The mean difference is \(-2\) both times, but pairing removes the cow-to-cow variation from the standard error, which falls from 0.923 to 0.463. The same difference, measured against half the noise, gives twice the \(t\).

Rounding notes. The source prints \(t = -2.16\) in (i), from \(S\) rounded to 1.85, and \(t = -4.34\) in (ii), from \(s\) rounded to 1.22; unrounded, \(-2.17\) and \(-4.32\). Printing note. In (ii) the source's table value 2.015 is the 5% point for 5 d.f.; for 7 d.f. it is 1.895. The verdict is the same.

D. t-test for a Correlation Coefficient

WORKED PROBLEM 16 — income and expenditure

The correlation between income and expenditure of 20 families is 0.203. Is it significant?

Given \(n = 20\), \(r = 0.203\).

(1) \(H_0: \rho = 0\), the variables are uncorrelated. (2) \(H_1: \rho \ne 0\). Two-tailed.

(3) By the t-test for a correlation coefficient,

\[ t = \frac{r\sqrt{n-2}}{\sqrt{1 - r^2}} = \frac{0.203\sqrt{18}}{\sqrt{1 - 0.0412}} = \frac{0.203(4.2426)}{0.9792} = \frac{0.8613}{0.9792} = 0.88 \sim t_{18} . \]

(4) \(|t| = 0.88 < t_{0.025,18} = 2.101\), so \(H_0\) is accepted: income and expenditure are uncorrelated in the population, as far as this sample can tell.

Printing note. The source's “Given” line reads \(r = 0.23\); its working uses 0.203, as the question does. (With 0.23, \(t\) would be 1.00, with the same verdict.)

WORKED PROBLEM 17 — 27 pairs with r = 0.6

A sample of 27 pairs from a normal population gives \(r = 0.6\). Is it significant?

(1) \(H_0: \rho = 0\). (2) \(H_1: \rho \ne 0\).

(3)

\[ t = \frac{r\sqrt{n-2}}{\sqrt{1 - r^2}} = \frac{0.6\sqrt{25}}{\sqrt{1 - 0.36}} = \frac{3}{0.8} = 3.75 \sim t_{25} . \]

(4) \(|t| = 3.75 > t_{0.025,25} = 2.060\), so \(H_0\) is rejected: the variables are correlated.

WORKED PROBLEM 18 — n = 200, a test and confidence limits

A bivariate sample of 200 from a normal population gives \(r = 0.4\). (i) Is it significant? (ii) Find 95% and 99% confidence limits for \(\rho\).

(i) \(H_0: \rho = 0\) against \(H_1: \rho \ne 0\).

\[ t = \frac{0.4\sqrt{198}}{\sqrt{1 - 0.16}} = \frac{0.4(14.071)}{0.9165} = \frac{5.628}{0.9165} = 6.14 \sim t_{198} . \]

For 198 d.f. the \(t\) curve is almost the normal one: \(t_{0.025,198} = 1.972\), against the normal 1.96 the source uses. Either way \(|t| = 6.14\) is far beyond it, so \(H_0\) is rejected: the variables are correlated.

(ii) For a large sample, \(r\) has standard error approximately \((1 - r^2)/\sqrt n = 0.84/\sqrt{200} = 0.84/14.142 = 0.0594\), so

\[ \text{95\%:}\quad 0.4 \pm 1.96(0.0594) = 0.4 \pm 0.1164 \;\Rightarrow\; (0.2836,\; 0.5164), \] \[ \text{99\%:}\quad 0.4 \pm 2.58(0.0594) = 0.4 \pm 0.1532 \;\Rightarrow\; (0.2468,\; 0.5532). \]

A refinement. Away from \(\rho = 0\) the distribution of \(r\) is skewed, and Fisher's \(Z\) (Unit 3) is the more accurate route: \(Z = \tfrac12\ln\dfrac{1.4}{0.6} = 0.4236\), with standard error \(1/\sqrt{197} = 0.0712\), gives \(0.4236 \pm 1.96(0.0712) = (0.2840,\; 0.5633)\) on the \(Z\) scale, and transforming back with \(\tanh\), the 95% interval \((0.2766,\; 0.5104)\). It is not centred on \(r\); it leans towards zero, as the skewness requires.

Printing note. The source prints the limits' formula with \(\sqrt 2\) in the denominator; \(\sqrt{200}\) is meant, and its values use it.

E. χ² Test for Goodness of Fit

WORKED PROBLEM 19 — random digits

300 digits taken from a random number table gave these frequencies. Are the digits equally frequent?

Digit0123456789
\(O_i\)28293331263532303125
\(E_i\)30303030303030303030
\((O_i - E_i)^2/E_i\)0.13330.03330.30000.03330.53330.83330.133300.03330.8333

(1) \(H_0\): the digits are equally frequent (uniform distribution), so each \(E_i = 300/10 = 30\). (2) \(H_1\): they are not.

(3)

\[ \chi^2 = \sum\frac{(O_i - E_i)^2}{E_i} = \frac{4 + 1 + 9 + 1 + 16 + 25 + 4 + 0 + 1 + 25}{30} = \frac{86}{30} = 2.8667 \sim \chi^2_9 . \]

(4) \(\chi^2 = 2.87 < \chi^2_{0.05,9} = 16.92\), so \(H_0\) is accepted: the digits are equally frequent.

WORKED PROBLEM 20 — Mendel's 9 : 3 : 3 : 1

Of 560 beans, the four groups A, B, C, D contained 319, 101, 108 and 32. Theory predicts the ratio 9 : 3 : 3 : 1. Do the data support it?

Expected frequencies: \(560 \times \tfrac{9}{16} = 315\), \(560 \times \tfrac{3}{16} = 105\), 105, \(560 \times \tfrac{1}{16} = 35\).

Group\(O_i\)\(E_i\)\(O_i - E_i\)\((O_i - E_i)^2/E_i\)
A31931540.0508
B101105−40.1524
C10810530.0857
D3235−30.2571
Total56056000.5460

(1) \(H_0\): the data follow 9 : 3 : 3 : 1. (2) \(H_1\): they do not.

(3) \(\chi^2 = 0.5460 \sim \chi^2_3\) (4 classes, nothing estimated).

(4) \(0.546 < \chi^2_{0.05,3} = 7.81\), so \(H_0\) is accepted: the data support Mendel's theory.

WORKED PROBLEM 21 — aircraft accidents by day of the week

84 aircraft accidents were distributed over the week as below. Are accidents uniformly distributed over the days?

DaySunMonTueWedThuFriSat
\(O_i\)141681211914
\((O_i - 12)^2/12\)0.33331.33331.333300.08330.75000.3333

(1) \(H_0\): accidents are uniform over the week, \(E_i = 84/7 = 12\). (2) \(H_1\): they are not.

(3)

\[ \chi^2 = \frac{4 + 16 + 16 + 0 + 1 + 9 + 4}{12} = \frac{50}{12} = 4.1667 \sim \chi^2_6 . \]

(4) \(4.17 < \chi^2_{0.05,6} = 12.59\), so \(H_0\) is accepted: the accidents are uniformly distributed over the week.

Notes. The source labels the third day “Thu”; it is Tuesday. Its total 4.1665 is the sum of the four-place terms; exactly it is \(50/12 = 4.1667\).

WORKED PROBLEM 22 — fitting a binomial

Fit a binomial distribution to the data below and test the goodness of fit.

\(x\)0123456
\(f\)5182812764

Fitting. \(N = \sum f = 80\) and \(\sum fx = 192\), so \(\bar x = 2.4\). For a binomial with \(n = 6\), the mean is \(np\), so \(p = 2.4/6 = 0.4\) and \(q = 0.6\). Then \(P(x) = \binom{6}{x}(0.4)^x(0.6)^{6-x}\) and \(E = 80\,P(x)\):

\(x\)0123456
\(P(x)\)0.04670.18660.31100.27650.13820.03690.0041
\(E\)3.7314.9324.8822.1211.062.950.33
\(E\) rounded41525221130

Pooling. \(E\) for \(x = 0\) is below 5, so it joins \(x = 1\); \(E\) for \(x = 5\) and \(6\) are below 5 and even together (3.28) stay below it, so they join \(x = 4\). Four classes remain:

Class\(O\)\(E\)\(O - E\)\((O - E)^2/E\)
0–1231940.84
2282530.36
31222−104.55
4–6171430.64
Total808006.39

(1) \(H_0\): the binomial is a good fit. (2) \(H_1\): it is not.

(3) \(\chi^2 = 6.39\). Degrees of freedom, by §5: \(k = 7\) classes, \(K = 3\) lost to pooling (two cells into one loses 1, three into one loses 2), \(l = 1\) parameter (\(p\)) estimated: \(7 - 1 - 3 - 1 = 2\). Counting the four pooled classes gives the same, \(4 - 1 - 1 = 2\).

(4) \(\chi^2 = 6.39 > \chi^2_{0.05,2} = 5.99\), so \(H_0\) is rejected: the binomial is not a suitable fit.

Notes. The source prints \(P(6)\) as 0.0004 and \(E(4)\) as 1.06; they are 0.0041 and 11.06, and its rounded column (0 and 11) uses the right values. It prints \(P(1)\) as 0.1867; \(6(0.4)(0.6)^5 = 0.186624\). Using the unrounded expected frequencies (18.66, 24.88, 22.12, 14.34) gives \(\chi^2 = 6.52\) instead of 6.39; the verdict is the same.

WORKED PROBLEM 23 — boys and girls in families of five

320 families with 5 children each were classified by the number of boys. Are male and female births equally probable?

Boys012345
Families \(O\)1456110884012
\(E = 320\binom5x/32\)10501001005010
\((O - E)^2/E\)1.60.721.01.442.00.4

(1) \(H_0\): male and female births are equally probable, \(p = q = \tfrac12\), so \(P(x) = \binom5x(\tfrac12)^5 = \binom5x/32\). (2) \(H_1\): they are not.

(3) \(\chi^2 = 1.6 + 0.72 + 1.0 + 1.44 + 2.0 + 0.4 = 7.16\). Here \(p\) is given by the hypothesis, not estimated, so d.f. \(= 6 - 1 = 5\).

(4) \(7.16 < \chi^2_{0.05,5} = 11.07\), so \(H_0\) is accepted: male and female births are equally probable.

Printing note. The source prints the last term as “04”; it is \(2^2/10 = 0.4\).

WORKED PROBLEM 24 — fitting a Poisson

Fit a Poisson distribution to the data below and test the goodness of fit.

\(x\)012345
\(f\)142156692751

Fitting. \(N = 400\), \(\sum fx = 400\), so \(\bar x = 1\), and the Poisson mean is estimated by \(\lambda = 1\). \(P(x) = e^{-1}/x!\), \(E = 400\,P(x)\):

\(x\)012345
\(P(x)\)0.36790.36790.18390.06130.01530.0031
\(E\)147.15147.1573.5824.536.131.23
\(E\) rounded147147742561

Pooling. \(E\) for \(x = 5\) is below 5, so it joins \(x = 4\): \(O = 6\), \(E = 7\).

\(x\)\(O\)\(E\)\((O - E)^2/E\)
01421470.1701
11561470.5510
269740.3378
327250.1600
4–5670.1429
Total4004001.3618

(1) \(H_0\): the Poisson is a good fit. (2) \(H_1\): it is not.

(3) \(\chi^2 = 1.3618\), with d.f. \(= 6 - 1 - 1 - 1 = 3\) (\(K = 1\) lost to pooling, \(l = 1\) for \(\lambda\)).

(4) \(1.36 < \chi^2_{0.05,3} = 7.81\), so \(H_0\) is accepted: the Poisson fits well.

Rounding note. The \(\chi^2\) above uses the expected frequencies rounded to whole numbers, as the source does. With unrounded ones, and the last class taken as “4 or more” (\(E = 400\,P(X \ge 4) = 7.60\)) so that the expected total is exactly 400, \(\chi^2 = 1.58\); the verdict is the same. The source's \(E = 147.16\) is \(400 \times 0.3679\); from the unrounded probability it is 147.15.

F. χ² Test for Independence of Attributes

WORKED PROBLEM 25 — the 2 × 2 formula

For the \(2 \times 2\) contingency table with cells \(a, b\) / \(c, d\) and \(N = a + b + c + d\), prove that \(\chi^2 = \dfrac{N(ad - bc)^2}{(a+b)(c+d)(a+c)(b+d)}\).

This is proved step by step in §6. In outline: each expected frequency is (row total)(column total)/\(N\); every one of the four deviations \(O - E\) works out to \(\pm(ad - bc)/N\), so \(\chi^2\) is \((ad - bc)^2/N^2\) times the sum of the four reciprocals \(1/E\); and that sum simplifies, because \((a+c) + (b+d) = N\) and \((a+b) + (c+d) = N\), to \(N^3/[(a+b)(c+d)(a+c)(b+d)]\). Multiplying gives the formula.

Printing note. The source's proof prints \((ad + bc)^2\) in its penultimate line, for \((ad - bc)^2\).

WORKED PROBLEM 26 — two treatments on 500 plots

Two treatments were applied to 500 agricultural plots, with the results below. Test whether the treatments are independent.

Treatment I (rows), II (columns)B₁B₂Total
A₁208 (\(a\))92 (\(b\))300
A₂32 (\(c\))168 (\(d\))200
Total240260500

(1) \(H_0\): treatments I and II are independent. (2) \(H_1\): they are not.

(3) By the \(2 \times 2\) formula,

\[ \chi^2 = \frac{500(208 \times 168 - 92 \times 32)^2}{300 \times 200 \times 240 \times 260} = \frac{500(34944 - 2944)^2}{3.744 \times 10^9} = \frac{500(32000)^2}{3.744 \times 10^9} = 136.75 \sim \chi^2_1 . \]

(4) \(136.75 > \chi^2_{0.05,1} = 3.84\), so \(H_0\) is rejected: the two treatments are dependent.

WORKED PROBLEM 27 — rural and urban votes

Sample polls of votes for two candidates A and B in rural and urban areas gave the results below. Is the nature of the area related to voting?

AreaABTotal
Rural6203801000
Urban5504501000
Total11708302000

(1) \(H_0\): the area is independent of voting. (2) \(H_1\): they are related.

(3)

\[ \chi^2 = \frac{2000(620 \times 450 - 380 \times 550)^2}{1000 \times 1000 \times 1170 \times 830} = \frac{2000(279000 - 209000)^2}{9.711 \times 10^{11}} = \frac{2000(70000)^2}{9.711 \times 10^{11}} = 10.09 \sim \chi^2_1 . \]

(4) \(10.09 > \chi^2_{0.05,1} = 3.84\), so \(H_0\) is rejected: the nature of the area is related to voting.

WORKED PROBLEM 28 — physical and mental ability, 3 × 3 at 1%

The physical and mental abilities of 1000 students are classified below. Test at the 1% level whether they are independent.

Physical (rows), mental (columns)HighAverageBelow averageTotal
High39271480
Average260252178690
Below average419198230
Total3403702901000

(1) \(H_0\): physical and mental abilities are independent. (2) \(H_1\): they are not.

(3) Expected frequencies \(E_{ij} = (A_i)(B_j)/N\): for example \(E_{11} = 80 \times 340/1000 = 27.2\).

\(O_{ij}\)392714260252178419198
\(E_{ij}\)27.229.623.2234.6255.3200.178.285.166.7
\(O - E\)11.8−2.6−9.225.4−3.3−22.1−37.25.931.3
\((O - E)^2/E\)5.11910.22843.64832.75000.04272.440817.69620.409014.6880
\[ \chi^2 = \sum_i\sum_j\frac{(O_{ij} - E_{ij})^2}{E_{ij}} = 47.0225 \sim \chi^2_{(3-1)(3-1)} = \chi^2_4 . \]

(4) \(47.02 > \chi^2_{0.01,4} = 13.28\), so \(H_0\) is rejected: physical and mental abilities are related. Most of the statistic comes from the “below average” physical row, where far fewer students than expected are mentally high and far more are below average.

WORKED PROBLEM 29 — two researchers' sampling techniques

Two researchers used different sampling techniques on the same group of students and classified them by intelligence level. Are the techniques significantly different?

ResearcherBelow averageAverageAbove averageGeniusTotal
I86604410200
II4033252100
Total126936912300

(1) \(H_0\): there is no significant difference between the two techniques (classification is independent of researcher). (2) \(H_1\): there is.

(3) Expected frequencies: row I 84, 62, 46, 8; row II 42, 31, 23, 4. The expected frequency 4, for researcher II's geniuses, is below 5, so cells must be pooled. In a contingency table the pooling has to keep the table rectangular, so the “Above average” and “Genius” columns are merged in both rows, giving a \(2 \times 3\) table:

ResearcherBelow averageAverageAbove average or genius
I: \(O\) (\(E\))86 (84)60 (62)54 (54)
II: \(O\) (\(E\))40 (42)33 (31)27 (27)
\[ \chi^2 = \frac{2^2}{84} + \frac{2^2}{62} + 0 + \frac{2^2}{42} + \frac{2^2}{31} + 0 = 0.0476 + 0.0645 + 0.0952 + 0.1290 = 0.336 \sim \chi^2_{(2-1)(3-1)} = \chi^2_2 . \]

(4) \(0.336 < \chi^2_{0.05,2} = 5.99\), so \(H_0\) is accepted: the two sampling techniques do not differ significantly.

Correction note. The source pools only researcher II's last two cells (25 + 2 against 23 + 4) and keeps researcher I's four cells, including the “Genius” term \((10 - 8)^2/8 = 0.5\); it gets \(\chi^2 = 0.923\) and takes \(3 - 1 = 2\) d.f. Pooling one row only leaves a table that is no longer a contingency table, so its degrees of freedom are not \((r-1)(s-1)\) of anything; merging whole columns is the regular procedure. Without any pooling \(\chi^2 = 2.10\) on 3 d.f. All three versions accept \(H_0\). (The source also prints \(4/62 = 0.0645\) as 0.064.)

G. F-test for Equality of Variances

WORKED PROBLEM 30 — standard deviations 2.9 and 2.6

Two samples of sizes 9 and 12 from normal populations have standard deviations 2.9 and 2.6. Test whether the population variances differ.

\[ S_1^2 = \frac{n_1 s_1^2}{n_1 - 1} = \frac{9(2.9)^2}{8} = \frac{75.69}{8} = 9.46, \qquad S_2^2 = \frac{n_2 s_2^2}{n_2 - 1} = \frac{12(2.6)^2}{11} = \frac{81.12}{11} = 7.37 . \]

(1) \(H_0: \sigma_1^2 = \sigma_2^2\). (2) \(H_1: \sigma_1^2 \ne \sigma_2^2\).

(3) \(S_1^2\) is the larger, so it goes on top (§8):

\[ F = \frac{S_1^2}{S_2^2} = \frac{9.46}{7.37} = 1.28 \sim F_{(8,\,11)} . \]

(4) \(1.28 < F_{0.05}(8, 11) = 2.95\), so \(H_0\) is accepted: the population variances do not differ significantly.

Notes. The source's \(H_1\) reads “difference between sample means”; the test is of variances. Strictly, with a two-sided \(H_1\) at 5% the upper 2.5% point, \(F_{0.025}(8, 11) = 3.66\), is the comparison; \(F = 1.28\) is below both.

WORKED PROBLEM 31 — sums of squares given, at 1%

A sample of 8 has sum of squares of deviations from its mean 84.4; another of 10 has 102.6. Is the difference between the variances significant at 1%?

Since \(n s^2 = \sum(x - \bar x)^2\), the unbiased estimate is the sum of squares over \(n - 1\):

\[ S_1^2 = \frac{84.4}{7} = 12.06, \qquad S_2^2 = \frac{102.6}{9} = 11.40 . \]

(1) \(H_0: \sigma_1^2 = \sigma_2^2\). (2) \(H_1: \sigma_1^2 \ne \sigma_2^2\).

(3) \(F = \dfrac{12.06}{11.40} = 1.06 \sim F_{(7,\,9)}\).

(4) \(1.06 < F_{0.01}(7, 9) = 5.61\), so \(H_0\) is accepted.

Note. The source reads the table value as 5.62; to three decimals it is 5.613.

WORKED PROBLEM 32 — equality of variances from the raw data

Two samples of sizes 10 and 12 are given. Test the equality of the population variances.

\(x\)10616171312815914\(\sum x = 120,\ \sum x^2 = 1560\)
\(y\)713221512141882123107\(\sum y = 170,\ \sum y^2 = 2774\)
\[ S_1^2 = \frac{1}{n_1 - 1}\left[\sum x^2 - \frac{(\sum x)^2}{n_1}\right] = \frac{1560 - 1440}{9} = \frac{120}{9} = 13.33, \] \[ S_2^2 = \frac{1}{n_2 - 1}\left[\sum y^2 - \frac{(\sum y)^2}{n_2}\right] = \frac{2774 - 2408.33}{11} = \frac{365.67}{11} = 33.24 . \]

(1) \(H_0: \sigma_1^2 = \sigma_2^2\). (2) \(H_1: \sigma_1^2 \ne \sigma_2^2\).

(3) \(S_2^2\) is the larger, so it goes on top and the degrees of freedom follow it:

\[ F = \frac{S_2^2}{S_1^2} = \frac{33.24}{13.33} = 2.49 \sim F_{(11,\,9)} . \]

(4) \(2.49 < F_{0.05}(11, 9) = 3.10\), so \(H_0\) is accepted: the population variances may be regarded as equal.

Rounding note. The source rounds \(\bar y = 14.1667\) to 14.17 and gets \(S_2^2 = 33.14\); exactly it is 33.24. \(F\) is 2.49 either way. Its intermediate \(s_2^2\) is printed 30.398; with \(\bar y = 14.17\) it is 30.378 (which is what its 33.14 uses), and exactly 30.472.

WORKED PROBLEM 33 — unbiased estimates given

Two samples of sizes 10 and 15 give unbiased estimates of the population variances, 5 and 9. Can the population variances be regarded as equal?

Unbiased estimates are the \(S^2\) themselves: \(S_1^2 = 5\) (9 d.f.), \(S_2^2 = 9\) (14 d.f.). Nothing is converted.

(1) \(H_0: \sigma_1^2 = \sigma_2^2\). (2) \(H_1: \sigma_1^2 \ne \sigma_2^2\).

(3) \(F = \dfrac{S_2^2}{S_1^2} = \dfrac{9}{5} = 1.80 \sim F_{(14,\,9)}\).

(4) \(1.80 < F_{0.05}(14, 9) = 3.03\), so \(H_0\) is accepted: the population variances may be regarded as equal.

Correction note. The source treats 5 and 9 as \(s^2\) (divisor \(n\)) and converts them: \(S_1^2 = 10(5)/9 = 5.56\), \(S_2^2 = 15(9)/14 = 9.64\), \(F = 1.73\), against a table value of 3.02. The question says the estimates are already unbiased, so the conversion inflates both. The verdict happens to be the same.

WORKED PROBLEM 34 — the same normal population? both tests

Sample I: size 10, mean 15, sum of squares of deviations from the mean 90. Sample II: size 12, mean 14, sum of squares 108. Can the samples be regarded as drawn from the same normal population, at 5%?

“The same normal population” means the same variance and the same mean, so two tests are needed, the F-test first, because the \(t\)-test for means assumes equal variances (§3).

I. F-test. \(H_0: \sigma_1^2 = \sigma_2^2\) against \(H_1: \sigma_1^2 \ne \sigma_2^2\).

\[ S_1^2 = \frac{90}{9} = 10, \qquad S_2^2 = \frac{108}{11} = 9.82, \qquad F = \frac{10}{9.82} = 1.02 \sim F_{(9,\,11)} . \]

\(1.02 < F_{0.05}(9, 11) = 2.90\): \(H_0\) is accepted; the variances may be taken as equal.

II. t-test. \(H_0: \mu_1 = \mu_2\) against \(H_1: \mu_1 \ne \mu_2\). Since \(n_1 s_1^2 = \sum(x - \bar x)^2 = 90\) and \(n_2 s_2^2 = 108\),

\[ S = \sqrt{\frac{90 + 108}{10 + 12 - 2}} = \sqrt{9.9} = 3.15, \qquad t = \frac{15 - 14}{3.15\sqrt{\dfrac1{10} + \dfrac1{12}}} = \frac{1}{3.15(0.4282)} = \frac{1}{1.347} = 0.74 \sim t_{20} . \]

\(|t| = 0.74 < t_{0.025,20} = 2.086\): \(H_0\) is accepted; the means may be taken as equal.

Conclusion. Both hypotheses are accepted, so the two samples may be regarded as drawn from the same normal population.

Where this goes next. Small samples cost you the CLT, but every test in this unit still assumed the parent population was normal. Unit 5 asks what is left when even that is unavailable — ordinal data, an obviously skewed sample, an outlier you cannot justify removing. Replacing the observations by their ranks buys a test valid for any continuous distribution; the price is a little power in the case where the data really was normal after all.

→ Unit 5 — Non-parametric Tests