Analysis of Variance (ANOVA), due to R. A. Fisher (1925), is a statistical technique to test whether three or more population means are equal by partitioning the total variation in the data into components attributable to different sources.
The fundamental identity:
If the treatment SS is large compared to error SS, the means differ; this is tested by an F-ratio.
A \(t\)-test compares two means; for three or more, testing every pair separately inflates the chance of a false difference, and ANOVA tests them all at once. Fisher developed it in the 1920s for agricultural field trials. In his words, it is “the separation of the variance ascribable to one group of causes from the variance ascribable to other groups”. (The textbook quotes the last words as “to the group”, which loses the point: one group of causes is set against the others.)
The variation in any experiment has two kinds of cause:
Example. Four fertilisers are each applied to six plots of paddy and the yields recorded. ANOVA splits the total variation in the 24 yields into a part due to the fertilisers and a part due to chance, and asks whether the first is larger than the second can explain.
Note on the assumptions. The textbook lists three: independence, normality and additivity. The F-test also needs the third assumption above, a common variance \(\sigma^2\) in every group; the linear model below assumes it in writing \(D(y) = \sigma^2 I\).
Let \(X_1, \ldots, X_n\) be independent \(N(0, \sigma^2)\), and suppose
\[ \sum_{i=1}^n X_i^2 = Q_1 + Q_2 + \cdots + Q_k , \]where each \(Q_j\) is a quadratic form in the \(X_i\) with rank \(r_j\). Then \(Q_1/\sigma^2, \ldots, Q_k/\sigma^2\) are independent \(\chi^2\) variates with \(r_1, \ldots, r_k\) degrees of freedom if and only if \(r_1 + r_2 + \cdots + r_k = n\).
ANOVA uses it directly: the total sum of squares is split into pieces whose degrees of freedom add up to the total, so the pieces are independent chi-squares, and their mean squares can be compared by an F-ratio. (The textbook's sum runs to \(k\) on the left; it runs over all \(n\) variables.)
For \(n\) observations and \(m\) unknown parameters,
\[ y = A\beta + \varepsilon, \qquad y = \begin{bmatrix} y_1 \\ \vdots \\ y_n \end{bmatrix},\; A = \begin{bmatrix} a_{11} & \cdots & a_{1m} \\ \vdots & & \vdots \\ a_{n1} & \cdots & a_{nm} \end{bmatrix},\; \beta = \begin{bmatrix} \beta_1 \\ \vdots \\ \beta_m \end{bmatrix},\; \varepsilon = \begin{bmatrix} \varepsilon_1 \\ \vdots \\ \varepsilon_n \end{bmatrix}, \]with \(A\) a known \(n \times m\) matrix of coefficients, and
Then \(E(y_i) = a_{i1}\beta_1 + \cdots + a_{im}\beta_m\), \(\text{Var}(y_i) = \sigma^2\) and \(\text{Cov}(y_i, y_j) = 0\) for \(i \ne j\); \(D(y) = \sigma^2 I\) is the dispersion matrix of \(y\). Assumptions 1 and 2 are enough for least squares to give the best linear unbiased estimates; normality is added for the F-tests.
This unit uses the fixed effects model throughout.
One factor with \(k\) treatments (or levels). Within treatment \(i\), there are \(n_i\) observations. Total \(N = \sum n_i\).
with \(\sum_i n_i \alpha_i = 0\) and \(\epsilon_{ij} \sim N(0, \sigma^2)\) independently.
\(H_0: \alpha_1 = \alpha_2 = \cdots = \alpha_k = 0\) (all treatment means equal) vs \(H_1\): at least one differs.
Let \(T_i = \sum_j y_{ij}\) (treatment total), \(G = \sum_i T_i\) (grand total). Correction Factor \(C = G^2/N\).
| Source | SS | df | MS | F |
|---|---|---|---|---|
| Treatments | \(\text{SS}_{Tr}\) | \(k-1\) | \(\text{MS}_{Tr}\) | \(\text{MS}_{Tr}/\text{MS}_E\) |
| Error | \(\text{SS}_E\) | \(N-k\) | \(\text{MS}_E\) | — |
| Total | \(\text{SS}_T\) | \(N-1\) | — | — |
Reject \(H_0\) if \(F > F_{\alpha,\, k-1,\, N-k}\).
Yields (kg/plot) for 3 fertilizers, each in 4 plots:
| F1 | 22 | 26 | 24 | 28 |
|---|---|---|---|---|
| F2 | 30 | 34 | 32 | 36 |
| F3 | 20 | 22 | 24 | 22 |
\(T_1 = 100,\; T_2 = 132,\; T_3 = 88;\;G = 320;\; N = 12;\; C = 320^2/12 = 8533.33\).
\(\sum y^2 = 484 + 676 + 576 + 784 + 900 + 1156 + 1024 + 1296 + 400 + 484 + 576 + 484 = 8840\).
\(\text{SS}_T = 8840 - 8533.33 = 306.67\).
\(\text{SS}_{Tr} = (100^2 + 132^2 + 88^2)/4 - 8533.33 = (10000 + 17424 + 7744)/4 - 8533.33 = 8792 - 8533.33 = 258.67\).
\(\text{SS}_E = 306.67 - 258.67 = 48.0\).
F = (258.67/2)/(48/9) = 129.33/5.33 = 24.25. df = (2, 9). \(F_{0.05, 2, 9} = 4.26\) ⇒ reject \(H_0\); fertilizers differ significantly.
Yields for 3 varieties: V1 (n₁=3): 20, 24, 22; V2 (n₂=4): 28, 30, 32, 34; V3 (n₃=2): 18, 20.
\(T_1 = 66, T_2 = 124, T_3 = 38, G = 228, N = 9, C = 228^2/9 = 5776\).
\(\sum y^2 = 400 + 576 + 484 + 784 + 900 + 1024 + 1156 + 324 + 400 = 6048\).
\(\text{SS}_T = 6048 - 5776 = 272\).
\(\text{SS}_{Tr} = 66^2/3 + 124^2/4 + 38^2/2 - 5776 = 1452 + 3844 + 722 - 5776 = 242\).
\(\text{SS}_E = 272 - 242 = 30\). df: treatment 2, error 6.
\(F = (242/2)/(30/6) = 121/5 = 24.2 > F_{0.05, 2, 6} = 5.14\) ⇒ reject \(H_0\).
Here \(\mu_i = \mu + \alpha_i\) is the mean of treatment \(i\), \(\mu = \frac1N\sum n_i\mu_i\), and \(\alpha_i = \mu_i - \mu\), so \(\sum n_i\alpha_i = \sum n_i\mu_i - N\mu = 0\). Minimise the error sum of squares \(E = \sum_i\sum_j (y_{ij} - \mu - \alpha_i)^2\):
\[ \frac{\partial E}{\partial \mu} = -2\sum_i\sum_j (y_{ij} - \mu - \alpha_i) = 0 \;\Longrightarrow\; N\bar y_{..} - N\mu - \sum n_i\alpha_i = 0 \;\Longrightarrow\; \hat\mu = \bar y_{..}, \] \[ \frac{\partial E}{\partial \alpha_i} = -2\sum_j (y_{ij} - \mu - \alpha_i) = 0 \;\Longrightarrow\; n_i\bar y_{i.} - n_i\mu - n_i\alpha_i = 0 \;\Longrightarrow\; \hat\alpha_i = \bar y_{i.} - \bar y_{..} . \]The residual is \(\hat\varepsilon_{ij} = y_{ij} - \hat\mu - \hat\alpha_i = y_{ij} - \bar y_{i.}\).
Write each observation as estimate plus residual: \(y_{ij} - \bar y_{..} = (\bar y_{i.} - \bar y_{..}) + (y_{ij} - \bar y_{i.})\). Squaring and summing over \(i\) and \(j\),
\[ \sum_i\sum_j (y_{ij} - \bar y_{..})^2 = \sum_i n_i(\bar y_{i.} - \bar y_{..})^2 + \sum_i\sum_j (y_{ij} - \bar y_{i.})^2 \] \[ + \; 2\sum_i (\bar y_{i.} - \bar y_{..})\sum_j (y_{ij} - \bar y_{i.}) . \]The last term is zero because \(\sum_j (y_{ij} - \bar y_{i.}) = n_i\bar y_{i.} - n_i\bar y_{i.} = 0\). So TSS = SST + SSE.
Degrees of freedom. TSS has \(N - 1\) (the \(N\) deviations satisfy \(\sum\sum (y_{ij} - \bar y_{..}) = 0\)); SST has \(k - 1\) (the \(k\) deviations satisfy \(\sum n_i(\bar y_{i.} - \bar y_{..}) = 0\)); SSE has \(N - k\) (its \(N\) deviations satisfy one constraint \(\sum_j (y_{ij} - \bar y_{i.}) = 0\) in each of the \(k\) groups). The degrees of freedom add up as the sums of squares do: \((k - 1) + (N - k) = N - 1\).
Shortcut formulas. Expanding the squares gives TSS \(= \sum\sum y_{ij}^2 - G^2/N\) and SST \(= \sum T_i^2/n_i - G^2/N\), where \(G^2/N\) is the correction factor.
Averaging the model gives \(\bar y_{i.} = \mu + \alpha_i + \bar\varepsilon_{i.}\) and \(\bar y_{..} = \mu + \bar\varepsilon_{..}\) (using \(\sum n_i\alpha_i = 0\)), with \(E(\bar\varepsilon_{i.}^2) = \sigma^2/n_i\) and \(E(\bar\varepsilon_{..}^2) = \sigma^2/N\). Then
\[ \text{SST} = \sum n_i\big[\alpha_i + (\bar\varepsilon_{i.} - \bar\varepsilon_{..})\big]^2, \qquad \sum n_i(\bar\varepsilon_{i.} - \bar\varepsilon_{..})^2 = \sum n_i\bar\varepsilon_{i.}^2 - N\bar\varepsilon_{..}^2 , \]and the cross term has expectation zero, so
\[ E(\text{SST}) = \sum n_i\alpha_i^2 + \sum n_i\frac{\sigma^2}{n_i} - N\frac{\sigma^2}{N} = \sum n_i\alpha_i^2 + (k-1)\sigma^2 . \]Similarly \(\text{SSE} = \sum\sum (\varepsilon_{ij} - \bar\varepsilon_{i.})^2\) \(= \sum\sum\varepsilon_{ij}^2 - \sum n_i\bar\varepsilon_{i.}^2\), so \(E(\text{SSE}) = N\sigma^2 - k\sigma^2 = (N-k)\sigma^2\). Dividing by the degrees of freedom,
\[ E(\text{MS}_{Tr}) = \sigma^2 + \frac{1}{k-1}\sum n_i\alpha_i^2, \qquad E(\text{MS}_E) = \sigma^2 . \]The error mean square estimates \(\sigma^2\) whatever the treatments do; the treatment mean square estimates \(\sigma^2\) only when \(H_0\) holds, and something larger otherwise. Under \(H_0\), by Cochran's theorem, \(\text{SST}/\sigma^2\) and \(\text{SSE}/\sigma^2\) are independent \(\chi^2_{k-1}\) and \(\chi^2_{N-k}\), so
\[ F = \frac{\text{SST}/(k-1)}{\text{SSE}/(N-k)} = \frac{\text{MS}_{Tr}}{\text{MS}_E} \sim F_{k-1,\,N-k} . \]Large values of \(F\) count against \(H_0\), so the test uses the upper tail only. An \(F\) below 1 is never significant, and the ratio is never turned upside down to make it exceed 1.
When \(H_0\) is rejected, the next question is which treatments differ. Two means differ significantly at level \(\alpha\) when their difference exceeds the critical difference (least significant difference)
\[ \text{CD} = t_{\alpha/2,\,N-k}\,\sqrt{\text{MS}_E\Big(\frac{1}{n_i} + \frac{1}{n_j}\Big)} \;=\; t_{\alpha/2,\,N-k}\,\sqrt{\frac{2\,\text{MS}_E}{n}} \;\text{ when every } n_i = n . \]This is the two-sample \(t\)-test with the pooled error mean square in place of \(\sigma^2\) (which is why \(t\), on the error degrees of freedom, and not \(z\)); \(t_{\alpha/2}\) is the two-tailed point.
Wording. A non-significant \(F\) means the data show no evidence of a difference; it does not prove the means equal. “Accept \(H_0\)” in the textbook should be read as “do not reject \(H_0\)”.
Two factors A (rows, \(r\) levels) and B (columns, \(c\) levels). Each combination has one observation \(y_{ij}\).
with \(\sum \alpha_i = 0,\; \sum \beta_j = 0,\; \epsilon_{ij} \sim N(0, \sigma^2)\).
Let \(R_i = \sum_j y_{ij}\) (row totals), \(C_j = \sum_i y_{ij}\) (column totals), \(G = \sum y_{ij}\), \(N = rc\), \(C = G^2/N\).
| Source | SS | df | MS | F |
|---|---|---|---|---|
| Rows | \(\text{SS}_R\) | \(r-1\) | \(\text{MS}_R\) | \(\text{MS}_R/\text{MS}_E\) |
| Columns | \(\text{SS}_C\) | \(c-1\) | \(\text{MS}_C\) | \(\text{MS}_C/\text{MS}_E\) |
| Error | \(\text{SS}_E\) | \((r-1)(c-1)\) | \(\text{MS}_E\) | — |
| Total | \(\text{SS}_T\) | \(N-1\) | — | — |
3 fertilizers (rows) × 4 varieties (columns). Yields:
| F\V | V1 | V2 | V3 | V4 | Row total |
|---|---|---|---|---|---|
| F1 | 10 | 12 | 14 | 16 | 52 |
| F2 | 14 | 16 | 18 | 20 | 68 |
| F3 | 8 | 10 | 12 | 14 | 44 |
| Col total | 32 | 38 | 44 | 50 | 164 |
\(N = 12, C = 164^2/12 = 2241.33\).
\(\sum y^2 = 100 + 144 + 196 + 256 + 196 + 256 + 324 + 400 + 64 + 100 + 144 + 196 = 2376\).
\(\text{SS}_T = 2376 - 2241.33 = 134.67\).
\(\text{SS}_R = (52^2 + 68^2 + 44^2)/4 - 2241.33 = (2704 + 4624 + 1936)/4 - 2241.33 = 9264/4 - 2241.33 = 2316 - 2241.33 = 74.67\).
\(\text{SS}_C = (32^2 + 38^2 + 44^2 + 50^2)/3 - 2241.33 = 6904/3 - 2241.33 = 2301.33 - 2241.33 = 60\).
\(\text{SS}_E = 134.67 - 74.67 - 60 = 0\) (linear data give zero error). df: rows 2, cols 3, error 6.
4 students (rows) × 3 subjects (cols). \(\text{SS}_R = 70,\; \text{SS}_C = 30,\; \text{SS}_E = 60\). df: 3, 2, 6.
\(F_R = (70/3)/(60/6) = 23.33/10 = 2.33;\; F_{0.05, 3, 6} = 4.76\) ⇒ accept (no student effect).
\(F_C = (30/2)/(60/6) = 15/10 = 1.5;\; F_{0.05, 2, 6} = 5.14\) ⇒ accept (no subject effect).
Write \(k\) treatments (rows, \(i\)) and \(h\) varieties or blocks (columns, \(j\)), \(N = hk\); in the notation above \(r = k\), \(c = h\). With \(\sum\alpha_i = 0\) and \(\sum\beta_j = 0\), minimising \(\sum\sum (y_{ij} - \mu - \alpha_i - \beta_j)^2\) gives, exactly as in one-way,
\[ \hat\mu = \bar y_{..}, \qquad \hat\alpha_i = \bar y_{i.} - \bar y_{..}, \qquad \hat\beta_j = \bar y_{.j} - \bar y_{..}, \]and residual \(y_{ij} - \bar y_{i.} - \bar y_{.j} + \bar y_{..}\). Then
\[ y_{ij} - \bar y_{..} = (\bar y_{i.} - \bar y_{..}) + (\bar y_{.j} - \bar y_{..}) + (y_{ij} - \bar y_{i.} - \bar y_{.j} + \bar y_{..}), \]and on squaring and summing, all three cross products vanish (each contains a sum of deviations from a mean over a full row or column):
\[ \underbrace{\sum\sum (y_{ij} - \bar y_{..})^2}_{\text{TSS}} = \underbrace{h\sum_i (\bar y_{i.} - \bar y_{..})^2}_{\text{SST}} + \underbrace{k\sum_j (\bar y_{.j} - \bar y_{..})^2}_{\text{SSV}} \] \[ + \underbrace{\sum\sum (y_{ij} - \bar y_{i.} - \bar y_{.j} + \bar y_{..})^2}_{\text{SSE}} . \]Degrees of freedom: \(hk - 1\) \(= (k - 1) + (h - 1) + (h-1)(k-1)\), the error taking what is left over.
With \(\bar y_{i.} = \mu + \alpha_i + \bar\varepsilon_{i.}\), \(\bar y_{.j} = \mu + \beta_j + \bar\varepsilon_{.j}\), \(\bar y_{..} = \mu + \bar\varepsilon_{..}\) and \(E(\bar\varepsilon_{i.}^2) = \sigma^2/h\), \(E(\bar\varepsilon_{.j}^2) = \sigma^2/k\), \(E(\bar\varepsilon_{..}^2) = \sigma^2/hk\):
\[ E(\text{SST}) = h\sum\alpha_i^2 + (k-1)\sigma^2, \qquad E(\text{SSV}) = k\sum\beta_j^2 + (h-1)\sigma^2 . \]For the error, the residual equals \(\varepsilon_{ij} - \bar\varepsilon_{i.} - \bar\varepsilon_{.j} + \bar\varepsilon_{..}\); expanding its square and taking expectations term by term,
\[ E(\text{SSE}) = \big(hk + k + h + 1 - 2k - 2h\big)\sigma^2 = (h-1)(k-1)\sigma^2 . \]So \(\text{MS}_E\) always estimates \(\sigma^2\), and
\[ F_T = \frac{\text{MS}_T}{\text{MS}_E} \sim F_{k-1,\,(h-1)(k-1)} \;\text{ under } H_T: \alpha_i = 0, \] \[ F_V = \frac{\text{MS}_V}{\text{MS}_E} \sim F_{h-1,\,(h-1)(k-1)} \;\text{ under } H_V: \beta_j = 0 . \]Shortcuts: SST \(= \frac1h\sum T_{i.}^2 - \text{CF}\), SSV \(= \frac1k\sum T_{.j}^2 - \text{CF}\), SSE \(=\) TSS \(-\) SST \(-\) SSV.
In a factorial experiment several factors are varied together, so that both their individual (main) effects and their interactions can be estimated from the same trials — far more efficient than studying one factor at a time. In a \(2^k\) design each of \(k\) factors is set at two levels, low and high.
For a \(2^2\) design with factors A and B, the four treatment combinations are written \((1),\ a,\ b,\ ab\) (a letter present = that factor at its high level). The effects are contrasts:
A significant interaction means the effect of one factor depends on the level of the other — then the main effects must not be interpreted in isolation.
Yields: \((1) = 20,\ a = 40,\ b = 30,\ ab = 54\).
Factor A has the dominant effect; the interaction is negligible.
Yates' algorithm computes all contrasts systematically. List the responses in standard order \((1), a, b, ab\) and repeat, \(k\) times, the operation "sum the pairs, then difference the pairs":
| Std order | Response | Column 1 | Column 2 (contrast) | Identifies |
|---|---|---|---|---|
| (1) | 20 | 60 | 144 | Grand total |
| a | 40 | 84 | 44 | A |
| b | 30 | 20 | 24 | B |
| ab | 54 | 24 | 4 | AB |
The final column reproduces the contrasts (44, 24, 4) and the grand total (144). The same "add-then-subtract" scheme with three passes handles a \(2^3\) design (factors A, B, C and interactions AB, AC, BC, ABC).
Two problems in the textbook's order, one-way then two-way, followed by its two exercises with the answers checked. Each follows the same steps: hypotheses, totals, correction factor, sums of squares, ANOVA table, and the comparison with the tabulated \(F\).
Source note. Every figure was recomputed exactly, and the tabulated values of \(F\) and \(t\) were computed rather than read from a table. Both worked problems agree with the textbook; one exercise answer is corrected.
Three processes A, B and C are tested to see whether their outputs are equivalent:
| Process | Output | \(T_i\) | \(n_i\) | \(T_i^2/n_i\) | \(\sum_j y_{ij}^2\) | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| A | 10 | 12 | 13 | 11 | 10 | 14 | 15 | 13 | 98 | 8 | 1200.5 | 1224 |
| B | 9 | 11 | 10 | 12 | 13 | 55 | 5 | 605 | 615 | |||
| C | 11 | 10 | 15 | 14 | 12 | 13 | 75 | 6 | 937.5 | 955 | ||
| Total | 228 | 19 | 2743 | 2794 | ||||||||
Hypotheses. \(H_0: \mu_1 = \mu_2 = \mu_3\) (the processes have the same mean output) against \(H_1\): not all equal.
Sums of squares. \(G = 228\), \(N = 19\), \(\text{CF} = 228^2/19 = 51984/19 = 2736\).
\[ \text{TSS} = 2794 - 2736 = 58, \qquad \text{SST} = 2743 - 2736 = 7, \qquad \text{SSE} = 58 - 7 = 51 . \]| Source | SS | d.f. | MS | \(F\) |
|---|---|---|---|---|
| Between processes | 7 | 2 | 3.5 | \(3.5/3.1875 = 1.098\) |
| Within processes (error) | 51 | 16 | 3.1875 | |
| Total | 58 | 18 |
Decision. \(F_{0.05}(2, 16) = 3.63\), and \(1.098 < 3.63\) (\(p = 0.36\)), so \(H_0\) is not rejected: the three processes show no significant difference in mean output (means 12.25, 11.0 and 12.5).
Note. The textbook concludes that the outputs “are equal”. A non-significant result does not show that; it shows only that these 19 observations give no evidence of a difference.
Four doctors each try five medicines, A to E, one patient per combination. Test at the 1% level whether the medicines differ, and whether the doctors differ.
| Doctor | A | B | C | D | E | \(T_{.j}\) | \(T_{.j}^2\) |
|---|---|---|---|---|---|---|---|
| 1 | 12 | 16 | 18 | 21 | 24 | 91 | 8281 |
| 2 | 16 | 25 | 20 | 23 | 28 | 112 | 12544 |
| 3 | 14 | 20 | 23 | 16 | 20 | 93 | 8649 |
| 4 | 15 | 24 | 23 | 25 | 36 | 123 | 15129 |
| \(T_{i.}\) | 57 | 85 | 84 | 85 | 108 | \(G = 419\) | 44603 |
| \(T_{i.}^2\) | 3249 | 7225 | 7056 | 7225 | 11664 | 36419 |
Hypotheses. \(H_M\): the five medicines have equal effects; \(H_D\): the four doctors have equal effects.
Sums of squares. \(k = 5\) medicines, \(h = 4\) doctors, \(N = 20\), \(\text{CF} = 419^2/20 = 8778.05\), \(\sum\sum y^2 = 9367\).
\[ \text{TSS} = 9367 - 8778.05 = 588.95, \qquad \text{SSM} = \frac{36419}{4} - 8778.05 = 326.70, \] \[ \text{SSD} = \frac{44603}{5} - 8778.05 = 142.55, \qquad \text{SSE} = 588.95 - 326.70 - 142.55 = 119.70 . \]| Source | SS | d.f. | MS | \(F\) | \(F_{0.01}\) |
|---|---|---|---|---|---|
| Medicines | 326.70 | 4 | 81.675 | 8.19 | 5.41 |
| Doctors | 142.55 | 3 | 47.52 | 4.76 | 5.95 |
| Error | 119.70 | 12 | 9.975 | ||
| Total | 588.95 | 19 |
Decision. Medicines: \(8.19 > 5.41\) (\(p = 0.002\)), so the medicines differ significantly. Doctors: \(4.76 < 5.95\), so at the 1% level the doctors show no significant difference. (At 5% they would: \(F_{0.05}(3, 12) = 3.49\) and \(p = 0.021\). The level must be fixed before the data are seen.)
Which medicines differ? The means are A 14.25, B 21.25, C 21.00, D 21.25, E 27.00. With \(t_{0.005}(12) = 3.055\),
\[ \text{CD}_{1\%} = 3.055\sqrt{\frac{2 \times 9.975}{4}} = 3.055 \times 2.233 = 6.82 . \]A is significantly below B, D and E (differences 7.0, 7.0, 12.75); C, at 6.75 above A, just misses. E's lead over B, C and D (5.75 to 6.0) is not significant at 1%, though it is at 5%, where \(\text{CD}_{5\%} = 2.179 \times 2.233 = 4.87\).
Misprint. The textbook divides by 9.9745 for the medicines' \(F\); the error mean square is \(119.7/12 = 9.975\). The ratio rounds to 8.19 either way.
| City | 1 | 2 | 3 | 4 | 5 | 6 | 7 |
|---|---|---|---|---|---|---|---|
| A | 82 | 79 | 73 | 69 | 69 | 63 | 61 |
| B | 84 | 82 | 80 | 79 | 76 | 68 | 62 |
| C | 88 | 84 | 80 | 68 | 68 | 66 | 66 |
| D | 79 | 77 | 76 | 74 | 72 | 68 | 64 |
| Fertiliser | I | II | III | IV | V |
|---|---|---|---|---|---|
| 1 | 1.9 | 2.2 | 2.6 | 1.8 | 2.1 |
| 2 | 2.5 | 1.9 | 2.3 | 2.6 | 2.2 |
| 3 | 1.7 | 1.9 | 2.2 | 2.0 | 2.1 |
| 4 | 2.1 | 1.8 | 2.5 | 2.3 | 2.4 |
Additional worked problems with step-by-step procedures to support self-study, matching this unit's topics.
Assumptions: observations are independent, normally distributed, with equal variances. Three principles of design: replication, randomization, local control.
Grain yield of rice (kg/ha) from 7 insecticide treatments, each with 4 replications (a one-way layout). Treatment totals (T) and grand total \(G = 57110\), \(N = 28\):
| Treatment | R1 | R2 | R3 | R4 | Total |
|---|---|---|---|---|---|
| Dol-mix | 2537 | 2069 | 2104 | 1797 | 8507 |
| Ferterra | 3366 | 2591 | 2211 | 2544 | 10712 |
| DDT+γ-BHC | 2536 | 2459 | 2827 | 2385 | 10207 |
| Standard | 2387 | 2453 | 1556 | 2116 | 8512 |
| Dimecron-Boom | 1997 | 1679 | 1649 | 1859 | 7184 |
| Dimecron-Knap | 1796 | 1704 | 1904 | 1320 | 6724 |
| Control | 1401 | 1516 | 1270 | 1077 | 5264 |
Correction factor \(C = G^2/N = 57110^2/28 = 116484004\).
\(\text{TSS} = \sum y^2 - C = 7577412\); \(\text{Treatment SS} = \sum T_i^2/4 - C = 5587175\); \(\text{Error SS} = 7577412 - 5587175 = 1990238\).
| Source | df | SS | MS | \(F\) | \(F_{0.05}\) |
|---|---|---|---|---|---|
| Treatment | 6 | 5587174 | 931196 | 9.83 | 2.57 |
| Error | 21 | 1990238 | 94773 | ||
| Total | 27 | 7577412 |
Since \(F = 9.83 > 2.57\), treatments differ significantly. SE of difference \(=\sqrt{\frac{2\times94773}{4}} = 217.68\); \(CD_{0.05} = 217.68\times 2.08 = 452.70\) kg/ha, \(CD_{0.01} = 217.68\times 2.831 = 616.33\) kg/ha. Compared with the control, all treatments except Dimecron-Knap give a significant yield increase.