What is new in this unit is what happens when a cell holds more than one observation. That single change buys an interaction term and a pure error term, and it costs the orthogonality of the design as soon as the cell frequencies stop being proportional — which is the hardest idea in the unit and the one section 3 is about.
With one observation per cell the model is
\[ y_{ij} = \mu + \alpha_i + \beta_j + e_{ij}, \]and the residual sum of squares has \((a-1)(b-1)\) degrees of freedom. That residual is doing double duty: it carries the genuine experimental error and any interaction between the two factors, and there is no way to tell them apart. The two-way analysis with one observation per cell therefore has to assume additivity, and cannot test it.
With \(m > 1\) observations per cell the model becomes
\[ y_{ijk} = \mu + \alpha_i + \beta_j + (\alpha\beta)_{ij} + e_{ijk}, \qquad i = 1,\ldots,a,\; j = 1,\ldots,b,\; k = 1,\ldots,m, \]with the side conditions \(\sum_i \alpha_i = \sum_j \beta_j = 0\) and \(\sum_i(\alpha\beta)_{ij} = \sum_j(\alpha\beta)_{ij} = 0\). Now the variation within a cell estimates \(\sigma^{2}\) on its own, because the \(m\) observations in a cell differ only by error — every fixed effect is the same for all of them. The interaction can therefore be separated and tested.
Write \(G\) for the grand total, \(A_i\) for the \(i\)th row total (over \(bm\) observations), \(B_j\) for the \(j\)th column total (over \(am\)), \(T_{ij}\) for the cell total (over \(m\)), and \(N = abm\). With the correction factor \(\text{CF} = G^{2}/N\):
\[ SS_{\text{total}} = \sum_{ijk} y_{ijk}^{2} - \text{CF}, \qquad SS_A = \frac{\sum_i A_i^{2}}{bm} - \text{CF}, \qquad SS_B = \frac{\sum_j B_j^{2}}{am} - \text{CF}, \] \[ SS_{\text{cells}} = \frac{\sum_{ij}T_{ij}^{2}}{m} - \text{CF}, \qquad SS_{AB} = SS_{\text{cells}} - SS_A - SS_B, \qquad SS_E = SS_{\text{total}} - SS_{\text{cells}}. \]The last two lines are the whole idea. \(SS_{\text{cells}}\) measures how much the \(ab\) cell means differ from one another; \(SS_A\) and \(SS_B\) are the parts of that explained by the two factors acting separately; whatever is left over is interaction. And the error is what remains within cells, on \(ab(m-1)\) degrees of freedom.
| Source | df | Mean square | \(F\) |
|---|---|---|---|
| Factor A | \(a-1\) | \(SS_A/(a-1)\) | \(MS_A/MS_E\) |
| Factor B | \(b-1\) | \(SS_B/(b-1)\) | \(MS_B/MS_E\) |
| Interaction AB | \((a-1)(b-1)\) | \(SS_{AB}/[(a-1)(b-1)]\) | \(MS_{AB}/MS_E\) |
| Error | \(ab(m-1)\) | \(SS_E/[ab(m-1)]\) | |
| Total | \(abm-1\) |
Test the interaction first. If it is significant, the main effect \(F\) tests are of limited interest: saying “factor A matters” is not meaningful when how much it matters depends on the level of B. The right report then describes the cell means, not the marginal ones.
Given. \(a = 3\) levels of A, \(b = 2\) of B, \(m = 3\) observations per cell, \(N = 18\).
| B₁ | B₂ | Row total \(A_i\) | |
|---|---|---|---|
| A₁ | 7, 9, 8 (24) | 12, 10, 14 (36) | 60 |
| A₂ | 10, 12, 11 (33) | 15, 17, 16 (48) | 81 |
| A₃ | 6, 8, 7 (21) | 9, 11, 13 (33) | 54 |
| Column total \(B_j\) | 78 | 117 | 195 |
Step 1 — the correction factor.
\[ \text{CF} = \frac{195^{2}}{18} = \frac{38025}{18} = 2112.5. \]Step 2 — the total. The raw sum of squares of all 18 observations is \(2289\), so
\[ SS_{\text{total}} = 2289 - 2112.5 = 176.5. \]Step 3 — factor A. Each row total covers \(bm = 6\) observations:
\[ SS_A = \frac{60^{2} + 81^{2} + 54^{2}}{6} - 2112.5 = \frac{3600 + 6561 + 2916}{6} - 2112.5 = \frac{13077}{6} - 2112.5, \] \[ = 2179.5 - 2112.5 = 67. \]Step 4 — factor B. Each column total covers \(am = 9\):
\[ SS_B = \frac{78^{2} + 117^{2}}{9} - 2112.5 = \frac{6084 + 13689}{9} - 2112.5 = 2197 - 2112.5 = 84.5. \]Step 5 — the cells. Each cell total covers \(m = 3\):
\[ SS_{\text{cells}} = \frac{24^{2}+36^{2}+33^{2}+48^{2}+21^{2}+33^{2}}{3} - 2112.5 = \frac{6795}{3} - 2112.5 = 2265 - 2112.5 = 152.5. \]Step 6 — interaction by subtraction.
\[ SS_{AB} = 152.5 - 67 - 84.5 = 1.0. \]Step 7 — error by subtraction.
\[ SS_E = 176.5 - 152.5 = 24, \qquad \text{on } ab(m-1) = 3 \times 2 \times 2 = 12 \text{ degrees of freedom.} \]Step 8 — the table.
| Source | SS | df | MS | \(F\) | \(p\) |
|---|---|---|---|---|---|
| A | 67.0 | 2 | 33.5 | 16.75 | 0.000337 |
| B | 84.5 | 1 | 84.5 | 42.25 | 0.000029 |
| AB | 1.0 | 2 | 0.5 | 0.25 | 0.782758 |
| Error | 24.0 | 12 | 2.0 | ||
| Total | 176.5 | 17 |
Interpretation. The interaction is nowhere near significant \((F = 0.25)\), so the main effects can be read as they stand: both factors matter, and B rather more strongly than A relative to its single degree of freedom. Note that the interaction mean square, \(0.5\), is smaller than the error mean square, \(2.0\) — which is exactly what an \(F\) below 1 means and is perfectly ordinary when the true interaction is zero.
The cell means are \(8, 12\) for A₁, \(11, 16\) for A₂ and \(7, 11\) for A₃. The B₂ − B₁ gap is \(4\), \(5\), \(4\): nearly constant, which is additivity, and a picture says it faster than the table does.
Real experiments lose observations, so cells end up with \(n_{ij}\) observations rather than a common \(m\). There is one case in which nothing goes wrong: when the frequencies are proportional,
\[ n_{ij} = \frac{n_{i\cdot}\,n_{\cdot j}}{N} \qquad \text{for every cell.} \]The condition says that the row and column classifications are independent in the table of counts — it is the same arithmetic as the expected frequencies of a chi-square test of independence. When it holds, the row and column contrasts remain orthogonal, the three sums of squares still add up to \(SS_{\text{cells}}\), and the analysis proceeds as before with the divisors changed from \(bm\), \(am\) and \(m\) to \(n_{i\cdot}\), \(n_{\cdot j}\) and \(n_{ij}\).
Given. A \(2 \times 2\) layout with \(N = 15\):
| B₁ | B₂ | \(n_{i\cdot}\) | \(A_i\) | |
|---|---|---|---|---|
| A₁ | 5, 7, 6, 6 (\(n=4\), total 24) | 9, 11 (\(n=2\), total 20) | 6 | 44 |
| A₂ | 8, 10, 9, 9, 11, 7 (\(n=6\), total 54) | 13, 15, 14 (\(n=3\), total 42) | 9 | 96 |
| \(n_{\cdot j}\) | 10 | 5 | 15 | |
| \(B_j\) | 78 | 62 | 140 |
Step 1 — check proportionality.
\[ \frac{n_{1\cdot}n_{\cdot 1}}{N} = \frac{6 \times 10}{15} = 4 = n_{11}, \qquad \frac{6 \times 5}{15} = 2 = n_{12}, \] \[ \frac{9 \times 10}{15} = 6 = n_{21}, \qquad \frac{9 \times 5}{15} = 3 = n_{22}. \checkmark \]All four agree, so the design is orthogonal and the ordinary formulae apply.
Step 2 — the correction factor and the total. The raw sum of squares is \(1434\), so
\[ \text{CF} = \frac{140^{2}}{15} = \frac{19600}{15} = 1306.666667, \qquad SS_{\text{total}} = 1434 - 1306.666667 = 127.333333. \]Step 3 — each term with its own divisor.
\[ SS_A = \frac{44^{2}}{6} + \frac{96^{2}}{9} - \text{CF} = 322.666667 + 1024 - 1306.666667 = 40, \] \[ SS_B = \frac{78^{2}}{10} + \frac{62^{2}}{5} - \text{CF} = 608.4 + 768.8 - 1306.666667 = 70.533333, \] \[ SS_{\text{cells}} = \frac{24^{2}}{4} + \frac{20^{2}}{2} + \frac{54^{2}}{6} + \frac{42^{2}}{3} - \text{CF} = 144 + 200 + 486 + 588 - 1306.666667 = 111.333333. \]Step 4 — interaction and error.
\[ SS_{AB} = 111.333333 - 40 - 70.533333 = 0.8, \qquad SS_E = 127.333333 - 111.333333 = 16, \]on \(N - ab = 15 - 4 = 11\) degrees of freedom, so \(MS_E = 16/11 = 1.454545\).
Step 5 — the tests.
| Source | SS | df | MS | \(F\) | \(p\) |
|---|---|---|---|---|---|
| A | 40.000000 | 1 | 40.000000 | 27.500000 | 0.000275 |
| B | 70.533333 | 1 | 70.533333 | 48.491667 | 0.000024 |
| AB | 0.800000 | 1 | 0.800000 | 0.550000 | 0.473855 |
| Error | 16.000000 | 11 | 1.454545 | ||
| Total | 127.333333 | 14 |
Interpretation. The three components \(40 + 70.533333 + 0.8 = 111.333333\) reproduce \(SS_{\text{cells}}\) exactly. That addition is the orthogonality, and it is worth checking every time — when it fails, the frequencies were not proportional and the next section applies.
When the \(n_{ij}\) are not proportional, the row classification and the column classification are no longer orthogonal: part of the variation could be attributed to either. Applying the ordinary formulae anyway does not merely lose a little accuracy — it can produce an arithmetic impossibility, and Example 1.3 produces one.
The repair is to abandon subtraction and fit models. Write \(R(\cdot)\) for the reduction in the residual sum of squares achieved by fitting a set of terms. Then the meaningful quantities are the adjusted sums of squares
\[ R(A \mid \mu, B) = R(\mu, A, B) - R(\mu, B), \qquad R(B \mid \mu, A) = R(\mu, A, B) - R(\mu, A), \]each asking “what does this factor add once the other is already in the model?” — exactly the question a \(t\) statistic answers in a multiple regression. They do not add up to \(R(\mu, A, B)\), and they are not meant to. Fitting the additive model requires solving the normal equations, which is why the method is called the method of fitting constants.
Given. The same kind of layout, \(N = 15\), but now with cell counts \(4, 2, 3, 6\):
| B₁ | B₂ | \(n_{i\cdot}\) | \(A_i\) | |
|---|---|---|---|---|
| A₁ | 5, 7, 6, 6 (\(n=4\), mean 6) | 9, 11 (\(n=2\), mean 10) | 6 | 44 |
| A₂ | 8, 10, 9 (\(n=3\), mean 9) | 13, 15, 14, 12, 16, 14 (\(n=6\), mean 14) | 9 | 111 |
| \(n_{\cdot j}\) | 7 | 8 | 15 | |
| \(B_j\) | 51 | 104 | 155 |
Step 1 — confirm the frequencies are disproportionate. \(n_{1\cdot}n_{\cdot 1}/N = 6 \times 7/15 = 2.8\), which is not \(n_{11} = 4\). The design is not orthogonal.
Step 2 — compute the naive sums of squares anyway. With \(\text{CF} = 155^{2}/15 = 1601.666667\) and a raw sum of squares of \(1779\),
\[ SS_{\text{total}} = 177.333333, \qquad SS_A = \frac{44^{2}}{6} + \frac{111^{2}}{9} - \text{CF} = 90, \] \[ SS_B = \frac{51^{2}}{7} + \frac{104^{2}}{8} - \text{CF} = 371.571429 + 1352 - 1601.666667 = 121.904762, \] \[ SS_{\text{cells}} = \frac{24^{2}}{4} + \frac{20^{2}}{2} + \frac{27^{2}}{3} + \frac{84^{2}}{6} - \text{CF} = 161.333333. \]Step 3 — watch it fail.
\[ SS_A + SS_B = 90 + 121.904762 = 211.904762 \;>\; 161.333333 = SS_{\text{cells}}, \]so the interaction “by subtraction” would be \(161.333333 - 211.904762 = -50.571429\) — a negative sum of squares, which is impossible. The two factors have been credited with the same variation twice.
Step 4 — fit the additive model instead. Take \(\mu, \alpha, \beta\) with \(\alpha_2 = \beta_2 = 0\), so the four fitted cell means are \(\mu + \alpha + \beta\), \(\mu + \alpha\), \(\mu + \beta\), \(\mu\). Weighted least squares on the cell means with weights \(n_{ij}\) gives the normal equations
\[ 15\mu + 6\alpha + 7\beta = 155, \qquad 6\mu + 6\alpha + 4\beta = 44, \qquad 7\mu + 4\alpha + 7\beta = 51. \]Each is a weighted total: the first sums all fifteen observations, the second the six in row A₁, the third the seven in column B₁.
Step 5 — solve them. Subtracting the second from the first gives \(9\mu + 3\beta = 111\), that is \(\beta = 37 - 3\mu\). The second then gives \(\alpha = \mu - \tfrac{52}{3}\). Substituting both into the third,
\[ 7\mu + 4\left(\mu - \tfrac{52}{3}\right) + 7(37 - 3\mu) = 51 \;\Longrightarrow\; -10\mu = -\tfrac{416}{3} \;\Longrightarrow\; \mu = \frac{208}{15} = 13.866667, \] \[ \alpha = -\frac{52}{15} = -3.466667, \qquad \beta = -\frac{23}{5} = -4.6. \]Step 6 — the fitted cell means.
\[ 5.800000, \qquad 10.400000, \qquad 9.266667, \qquad 13.866667, \]against the observed \(6, 10, 9, 14\).
Step 7 — the reduction due to the additive model.
\[ R(\mu, A, B) = \sum_{ij} n_{ij}\,\hat c_{ij}^{2} - \text{CF} = 1762.2 - 1601.666667 = 160.533333. \]Step 8 — the adjusted sums of squares.
\[ R(A \mid \mu, B) = 160.533333 - 121.904762 = 38.628571, \] \[ R(B \mid \mu, A) = 160.533333 - 90 = 70.533333. \]Step 9 — interaction and error.
\[ SS_{AB} = SS_{\text{cells}} - R(\mu, A, B) = 161.333333 - 160.533333 = 0.8, \] \[ SS_E = SS_{\text{total}} - SS_{\text{cells}} = 177.333333 - 161.333333 = 16, \]on \(15 - 4 = 11\) degrees of freedom, \(MS_E = 1.454545\).
| Source | SS | df | MS | \(F\) | \(p\) |
|---|---|---|---|---|---|
| A adjusted for B | 38.628571 | 1 | 38.628571 | 26.557143 | 0.000317 |
| B adjusted for A | 70.533333 | 1 | 70.533333 | 48.491667 | 0.000024 |
| AB | 0.800000 | 1 | 0.800000 | 0.550000 | 0.473855 |
| Error | 16.000000 | 11 | 1.454545 |
Interpretation. Compare \(SS_A = 90\) unadjusted with \(R(A \mid \mu, B) = 38.628571\) adjusted: more than half of what A appeared to explain was really B's, reaching A only because the larger cells happened to sit in the corners of the table where B is large. The adjusted figure is the honest one, and it is the only one that can be reported. Note also that the three adjusted terms do not add to \(SS_{\text{cells}}\), and that this is not a defect: in a non-orthogonal design there is no unique partition, and insisting on one is what produced the negative number in Step 3.
A significant treatment \(F\) says only that the \(t\) treatment means are not all equal. It does not say which differ. The obvious repair — run a \(t\) test on every pair — inflates the error rate exactly as \(p\) separate tests do in multivariate analysis: with \(t\) treatments there are \(\binom{t}{2}\) pairs, and at \(t = 5\) that is ten tests.
Three standard answers, in increasing order of conservatism:
| Test | Critical difference | Protects |
|---|---|---|
| Fisher's LSD | \(t_{\alpha/2,\,f}\sqrt{2\,MS_E/r}\) | each comparison at \(\alpha\); nothing at the family level |
| Duncan's multiple range | \(R_p = r_p\sqrt{MS_E/r}\), \(p\) = number of means spanned | each range of \(p\) means at \(1-(1-\alpha)^{p-1}\) |
| Tukey's HSD | \(q_{\alpha}(t,f)\sqrt{MS_E/r}\) | the whole family at \(\alpha\) |
Here \(f\) is the error degrees of freedom, \(r\) the number of replicates per treatment, and \(q_\alpha(p, f)\) the upper \(\alpha\) point of the studentized range — the distribution of \(\left(\max_i \bar y_i - \min_i \bar y_i\right)/s_{\bar y}\) for \(p\) means.
Fisher's LSD is the plain two-sample \(t\) test using the pooled \(MS_E\), so it is only honest when the overall \(F\) has already rejected — the “protected” LSD. Used without that protection it is the multiple-testing error in its purest form.
Duncan's test sits between the two. Order the means; a set of \(p\) adjacent means is declared homogeneous unless its range exceeds \(R_p\). The critical values grow with \(p\), which is what stops a wide span from being declared significant merely because it is wide. Duncan takes the protection level for a span of \(p\) means to be \(\alpha_p = 1 - (1-\alpha)^{p-1}\) rather than \(\alpha\), which is why \(R_p\) rises far more slowly than Tukey's single critical value and why the test is more powerful and less conservative.
One identity worth knowing: at \(p = 2\) Duncan's protection level is \(\alpha\) itself, and \(q_\alpha(2, f) = \sqrt 2\,t_{\alpha/2, f}\), so \(R_2\) is the LSD. Duncan's test begins where Fisher's does and diverges from it only for wider spans.
Given. Four treatments, \(r = 5\) replicates each, in a completely randomised design.
| Treatment | Observations | Total | Mean |
|---|---|---|---|
| T₁ | 12, 14, 11, 13, 15 | 65 | 13 |
| T₂ | 18, 20, 17, 19, 21 | 95 | 19 |
| T₃ | 13, 15, 12, 14, 16 | 70 | 14 |
| T₄ | 20, 22, 19, 21, 23 | 105 | 21 |
| 335 |
Step 1 — the analysis of variance. With \(\text{CF} = 335^{2}/20 = 5611.25\) and a raw sum of squares of \(5875\),
\[ SS_{\text{total}} = 263.75, \qquad SS_{\text{tr}} = \frac{65^{2}+95^{2}+70^{2}+105^{2}}{5} - \text{CF} = 5835 - 5611.25 = 223.75, \] \[ SS_E = 263.75 - 223.75 = 40 \quad \text{on } 16 \text{ df}, \qquad MS_E = 2.5. \] \[ F = \frac{223.75/3}{2.5} = \frac{74.583333}{2.5} = 29.833333 \quad (p = 0.000001). \]The \(F\) test rejects, so a multiple comparison is licensed.
Step 2 — the two standard errors.
\[ s_{\bar y} = \sqrt{\frac{MS_E}{r}} = \sqrt{\frac{2.5}{5}} = \sqrt{0.5} = 0.707107, \qquad s_{\text{diff}} = \sqrt{\frac{2\,MS_E}{r}} = \sqrt{1} = 1. \]Step 3 — Fisher's LSD. With \(t_{16,\,0.025} = 2.119905\),
\[ \text{LSD} = 2.119905 \times 1 = 2.120. \]Step 4 — Duncan's significant ranges. The studentized range points at the protection levels \(\alpha_p = 1 - (0.95)^{p-1}\) are
| \(p\) | \(\alpha_p\) | \(r_p\) | \(R_p = r_p \times 0.707107\) |
|---|---|---|---|
| 2 | 0.0500 | 2.9980 | 2.120 |
| 3 | 0.0975 | 3.1438 | 2.223 |
| 4 | 0.1426 | 3.2349 | 2.287 |
\(R_2 = 2.120\) is the LSD, as the identity in section 4 requires: \(\sqrt 2 \times 2.119905 = 2.998 = r_2\). \(\checkmark\) The critical differences are given to three decimals because they are products of two rounded figures, and the fourth decimal of such a product is not reliable.
Step 5 — order the means and compare. Ordered: \(13\) (T₁), \(14\) (T₃), \(19\) (T₂), \(21\) (T₄).
| Comparison | Means spanned \(p\) | Range | \(R_p\) | Verdict |
|---|---|---|---|---|
| T₄ − T₁ | 4 | 8 | 2.287 | significant |
| T₄ − T₃ | 3 | 7 | 2.223 | significant |
| T₂ − T₁ | 3 | 6 | 2.223 | significant |
| T₄ − T₂ | 2 | 2 | 2.120 | not significant |
| T₂ − T₃ | 2 | 5 | 2.120 | significant |
| T₃ − T₁ | 2 | 1 | 2.120 | not significant |
Step 6 — the homogeneous groups.
\[ \underline{\text{T}_1 \quad \text{T}_3} \qquad\qquad \underline{\text{T}_2 \quad \text{T}_4} \]Step 7 — Tukey's HSD, for comparison. \(q_{0.05}(4, 16) = 4.0461\), so
\[ \text{HSD} = 4.0461 \times 0.707107 = 2.861. \]Against that, T₄ − T₂ = 2 is still not significant, but the conclusion for every pair happens to be unchanged here.
Interpretation. All three tests agree on this data set, which is the comfortable case. They need not: Tukey's critical difference is \(2.861\) against Duncan's \(2.120\) to \(2.287\), so a difference of, say, \(2.5\) between adjacent means would be significant by Duncan and not by Tukey. The choice is a choice about which error you would rather make, and it must be made before looking at the data. The honest report names the test used.
Sometimes a nuisance variable \(x\) is measured on each unit and is known to affect the response — initial weight in a feeding trial, previous marks in a teaching experiment, soil nitrogen in a field trial. Blocking on it is one answer, but blocking needs the variable before the design is laid out and it throws away the quantitative detail. Analysis of covariance instead fits a regression on \(x\) and compares the treatments at a common value of it. The model for a one-way layout is
\[ y_{ij} = \mu + \tau_i + \beta\left(x_{ij} - \bar x_{\cdot\cdot}\right) + e_{ij}, \]which is an analysis of variance and a regression in one, and the analysis is a two-way bookkeeping of sums of squares and products.
The arithmetic. Compute \(S_{yy}\), \(S_{xx}\) and \(S_{xy}\) for each line of the ordinary analysis of variance table — treatments, error, total. Then the error sum of squares adjusted for the covariate is
\[ E_{yy}' = E_{yy} - \frac{E_{xy}^{2}}{E_{xx}}, \qquad \hat\beta = \frac{E_{xy}}{E_{xx}}, \]on one degree of freedom fewer, because \(\beta\) has been estimated. For the treatment test, do the same for treatments-plus-error and subtract:
\[ SS_{\text{tr}}' = \left[(T+E)_{yy} - \frac{(T+E)_{xy}^{2}}{(T+E)_{xx}}\right] - E_{yy}', \]on \(t - 1\) degrees of freedom, and
\[ F = \frac{SS_{\text{tr}}'/(t-1)}{E_{yy}'/(N - t - 1)}. \]Adjusted treatment means. Each treatment mean is moved to what it would have been had that treatment's units started at the average \(x\):
\[ \bar y_{i}' = \bar y_{i} - \hat\beta\left(\bar x_i - \bar x\right). \]The same machinery applies to a two-way layout: every line of the table gets its \(S_{xx}\), \(S_{xy}\), \(S_{yy}\), and each effect is tested after the same adjustment, with the error line supplying \(\hat\beta\).
The assumption to state. The regression slope \(\beta\) is taken to be the same for every treatment — parallel regression lines. If the slopes differ, the treatments cannot be compared at a single \(x\), and the right analysis is one that includes a treatment-by-covariate interaction. The covariate must also be unaffected by the treatment: a variable measured after treatment has been applied may itself have been changed by it, and adjusting for it then removes part of the effect being measured.