The Foundation section gives the \(2^{2}\) contrasts one at a time. Writing them out that way for \(k = 4\) means fifteen contrasts of sixteen terms each, and the sign errors are certain. There is a generating form that produces every one of them mechanically. For \(2^{3}\),
\[ [A] = (a-1)(b+1)(c+1), \qquad [AB] = (a-1)(b-1)(c+1), \qquad [ABC] = (a-1)(b-1)(c-1), \]expanded formally, with \(1\) standing for the combination \((1)\) and each product of letters standing for the combination containing them. The rule: a factor in the effect contributes \((\text{letter} - 1)\), a factor not in it contributes \((\text{letter} + 1)\). Expanding \([AB]\) gives
\[ (a-1)(b-1)(c+1) = abc + ab + c + (1) - ac - bc - a - b, \]which is the \(AB\) contrast with its eight signs, obtained without thinking about any of them. Equivalently, and more usefully at the desk: a combination enters with \(+\) when it shares an even number of letters with the effect, and with \(-\) when it shares an odd number.
The conversion to an effect and a sum of squares is the Foundation one with \(r\) replicates carried through:
\[ \text{effect} = \frac{[\,\cdot\,]}{r\,2^{k-1}}, \qquad SS = \frac{[\,\cdot\,]^{2}}{r\,2^{k}}, \]each on one degree of freedom, the \(SS\) divisor being \(\sum c_i^{2}\) for a contrast whose \(r2^{k}\) coefficients are all \(\pm 1\).
The algorithm itself is in the Foundation section, worked on a \(2^{2}\), which ends with the remark that three passes handle a \(2^{3}\). Example 2.1 below is that run, because the pattern of which column holds which effect is not obvious until it has been seen once.
Two checks that cost nothing, and that the two-pass case is too small to make the point of:
Run both. A single transposed pair in the standard order passes unnoticed otherwise, and it corrupts every effect from that pass onward.
Given. Three factors at two levels each, \(r = 2\) replicates, 16 observations, laid out in two blocks of eight (one complete replicate per block).
| Combination | Replicate I | Replicate II | Total |
|---|---|---|---|
| (1) | 9 | 11 | 20 |
| a | 16 | 16 | 32 |
| b | 12 | 14 | 26 |
| ab | 21 | 23 | 44 |
| c | 11 | 13 | 24 |
| ac | 18 | 20 | 38 |
| bc | 15 | 15 | 30 |
| abc | 25 | 27 | 52 |
| Replicate total | 127 | 139 | 266 |
Step 1 — run Yates's algorithm on the totals.
| Combination | Total | (1) | (2) | (3) | Effect | SS |
|---|---|---|---|---|---|---|
| (1) | 20 | 52 | 122 | 266 = \(G\) | ||
| a | 32 | 70 | 144 | 66 = [A] | 8.25 | 272.25 |
| b | 26 | 62 | 30 | 38 = [B] | 4.75 | 90.25 |
| ab | 44 | 82 | 36 | 14 = [AB] | 1.75 | 12.25 |
| c | 24 | 12 | 18 | 22 = [C] | 2.75 | 30.25 |
| ac | 38 | 18 | 20 | 6 = [AC] | 0.75 | 2.25 |
| bc | 30 | 14 | 6 | 2 = [BC] | 0.25 | 0.25 |
| abc | 52 | 22 | 8 | 2 = [ABC] | 0.25 | 0.25 |
Column (1) is formed from the totals: \(20+32 = 52\), \(26+44 = 70\), \(24+38 = 62\), \(30+52 = 82\), then \(32-20 = 12\), \(44-26 = 18\), \(38-24 = 14\), \(52-30 = 22\). Columns (2) and (3) repeat the operation on the column before.
Step 2 — check one contrast the long way.
\[ [A] = (32 + 44 + 38 + 52) - (20 + 26 + 24 + 30) = 166 - 100 = 66. \checkmark \]Step 3 — effects and sums of squares. With \(r = 2\) and \(k = 3\), the divisors are \(r2^{k-1} = 8\) for an effect and \(r2^{k} = 16\) for a sum of squares:
\[ A = \frac{66}{8} = 8.25, \qquad SS_A = \frac{66^{2}}{16} = \frac{4356}{16} = 272.25, \]and likewise down the table. Their sum is
\[ 272.25 + 90.25 + 12.25 + 30.25 + 2.25 + 0.25 + 0.25 = 407.75. \]Step 4 — the rest of the analysis of variance. With \(\text{CF} = 266^{2}/16 = 4422.25\) and a raw sum of squares of \(4842\),
\[ SS_{\text{total}} = 4842 - 4422.25 = 419.75, \qquad SS_{\text{rep}} = \frac{127^{2} + 139^{2}}{8} - \text{CF} = 4431.25 - 4422.25 = 9, \] \[ SS_E = 419.75 - 9 - 407.75 = 3 \quad \text{on } (r-1)(2^{k}-1) = 7 \text{ df}, \qquad MS_E = \frac{3}{7} = 0.428571. \]Step 5 — the table.
| Source | SS | df | \(F\) | \(p\) |
|---|---|---|---|---|
| Replicates | 9.00 | 1 | ||
| A | 272.25 | 1 | 635.250000 | 0.000000 |
| B | 90.25 | 1 | 210.583333 | 0.000002 |
| C | 30.25 | 1 | 70.583333 | 0.000067 |
| AB | 12.25 | 1 | 28.583333 | 0.001068 |
| AC | 2.25 | 1 | 5.250000 | 0.055702 |
| BC | 0.25 | 1 | 0.583333 | 0.469964 |
| ABC | 0.25 | 1 | 0.583333 | 0.469964 |
| Error | 3.00 | 7 | ||
| Total | 419.75 | 15 |
Interpretation. A, B, C and the AB interaction are significant; AC is marginal at \(5\%\) and BC and ABC are not. The pattern — large main effects, one two-factor interaction, negligible three-factor interaction — is the usual one, and it is what makes the fractional designs of Unit 3 worth having: if three-factor interactions are reliably negligible, the runs spent estimating them are wasted.
Read the AB interaction from the data rather than the table. The effect of A at low B is \(\left[(a - (1)) + (ac - c)\right]/(2r) = \left[(32-20)+(38-24)\right]/4 = 26/4 = 6.5\); at high B it is \(\left[(ab - b)+(abc - bc)\right]/4 = \left[(44-26)+(52-30)\right]/4 = 40/4 = 10\). The difference is \(3.5\), and half of it, \(1.75\), is exactly the AB effect computed above. That is what an interaction effect is: half the change in one factor's effect when the other factor moves.
| \(k\) | Runs per replicate | Main effects | 2-factor | 3-factor | 4-factor | Total effects |
|---|---|---|---|---|---|---|
| 2 | 4 | 2 | 1 | — | — | 3 |
| 3 | 8 | 3 | 3 | 1 | — | 7 |
| 4 | 16 | 4 | 6 | 4 | 1 | 15 |
The count is \(\binom{k}{j}\) effects with \(j\) letters, and they sum to \(2^{k} - 1\), which is the treatment degrees of freedom — as it must be. Every entry in the table is one degree of freedom, and every one is estimated by a contrast of all \(r2^{k}\) observations, so all effects are estimated with the same precision:
\[ \operatorname{Var}(\text{any effect}) = \frac{\sigma^{2}}{r\,2^{k-2}}. \]Why. An effect is \(\left[\,\cdot\,\right]/(r2^{k-1})\) and the contrast is a sum of \(r2^{k}\) terms each with coefficient \(\pm 1\), so its variance is \(r2^{k}\sigma^{2}\); dividing by \((r2^{k-1})^{2}\) gives \(r2^{k}\sigma^{2}/(r^{2}2^{2k-2}) = \sigma^{2}/(r2^{k-2})\).
The arithmetic grows, the structure does not. At \(k = 4\) there are fifteen contrasts and Yates's algorithm takes four passes, but nothing conceptual is added. What does change is the balance of the design: eleven of the fifteen effects involve three or four factors, and those are usually the ones worth sacrificing — which is the subject of confounding and fractional replication in Unit 3.
Two levels can only fit a straight line between them. Three levels allow a curvature term, which matters whenever the response has an optimum in the range studied — a temperature that is too low and too high, a dose that is too small and too large. The cost is runs: \(3^{2} = 9\) against \(2^{2} = 4\).
With three levels each factor's \(2\) degrees of freedom split into a linear and a quadratic component, using the orthogonal polynomial coefficients for equally spaced levels:
\[ \text{linear } (-1, 0, +1), \qquad \text{quadratic } (+1, -2, +1), \]which are orthogonal to each other and to the mean. The interaction's \(4\) degrees of freedom then split into four single-degree components, obtained by multiplying the coefficients:
\[ A_LB_L, \qquad A_LB_Q, \qquad A_QB_L, \qquad A_QB_Q. \]For a contrast \(C = \sum_{ij} c_i d_j T_{ij}\) on cell totals \(T_{ij}\) over \(r\) replicates,
\[ SS = \frac{C^{2}}{r\left(\sum_i c_i^{2}\right)\left(\sum_j d_j^{2}\right)}. \]The two divisors are \(2\) for a linear coefficient set and \(6\) for a quadratic one.
Given. \(r = 2\) replicates, 18 observations. Cell totals:
| b₀ | b₁ | b₂ | \(A_i\) | |
|---|---|---|---|---|
| a₀ | 10 | 14 | 18 | 42 |
| a₁ | 16 | 22 | 26 | 64 |
| a₂ | 20 | 24 | 30 | 74 |
| \(B_j\) | 46 | 60 | 74 | 180 |
Each cell holds two observations differing by 2, so the within-cell sum of squares is \(2\) per cell and \(SS_E = 9 \times 2 = 18\) on \(9(r-1) = 9\) degrees of freedom, \(MS_E = 2\).
Step 1 — the ordinary analysis of variance. With \(\text{CF} = 180^{2}/18 = 1800\) and a raw sum of squares of \(1974\),
\[ SS_A = \frac{42^{2}+64^{2}+74^{2}}{6} - 1800 = \frac{11336}{6} - 1800 = 89.333333, \] \[ SS_B = \frac{46^{2}+60^{2}+74^{2}}{6} - 1800 = \frac{11192}{6} - 1800 = 65.333333, \] \[ SS_{\text{cells}} = \frac{3912}{2} - 1800 = 156, \qquad SS_{AB} = 156 - 89.333333 - 65.333333 = 1.333333, \] \[ SS_{\text{total}} = 1974 - 1800 = 174, \qquad SS_E = 174 - 156 = 18. \]Step 2 — split A. Apply the coefficients to the row totals \(42, 64, 74\):
\[ C_{A_L} = -42 + 0 + 74 = 32, \qquad SS_{A_L} = \frac{32^{2}}{2 \times 2 \times 3} = \frac{1024}{12} = 85.333333, \] \[ C_{A_Q} = 42 - 2(64) + 74 = -12, \qquad SS_{A_Q} = \frac{(-12)^{2}}{2 \times 6 \times 3} = \frac{144}{36} = 4. \]The divisors are \(r \sum c_i^{2} \times 3\), the last factor being the three levels of B that each row total spans. The two components add to \(89.333333 = SS_A\). \(\checkmark\)
Step 3 — split B. On the column totals \(46, 60, 74\):
\[ C_{B_L} = -46 + 74 = 28, \qquad SS_{B_L} = \frac{784}{12} = 65.333333, \] \[ C_{B_Q} = 46 - 120 + 74 = 0, \qquad SS_{B_Q} = 0. \]B is exactly linear: its three totals are equally spaced, \(46, 60, 74\), stepping by 14 each time.
Step 4 — split the interaction. Each component is \(\sum_{ij}c_id_jT_{ij}\) over the nine cell totals. For \(A_QB_Q\), for instance, the nine coefficients are the outer product of \((1,-2,1)\) with itself:
\[ C_{A_QB_Q} = 10 - 28 + 18 - 32 + 88 - 52 + 20 - 48 + 30 = 6, \] \[ SS_{A_QB_Q} = \frac{6^{2}}{2 \times 6 \times 6} = \frac{36}{72} = 0.5. \]| Component | Contrast | Divisor | SS | \(F\) (1, 9) | \(p\) |
|---|---|---|---|---|---|
| \(A_L\) | 32 | 12 | 85.333333 | 42.666667 | 0.000107 |
| \(A_Q\) | −12 | 36 | 4.000000 | 2.000000 | 0.190947 |
| \(B_L\) | 28 | 12 | 65.333333 | 32.666667 | 0.000289 |
| \(B_Q\) | 0 | 36 | 0.000000 | 0.000000 | 1.000000 |
| \(A_LB_L\) | 2 | 8 | 0.500000 | 0.250000 | 0.629071 |
| \(A_LB_Q\) | 2 | 24 | 0.166667 | 0.083333 | 0.779368 |
| \(A_QB_L\) | −2 | 24 | 0.166667 | 0.083333 | 0.779368 |
| \(A_QB_Q\) | 6 | 72 | 0.500000 | 0.250000 | 0.629071 |
Step 5 — the checks. The four interaction components sum to \(0.5 + 0.166667 + 0.166667 + 0.5 = 1.333333 = SS_{AB}\), and the eight components together sum to \(156 = SS_{\text{cells}}\). \(\checkmark\)
Interpretation. Almost everything is in \(A_L\) and \(B_L\): the response rises nearly linearly in both factors, with no useful curvature and no interaction. Had \(A_Q\) been large, the reading would have been that A has an interior optimum, and the natural next step would be the response surface methods of Unit 4 rather than more factorial runs.
The split also answers a question the \(2\)-degree-of-freedom \(F\) for A cannot. \(F_A = (89.333333/2)/2 = 22.333333\) is significant, but that test spreads its power over both components. Testing \(A_L\) alone gives \(F = 42.666667\) on one degree of freedom — nearly twice as sensitive, because the test is aimed at the alternative that actually holds.
How to read a two-factor interaction — as the failure of the lines to be parallel, and as the rule that main effects must not then be reported alone — is already covered, with its plot. Example 2.1 checked it arithmetically on this data set: the effect of A was \(6.5\) at low B and \(10\) at high B, and half the difference is the \(1.75\) the contrast gave.
Three factors bring something that two cannot. \(ABC\) measures how the AB interaction itself changes between the levels of C — half the difference between \(AB\) computed at high C and \(AB\) computed at low C. It is a difference of differences of differences, and that is exactly why it is hard to explain to anyone and easy to mistrust.
Two practical consequences, and the whole of the next unit rests on them.