Suppose a \(2^{3}\) experiment must be run in blocks of four rather than eight. Some contrast has to be sacrificed: with two blocks of four there is one block contrast, and whatever treatment contrast has the same pattern of \(\pm\) signs cannot be told apart from it. Choosing which is the whole art, and the answer is always the same — sacrifice the highest-order interaction, because it is the effect least likely to matter.
The construction. To confound \(ABC\) in a \(2^{3}\), put into the principal block every treatment combination having an even number of letters in common with \(ABC\), and the rest into the other block:
\[ \text{Block 1: } (1),\ ab,\ ac,\ bc \qquad\qquad \text{Block 2: } a,\ b,\ c,\ abc. \]Check one: \(ab\) shares two letters with \(ABC\) — even — so it is in block 1; \(a\) shares one — odd — so it is in block 2. Every combination in block 1 carries a \(+\) sign in the \(ABC\) contrast and every one in block 2 a \(-\), which is exactly why the two cannot be separated.
Equivalently, modulo 2. Write each combination as an exponent vector \((x_1, x_2, x_3)\) with \(x_i \in \{0,1\}\). The principal block is the solutions of
\[ x_1 + x_2 + x_3 \equiv 0 \pmod 2, \]which is a subgroup of the \(2^{3}\) combinations under componentwise addition mod 2; the other block is its coset. That algebraic description is what generalises: to confound \(ABCD\) in a \(2^{4}\), take \(x_1+x_2+x_3+x_4 \equiv 0\); to make four blocks of four in a \(2^{4}\), impose two independent equations, and then a third effect — the generalised interaction, the symmetric difference of the two — is confounded automatically and must be checked before the design is adopted.
What is lost. The confounded effect has no unconfounded estimate at all: its sum of squares is indistinguishable from the block sum of squares, so it is simply not tested. Everything else is estimated exactly as in an unblocked design, with a smaller error mean square than an eight-unit block would have given.
Given. The \(2^{3}\) data of Example 2.1, now run as two replicates of two blocks of four, with \(ABC\) confounded in both replicates.
| Replicate | Block 1: (1), ab, ac, bc | Block 2: a, b, c, abc | Replicate total |
|---|---|---|---|
| I | 9 + 21 + 18 + 15 = 63 | 16 + 12 + 11 + 25 = 64 | 127 |
| II | 11 + 23 + 20 + 15 = 69 | 16 + 14 + 13 + 27 = 70 | 139 |
Step 1 — the sum of squares for blocks within replicates. There are four blocks and two replicates, so this carries \(2\) degrees of freedom:
\[ SS_{\text{blocks within reps}} = \frac{63^{2}+64^{2}+69^{2}+70^{2}}{4} - \frac{127^{2}+139^{2}}{8} = \frac{17726}{4} - \frac{35450}{8}, \] \[ = 4431.5 - 4431.25 = 0.25. \]Step 2 — compare with the \(ABC\) sum of squares. Example 2.1 gave \([ABC] = 2\) and
\[ SS_{ABC} = \frac{2^{2}}{16} = 0.25. \]The two are equal, and not by accident: the block contrast within a replicate is the \(ABC\) contrast, so their sums of squares are the same number computed twice.
Step 3 — the analysis of variance. \(ABC\) is not listed, because it has been absorbed into blocks.
| Source | SS | df |
|---|---|---|
| Replicates | 9.00 | 1 |
| Blocks within replicates (= \(ABC\)) | 0.25 | 2 |
| A, B, C, AB, AC, BC | 407.50 | 6 |
| Error | 3.00 | 6 |
| Total | 419.75 | 15 |
The six unconfounded effects total \(407.75 - 0.25 = 407.50\), and the error is \(419.75 - 9 - 0.25 - 407.50 = 3.00\) on \(6\) degrees of freedom, \(MS_E = 0.5\).
Interpretation. Compare with Example 2.1, where the same data analysed without blocking gave \(SS_E = 3\) on \(7\) degrees of freedom. Confounding has cost one degree of freedom of error and the whole of \(ABC\), and has bought blocks of four instead of eight. In a real experiment the blocks would be more homogeneous than a block of eight could be, so the error mean square would fall — which is the entire point, and is the one thing this constructed data set cannot demonstrate.
Total confounding loses an effect completely. If the experiment has several replicates anyway, a better arrangement is to confound a different effect in each — then every effect is estimable from the replicates in which it is not confounded.
For a \(2^{3}\) in four replicates of two blocks of four:
| Replicate | Effect confounded | Effects estimable in it |
|---|---|---|
| I | \(ABC\) | all except \(ABC\) |
| II | \(AB\) | all except \(AB\) |
| III | \(AC\) | all except \(AC\) |
| IV | \(BC\) | all except \(BC\) |
The three main effects are estimated from all four replicates and are therefore completely confounded with nothing. Each of the four interactions is estimated from the three replicates in which it is free, so its contrast is formed over \(r' = 3\) replicates and
\[ SS = \frac{\left[\,\cdot\,\right]^{2}}{r'\,2^{k}} \quad\text{with } r' = 3 \text{ in place of } r = 4. \]The relative information on a partially confounded effect is \(r'/r = 3/4 = 0.75\): its estimate has variance \(4/3\) times that of an unconfounded effect, because it rests on three quarters of the data. That is a far better bargain than losing an effect outright, and it is why partial confounding is preferred whenever the number of replicates allows it.
Degrees of freedom for the \(4 \times 8 = 32\) observations:
| Source | df |
|---|---|
| Replicates | 3 |
| Blocks within replicates | 4 |
| Treatment effects (7, each on 1 df) | 7 |
| Error | 17 |
| Total | 31 |
The four block-within-replicate degrees of freedom are one per replicate, and they carry the four partially confounded interactions — each appearing once, in the replicate where it was sacrificed. A balanced partial confounding is one in which every effect of a given order is confounded equally often, so that all of them carry the same relative information; the scheme above is balanced over the two-factor interactions.
Confounding reduces the block size but still runs all \(2^{k}\) combinations. Fractional replication runs only some of them. A one-half fraction of a \(2^{4}\) is eight runs: choose a defining contrast \(I = ABCD\) and take the eight combinations on one side of it.
What is lost is aliasing. Every effect is now indistinguishable from its product with the defining word. Multiplying formally and reducing squared letters to \(I\):
\[ A \cdot ABCD = A^{2}BCD = BCD, \]so \(A\) and \(BCD\) are aliases — one contrast estimates their sum, and there is no way to say which it came from. The full alias structure of the half fraction of \(2^{4}\) with \(I = ABCD\) is
| Estimates | Aliased with |
|---|---|
| \(A\) | \(BCD\) |
| \(B\) | \(ACD\) |
| \(C\) | \(ABD\) |
| \(D\) | \(ABC\) |
| \(AB\) | \(CD\) |
| \(AC\) | \(BD\) |
| \(AD\) | \(BC\) |
Seven contrasts from eight runs, as they must be.
Resolution is the length of the shortest word in the defining relation. Here the only word is \(ABCD\), so the design is of resolution IV, written \(2^{4-1}_{\text{IV}}\). The rule that makes resolution useful: in a resolution \(R\) design, an effect involving \(p\) factors is aliased only with effects involving at least \(R - p\) factors. At \(R = 4\): main effects \((p=1)\) are clear of everything up to three-factor interactions, but two-factor interactions are aliased with each other.
A one-quarter fraction needs two generators, and they bring a third word with them. Taking a \(2^{5}\) in eight runs with \(D = AB\) and \(E = AC\), the defining relation is
\[ I = ABD = ACE = BCDE, \]the last being the product of the first two, \(ABD \cdot ACE = A^{2}BCDE = BCDE\). The shortest word has three letters, so this is resolution III, and main effects are aliased with two-factor interactions:
| Estimates | Aliased with |
|---|---|
| \(A\) | \(BD\), \(CE\), \(ABCDE\) |
| \(B\) | \(AD\), \(CDE\), \(ABCE\) |
| \(C\) | \(AE\), \(BDE\), \(ABCD\) |
| \(D\) | \(AB\), \(BCE\), \(ACDE\) |
| \(E\) | \(AC\), \(BCD\), \(ABDE\) |
Five factors in eight runs is a bargain, and resolution III is the price. Such designs are used for screening — finding which of many factors matter at all — on the understanding that a follow-up experiment will separate the aliases of whatever turns out to be large.
Some factors are awkward to change. Irrigation method must be applied to a whole field; oven temperature to a whole batch; a teaching method to a whole class. The factor that must be applied to large units is the whole-plot factor; the one that can be applied within them is the sub-plot factor.
The design is then two experiments nested one inside the other. Whole plots are randomised within replicates and carry factor A; each whole plot is split into \(b\) sub-plots to which the levels of B are randomised separately.
The consequence is two error terms. A is compared between whole plots, so it is tested against the whole-plot error; B and \(A\times B\) are compared within whole plots, so they are tested against the sub-plot error, which is almost always the smaller of the two. The design therefore sacrifices precision on A to gain it on B and on the interaction — which is the right trade only when the sub-plot factor is the one of interest.
| Source | df, general | df, \(a=3\), \(b=4\), \(r=4\) | Tested against |
|---|---|---|---|
| Replicates | \(r-1\) | 3 | |
| A (whole plot) | \(a-1\) | 2 | whole-plot error |
| Whole-plot error | \((r-1)(a-1)\) | 6 | |
| B (sub plot) | \(b-1\) | 3 | sub-plot error |
| \(A \times B\) | \((a-1)(b-1)\) | 6 | sub-plot error |
| Sub-plot error | \(a(r-1)(b-1)\) | 27 | |
| Total | \(rab-1\) | 47 |
The degrees of freedom check: \(3+2+6+3+6+27 = 47 = 3 \times 4 \times 4 - 1\). \(\checkmark\)
The commonest mistake in practice is to analyse a split-plot experiment as if it were a two-factor factorial in randomised blocks, pooling the two errors. That inflates the precision claimed for A and understates it for B, in a direction that flatters the hard-to-change factor. If the randomisation was restricted, the analysis must say so.
When there are \(v\) treatments but blocks hold only \(k < v\) units, a block cannot contain every treatment. A balanced incomplete block design restores as much symmetry as possible: each treatment appears \(r\) times, each block holds \(k\) units, and every pair of treatments appears together in exactly \(\lambda\) blocks. Writing \(b\) for the number of blocks and \(N = bk = vr\) for the number of units,
\[ \text{(i) } vr = bk, \qquad \text{(ii) } \lambda(v-1) = r(k-1), \qquad \text{(iii) } b \ge v \;\text{ (Fisher's inequality)}. \]Why (i). Count the units two ways — by treatment and by block. Why (ii). Fix a treatment. It appears in \(r\) blocks, each containing \(k-1\) other units, so it meets \(r(k-1)\) other units in all; and it meets each of the other \(v-1\) treatments exactly \(\lambda\) times. Why (iii) is deeper and is not proved here: it says an incomplete design cannot have fewer blocks than treatments.
Because \(\lambda = r(k-1)/(v-1)\) must be a whole number, the parameters cannot be chosen freely; and the relations are necessary, not sufficient, so a parameter set satisfying them need not correspond to an actual design.
A treatment total is not comparable with another, because the two appear in different blocks with different block effects. The repair is the adjusted treatment total
\[ Q_i = T_i - \frac{1}{k}\sum_{j \,:\, i \in j} B_j, \]which subtracts from \(T_i\) the average level of the blocks that treatment \(i\) happened to appear in. The \(Q_i\) always sum to zero, which is the first arithmetic check. Then
\[ SS_{\text{treatments (adjusted)}} = \frac{k\sum_i Q_i^{2}}{\lambda v}, \qquad SS_{\text{blocks (unadjusted)}} = \frac{\sum_j B_j^{2}}{k} - \text{CF}, \] \[ SS_E = SS_{\text{total}} - SS_{\text{blocks (unadj)}} - SS_{\text{treatments (adj)}}, \]on \(N - b - v + 1\) degrees of freedom. The adjusted treatment means and the standard error of a difference are
\[ \hat\mu_i = \bar y_{\cdot\cdot} + \frac{k\,Q_i}{\lambda v}, \qquad SE\left(\hat\mu_i - \hat\mu_{i'}\right) = \sqrt{\frac{2k\,MS_E}{\lambda v}}. \]The efficiency factor compares the design with a randomised block design of the same size:
\[ E = \frac{\lambda v}{rk} = \frac{v(k-1)}{k(v-1)} < 1. \]It is always less than 1 — incompleteness costs something — but in practice the smaller blocks reduce \(\sigma^{2}\) by more than \(E\) costs, which is why the design is used at all.
Recovery of inter-block information. The intra-block analysis treats block effects as fixed and throws away the differences between block totals. If blocks are regarded as random, those differences carry a second, independent estimate of each treatment contrast. Combining the two with weights inversely proportional to their variances gives the combined estimate, which is worth having when \(b\) is large relative to \(v\) and the block variance is not much bigger than the unit variance. When block differences are large, the inter-block estimate is nearly worthless and the combined estimate is essentially the intra-block one.
Given. Four treatments in six blocks of two — every pair of treatments together exactly once.
| Block | Contents | Block total \(B_j\) |
|---|---|---|
| 1 | A = 10, B = 12 | 22 |
| 2 | A = 11, C = 15 | 26 |
| 3 | A = 9, D = 18 | 27 |
| 4 | B = 13, C = 16 | 29 |
| 5 | B = 11, D = 17 | 28 |
| 6 | C = 14, D = 20 | 34 |
Step 1 — check the parameters.
\[ vr = 4 \times 3 = 12 = 6 \times 2 = bk, \qquad \lambda(v-1) = 1 \times 3 = 3 = 3 \times 1 = r(k-1), \qquad b = 6 \ge 4 = v. \checkmark \]Step 2 — treatment totals.
\[ T_A = 10+11+9 = 30, \quad T_B = 12+13+11 = 36, \quad T_C = 15+16+14 = 45, \quad T_D = 18+17+20 = 55, \]summing to \(G = 166\) over \(N = 12\) units, so \(\text{CF} = 166^{2}/12 = 2296.333333\).
Step 3 — the adjusted totals. Treatment A appears in blocks 1, 2, 3, whose totals are \(22, 26, 27\):
\[ Q_A = 30 - \frac{22+26+27}{2} = 30 - \frac{75}{2} = 30 - 37.5 = -7.5, \] \[ Q_B = 36 - \frac{22+29+28}{2} = 36 - 39.5 = -3.5, \qquad Q_C = 45 - \frac{26+29+34}{2} = 45 - 44.5 = 0.5, \] \[ Q_D = 55 - \frac{27+28+34}{2} = 55 - 44.5 = 10.5. \]Check: \(-7.5 - 3.5 + 0.5 + 10.5 = 0\). \(\checkmark\)
Step 4 — the adjusted treatment sum of squares.
\[ SS_{\text{tr(adj)}} = \frac{k\sum Q_i^{2}}{\lambda v} = \frac{2\left(56.25 + 12.25 + 0.25 + 110.25\right)}{1 \times 4} = \frac{2 \times 179}{4} = 89.5. \]Step 5 — blocks and total. The raw sum of squares of the twelve observations is \(2426\), so
\[ SS_{\text{total}} = 2426 - 2296.333333 = 129.666667, \] \[ SS_{\text{blocks(unadj)}} = \frac{22^{2}+26^{2}+27^{2}+29^{2}+28^{2}+34^{2}}{2} - \text{CF} = 2335 - 2296.333333 = 38.666667. \]Step 6 — error by subtraction.
\[ SS_E = 129.666667 - 38.666667 - 89.5 = 1.5, \]on \(N - b - v + 1 = 12 - 6 - 4 + 1 = 3\) degrees of freedom, so \(MS_E = 0.5\).
| Source | SS | df | MS | \(F\) | \(p\) |
|---|---|---|---|---|---|
| Blocks (unadjusted) | 38.666667 | 5 | 7.733333 | ||
| Treatments (adjusted) | 89.500000 | 3 | 29.833333 | 59.666667 | 0.003575 |
| Error | 1.500000 | 3 | 0.500000 | ||
| Total | 129.666667 | 11 |
Step 7 — adjusted means. With \(\bar y_{\cdot\cdot} = 166/12 = 13.833333\) and \(k/(\lambda v) = 2/4 = 0.5\),
\[ \hat\mu_i = 13.833333 + 0.5\,Q_i: \qquad 10.083333,\;\; 12.083333,\;\; 14.083333,\;\; 19.083333 \]for A, B, C, D. The unadjusted means were \(10, 12, 15, 18.333333\) — C has been pulled down and D pushed up, because D happened to appear in the three heaviest blocks.
Step 8 — the standard error and the efficiency.
\[ SE\left(\hat\mu_i - \hat\mu_{i'}\right) = \sqrt{\frac{2 \times 2 \times 0.5}{4}} = \sqrt{0.5} = 0.707107, \qquad E = \frac{\lambda v}{rk} = \frac{4}{6} = 0.666667. \]Interpretation. The design recovers two thirds of the information a complete block design would have given, and in exchange the blocks hold two units instead of four. That is the trade in one line: whether it is worth making depends entirely on how much more homogeneous a block of two is than a block of four, which is a question about the material and not about the arithmetic.
Two Latin squares of side \(s\) are orthogonal if, superimposed, every one of the \(s^{2}\) ordered pairs of symbols occurs exactly once. A complete set has \(s-1\) squares, and exists whenever \(s\) is a prime or a prime power.
The construction. Take \(v = s^{2}\) treatments arranged in an \(s \times s\) array. Its \(s\) rows give \(s\) blocks of size \(s\); its \(s\) columns give \(s\) more; and each of the \(s-1\) mutually orthogonal Latin squares contributes \(s\) more, one block per symbol. Using \(t\) of the available square-based systems gives
\[ v = s^{2}, \qquad k = s, \qquad b = s(t+2), \qquad r = t+2, \qquad \lambda = 1, \]a resolvable design, because each of the \(t+2\) systems is a complete replicate. With every square used, \(t = s-1\) and \(r = s+1\), giving the complete design with \(b = s(s+1)\) and \(\lambda = 1\). With \(t < s - 1\) the result is a lattice design, which is not balanced but partially balanced — the subject of Unit 4.
Worked at \(s = 3\). Nine treatments in a \(3 \times 3\) array; rows, columns and the two orthogonal Latin squares of side 3 give \(b = 3 \times 4 = 12\) blocks of \(k = 3\), with \(r = 4\) and
\[ \lambda = \frac{r(k-1)}{v-1} = \frac{4 \times 2}{8} = 1. \checkmark \]Every pair of the nine treatments occurs together exactly once: two in the same row occur in that row-block and nowhere else, and any other pair is brought together by exactly one of the two Latin squares.