Skip to the content

Topics Covered

Cluster Estimator Between-Cluster Mean Square Intra-Cluster Correlation Design Effect Optimum Cluster Size Unequal Cluster Sizes
On this page
  1. 1. Clusters of Equal Size
  2. 2. Choosing the Cluster Size under a Cost Constraint
  3. 3. Clusters of Unequal Size
  4. Key Take-aways
Where this unit starts. Cluster sampling is already described at Foundation level: Sampling Techniques, Unit 4, section 2 defines it, says when it is appropriate, sets it against stratified sampling in a table, and states that a cluster sample is generally less precise than a simple random sample of the same size but much cheaper. That description is not repeated here.

What that section does not contain is a single formula. This unit supplies them: the estimator and its exact variance, the intra-cluster correlation that explains why the loss of precision happens and predicts its size, the design effect, the choice of cluster size under a cost constraint, and the unequal-sizes case, where even the estimator has to be reconsidered.

1. Clusters of Equal Size

THE ESTIMATOR AND ITS VARIANCE

Let the population of \(NM\) units be divided into \(N\) clusters of \(M\) units each, and let \(n\) clusters be chosen by simple random sampling without replacement, every unit in a chosen cluster being measured. Write \(\bar y_i\) for the mean of cluster \(i\). The estimator of the population mean per unit is

\[ \bar y_{cl} = \frac{1}{n}\sum_{i \in s}\bar y_i, \]

simply the mean of the \(n\) cluster means. It is the sample mean of a simple random sample of \(n\) items from a population of \(N\) items — the items being cluster means — so its variance is immediate from the Foundation SRSWOR result:

\[ \operatorname{Var}\!\left(\bar y_{cl}\right) = \frac{N-n}{Nn}\,S_b^{2*}, \qquad S_b^{2*} = \frac{1}{N-1}\sum_{i=1}^{N}\left(\bar y_i - \bar{\bar y}\right)^{2}, \]

the variance among cluster means. It is convenient to work instead with the between-cluster mean square on a per-unit scale,

\[ S_b^{2} = \frac{M}{N-1}\sum_{i=1}^{N}\left(\bar y_i - \bar{\bar y}\right)^{2} = M\,S_b^{2*}, \qquad\text{so}\qquad \operatorname{Var}\!\left(\bar y_{cl}\right) = \frac{1-f}{nM}\,S_b^{2}, \quad f = \frac{n}{N}. \]

Written that way it invites the comparison that matters: a simple random sample of the same number of units, \(nM\), would have variance \(\frac{1-f}{nM}S^{2}\) with \(S^{2}\) the variance of the whole population. The two differ only in \(S_b^{2}\) against \(S^{2}\), and the ratio of those is the design effect.

THE INTRA-CLUSTER CORRELATION

Define \(\rho\) as the correlation between two units drawn from the same cluster:

\[ \rho = \frac{E\left(y_{ij} - \bar{\bar y}\right)\left(y_{ik} - \bar{\bar y}\right)} {E\left(y_{ij} - \bar{\bar y}\right)^{2}} = \frac{\sum_{i}\sum_{j \ne k}\left(y_{ij} - \bar{\bar y}\right) \left(y_{ik} - \bar{\bar y}\right)} {(M-1)(NM-1)S^{2}}, \qquad j \ne k. \]

The exact finite-population identity connecting it to the between-cluster mean square is

\[ S_b^{2} = S^{2}\left[1 + (M-1)\rho\right]\cdot\frac{NM-1}{M(N-1)}, \]

so that

\[ \boxed{\; \frac{\operatorname{Var}\!\left(\bar y_{cl}\right)} {\operatorname{Var}\!\left(\bar y_{SRS}\right)} = \left[1 + (M-1)\rho\right]\cdot\frac{NM-1}{M(N-1)}\;} \]

The familiar textbook design effect \(1 + (M-1)\rho\) is this with the second factor dropped, which is legitimate only when \(N\) is large: as \(N \to \infty\) the factor \((NM-1)/[M(N-1)] \to 1\). At the small \(N\) of a worked example it does not, and Example 3.1 shows the discrepancy rather than hiding it.

What the formula says. \(\rho\) measures how alike units within a cluster are. Villages, schools and households are positively internally correlated on almost every variable one would measure, so \(\rho > 0\), so the design effect exceeds 1, so cluster sampling loses precision. It is multiplied by \(M-1\): a cluster of 30 units with a modest \(\rho = 0.1\) carries a design effect of about \(3.9\), meaning the sample must be nearly four times as large for the same precision. Negative \(\rho\) — clusters deliberately made internally heterogeneous — would make cluster sampling better than simple random sampling, which is precisely the aim of stratification, and is why stratification and clustering are opposite design ideas.

EXAMPLE 3.1 — THE DESIGN EFFECT, COMPUTED TWO WAYS

Given. \(N = 4\) clusters of \(M = 3\) units, so \(NM = 12\); \(n = 2\) clusters are selected.

ClusterValues\(\bar y_i\)
14, 6, 55
210, 12, 1111
37, 9, 88
413, 15, 1414
Grand mean9.5

The clusters are deliberately tight internally — each spans only 2 units — and far apart from one another, which is the situation cluster sampling faces in practice.

Step 1 — the three mean squares.

\[ S_b^{2} = \frac{3}{3}\left[(5-9.5)^{2} + (11-9.5)^{2} + (8-9.5)^{2} + (14-9.5)^{2}\right], \] \[ = 20.25 + 2.25 + 2.25 + 20.25 = 45, \] \[ S_w^{2} = \frac{1}{N(M-1)}\sum_i\sum_j\left(y_{ij} - \bar y_i\right)^{2} = \frac{4 \times 2}{4 \times 2} = 1, \]

each cluster contributing \((-1)^{2} + 1^{2} + 0^{2} = 2\); and over all twelve units

\[ S^{2} = \frac{1}{11}\sum_{ij}\left(y_{ij} - 9.5\right)^{2} = 13. \]

Step 2 — the intra-cluster correlation. Summing the \(M(M-1) = 6\) cross-products within each cluster and dividing as the definition requires,

\[ \rho = \frac{131}{143} = 0.916084. \]

Step 3 — the design effect, from \(\rho\).

\[ 1 + (M-1)\rho = 1 + 2(0.916084) = 2.832168, \qquad \frac{NM-1}{M(N-1)} = \frac{11}{9} = 1.222222, \] \[ \left[1 + (M-1)\rho\right]\frac{NM-1}{M(N-1)} = 2.832168 \times 1.222222 = 3.461538. \]

Step 4 — the design effect, from the variances directly. With \(n = 2\), \(f = 1/2\),

\[ \operatorname{Var}\!\left(\bar y_{cl}\right) = \frac{1 - 0.5}{2 \times 3}\times 45 = \frac{0.5}{6}\times 45 = 3.75, \] \[ \operatorname{Var}\!\left(\bar y_{SRS}\right) = \left(1 - \frac{6}{12}\right)\frac{13}{6} = 0.5 \times 2.166667 = 1.083333, \] \[ \frac{3.75}{1.083333} = \frac{45}{13} = 3.461538. \checkmark \]

The two routes agree exactly, which is the check that the identity in section 1 is an identity and not an approximation.

Step 5 — read the size of the loss. The cluster sample of six units is as precise as a simple random sample of \(6/3.461538 = 1.73\) units. Note also what the large-\(N\) design effect alone would have claimed: \(2.832168\), understating the loss by \(18\%\), because \(N = 4\) is nowhere near large.

Interpretation. \(\rho = 0.92\) is extreme, and deliberately so: the within-cluster variance is \(1\) against a total variance of \(13\). Measuring three units in a cluster is close to measuring the same unit three times. The remedy is not to abandon clusters — they are cheap — but to take more clusters and fewer units in each, which is the calculation of section 2 and the two-stage design of Unit 4.

2. Choosing the Cluster Size under a Cost Constraint

THE TRADE-OFF, MADE ARITHMETIC

Precision argues for small clusters; cost argues for large ones. A simple and realistic cost function separates the two:

\[ C = c_1 n + c_2 nM, \]

where \(c_1\) is the cost of reaching a cluster — travel, permissions, listing — and \(c_2\) the cost of measuring one unit once there. The total number of units measured is \(nM\), so for a fixed budget \(C\),

\[ n = \frac{C}{c_1 + c_2M}, \qquad nM = \frac{CM}{c_1 + c_2M}. \]

The effective sample size is \(nM\) divided by the design effect, and that is what should be maximised. Substituting the large-\(N\) design effect,

\[ n_{\text{eff}} = \frac{nM}{1 + (M-1)\rho} = \frac{C\,M}{\left(c_1 + c_2M\right)\left[1 + (M-1)\rho\right]}. \]

The numerator grows linearly in \(M\) and the denominator grows quadratically, so \(n_{\text{eff}}\) rises, peaks and falls; differentiating and setting to zero gives

\[ M_{\text{opt}} = \sqrt{\frac{c_1}{c_2}\cdot\frac{1-\rho}{\rho}}. \]

Both factors read naturally. Expensive travel relative to measurement pushes towards larger clusters; high internal correlation pushes towards smaller ones, and as \(\rho \to 1\) the optimum goes to a single unit per cluster — exactly the conclusion Example 3.1 reached by inspection.

EXAMPLE 3.2 — WHAT A BUDGET BUYS AT EACH CLUSTER SIZE

Given. A budget of \(1000\), with \(c_1 = 40\) per cluster reached and \(c_2 = 5\) per unit measured.

\(M\)Cost per cluster \(c_1 + c_2M\)\(n = C/(c_1+c_2M)\) Units measured \(nM\)
14522.22222.22
25020.00040.00
35518.18254.55
46016.66766.67
56515.38576.92
67014.28685.71

Step 1 — note that raw sample size always favours bigger clusters. Measuring 86 units beats measuring 22, and if precision per unit were constant the answer would simply be "make \(M\) as large as possible". It is not constant, which is the whole point.

Step 2 — apply the optimum formula at two values of \(\rho\).

\[ \rho = 0.916084: \quad M_{\text{opt}} = \sqrt{\frac{40}{5}\cdot\frac{0.083916}{0.916084}} = \sqrt{8 \times 0.091603} = \sqrt{0.732824} = 0.856, \]

which is below 1, so the answer is \(M = 1\): take one unit from each of many clusters, and the design ceases to be cluster sampling at all.

\[ \rho = 0.1: \quad M_{\text{opt}} = \sqrt{8 \times \frac{0.9}{0.1}} = \sqrt{72} = 8.485, \]

so \(M = 8\) or \(9\).

Interpretation. The same budget and the same costs give opposite answers at the two correlations, and \(\rho\) is a property of the population that must be estimated from a pilot survey or from an earlier round. That is the practical lesson: a cluster size chosen without an estimate of \(\rho\) is chosen arbitrarily, and the cost of getting it wrong is measured in multiples of the sample size.

3. Clusters of Unequal Size

WHEN THE SIZES DIFFER, THE ESTIMATOR ITSELF IS IN QUESTION

Real clusters — villages, schools, households — are not of equal size. Let cluster \(i\) contain \(M_i\) units, with \(M_0 = \sum_i M_i\) and \(\bar M = M_0/N\). Three estimators of the mean per unit now present themselves, and they are genuinely different.

EstimatorFormUnbiased?Comment
Mean of cluster means \(\dfrac{1}{n}\sum_{i \in s}\bar y_i\) for the mean of the \(\bar y_i\), not for \(\bar{\bar Y}\) gives a small cluster the same weight as a large one
Ratio-to-size \(\dfrac{\sum_{i \in s} M_i\bar y_i}{\sum_{i \in s} M_i} = \dfrac{\sum_{i \in s} y_{i\cdot}}{\sum_{i \in s} M_i}\) biased, \(O(1/n)\) the natural estimator; a ratio estimator with \(M_i\) as auxiliary
Expansion \(\dfrac{N}{n\,M_0}\sum_{i \in s} y_{i\cdot}\) unbiased, if \(M_0\) is known higher variance when the \(M_i\) vary widely

The ratio-to-size estimator is the usual choice, and naming it a ratio estimator is not a coincidence — it is the estimator of Unit 2 with \(y\) the cluster total and \(x\) the cluster size, so everything proved there applies unchanged: the bias is \(O(1/n)\), the approximate variance is \(\frac{1-f}{n\bar M^{2}}\cdot\frac{1}{N-1}\sum_i M_i^{2} \left(\bar y_i - \bar{\bar Y}\right)^{2}\), and it is efficient exactly when the cluster total is nearly proportional to the cluster size — which is to say, when the cluster means do not vary much with size.

The better answer is usually to stop sampling clusters with equal probability. Selecting clusters with probability proportional to \(M_i\), as in Unit 1, and then taking a fixed number of units from each, makes every unit's inclusion probability equal, removes the size effect entirely and gives a self-weighting design. That combination — PPS at the first stage, a fixed sub-sample at the second — is the standard large-survey design, and it is the subject of Unit 4.

Key Take-aways