Skip to the content

Topics Covered

PPS Selection Lahiri's Method Ratio Estimator Regression Estimator Efficiency Comparison Separate vs Combined Cluster Sampling Two-Stage Sampling
On this page
  1. Section B — The Seven Practicals
  2. Practical 1 — PPS Sampling With and Without Replacement
  3. Practical 2 — Ratio Estimator in Simple Random Sampling, Compared with SRS
  4. Practical 3 — Separate and Combined Ratio Estimators
  5. Practical 4 — Regression Estimator, Compared with SRS and the Ratio Estimator
  6. Practical 5 — Separate and Combined Regression Estimators
  7. Practical 6 — Cluster Sampling with Equal Cluster Sizes
  8. Practical 7 — Sub-Sampling (Two-Stage) with Equal First-Stage Units
  9. How Marks Are Lost
  10. What the Practical Record Should Contain
About this course. STS-206 is a conventional practical in two sections: Section A, Design and Analysis of Experiments, which belongs to Design and Analysis of Experiments (STS-203), and Section B, Sampling Theory — the seven experiments below. Both sections are examined by hand, so every experiment is worked with full arithmetic and no software.
Before you start. Everything the BSc paper Sampling Techniques teaches is assumed and is not repeated: drawing a simple random sample with and without replacement, the sample mean and its variance, the finite population correction, stratified sampling and its allocations, and systematic sampling. Only the estimators that are new at this level are worked here.

Section B — The Seven Practicals

THE SEVEN, AND WHERE ELSE EACH IS WORKED

All seven are worked below, each set out as 1. Problem, 2. Aim, 3. Formula, 4. Calculation, 5. Result. Practicals 2 and 4 share one sample, and 3 and 5 share the same two strata, so the estimators can be compared directly.

#PracticalWorked hereAlso in
1PPS sampling with and without replacementPractical 1Unit 1, Examples 1.1, 1.2 and 1.3
2Ratio estimators in SRS, comparison with SRSPractical 2—
3Separate and combined ratio estimators, comparisonPractical 3Unit 2, Example 2.2
4Regression estimators in SRS, comparison with SRS and ratio estimatorsPractical 4—
5Separate and combined regression estimators, comparisonPractical 5—
6Cluster sampling with equal cluster sizesPractical 6Unit 3, Examples 3.1 and 3.2
7Sub-sampling (two-stage) with equal first-stage unitsPractical 7Unit 4, Example 4.1

Practical 1 — PPS Sampling With and Without Replacement

1. Problem

(a) Five units have sizes \(x_i = 12, 8, 20, 5, 15\). Draw a sample of three with probability proportional to size, with replacement, by the cumulative total method, using the random numbers \(37, 09, 52\); and show how Lahiri's method treats the pairs \((4, 9)\) and \((2, 6)\).

(b) A population of \(N = 5\) units has sizes \(X_i = 10, 20, 30, 25, 15\) and (for this demonstration only) known values \(Y_i = 12, 26, 33, 30, 19\). For a PPS sample of \(n = 2\) with replacement, find the variance of the Hansen–Hurwitz estimator of the total, compare it with equal-probability sampling, and estimate the total if the draws fall on units 2 and 3.

(c) \(N = 4\) units have \(p = (0.1, 0.2, 0.3, 0.4)\) and \(Y = (3, 6, 9, 12)\), total 30. Two units are drawn successively without replacement. Find the inclusion probabilities, the Horvitz–Thompson estimate from the sample \(\{2, 4\}\), and its Yates–Grundy variance.

2. Aim

To draw samples with probability proportional to size, with and without replacement, and to estimate a population total with the estimator that belongs to each draw.

3. Formula

\[ p_i = \frac{x_i}{X}, \qquad \hat Y_{HH} = \frac{1}{n}\sum \frac{y_i}{p_i}, \quad \operatorname{Var}(\hat Y_{HH}) = \frac{1}{n}\sum_i p_i\left(\frac{Y_i}{p_i} - Y\right)^{2}, \quad v(\hat Y_{HH}) = \frac{1}{n(n-1)}\sum\left(\frac{y_i}{p_i} - \hat Y_{HH}\right)^{2} \] \[ \hat Y_{HT} = \sum \frac{y_i}{\pi_i}, \qquad \operatorname{Var}(\hat Y_{HT}) = \sum_{i<j}\left(\pi_i\pi_j - \pi_{ij}\right)\left(\frac{Y_i}{\pi_i} - \frac{Y_j}{\pi_j}\right)^{2} \]

Which estimator goes with which draw. This is the part most often got wrong in the examination, so state it in the first line of the answer: Hansen–Hurwitz for PPS with replacement (it uses \(p_i\)); Horvitz–Thompson for PPS without replacement (it uses the inclusion probabilities \(\pi_i\), with the Yates–Grundy variance; Unit 1, section 4).

Applying it:

  1. Cumulative total method: form the cumulative totals; draw a random number in \(1, \dots, X\); the unit whose range contains it is selected. With replacement the ranges are unchanged between draws; without replacement the selected unit is struck out and the cumulative column is rebuilt.
  2. Lahiri's method: draw a unit number from \(1, \dots, N\) and an independent number from \(1, \dots, M\), \(M = \max x_i\); accept the unit if the second number does not exceed its size, otherwise discard the pair and draw again.
  3. Estimate with the estimator that belongs to the draw, and check the identities: \(\sum\pi_i = n\), \(\sum_{i<j}\pi_{ij} = n(n-1)/2\).

4. Calculation

(a) \(X = 60\):

Unit\(x_i\)CumulativeRange\(p_i = x_i/X\)
112121–120.200000
282013–200.133333
3204021–400.333333
454541–450.083333
5156046–600.250000

37 falls in 21–40, 09 in 1–12 and 52 in 46–60: the sample is units \(3, 1, 5\). By Lahiri's method, \(M = 20\): the pair \((4, 9)\) is rejected because \(9 > x_4 = 5\), and the pair \((2, 6)\) is accepted because \(6 \le x_2 = 8\).

(b) \(p_i = X_i/100 = 0.10, 0.20, 0.30, 0.25, 0.15\), and \(Y_i/p_i = 120, 130, 110, 120, 126.666667\), whose departures from \(Y = 120\) are \(0, 10, -10, 0, 6.666667\). So

\[ \sum_i p_i\left(\frac{Y_i}{p_i} - Y\right)^{2} = 0.20(100) + 0.30(100) + 0.15(44.444444) = 56.666667, \qquad \operatorname{Var}(\hat Y_{HH}) = \frac{1}{2}\cdot\frac{170}{3} = \frac{85}{3} = 28.333333. \]

Under simple random sampling with replacement, \(\bar Y = 24\), \(\sigma^{2} = (144 + 4 + 81 + 36 + 25)/5 = 58\), and \(\operatorname{Var}(N\bar y) = 25 \times 58/2 = 725\); the ratio is \(725/(85/3) = 25.588235\). From units 2 and 3, \(\hat Y_{HH} = \tfrac12(130 + 110) = 120\).

(c) For unit 1,

\[ \pi_1 = 0.1 + 0.2\!\left(\frac{0.1}{0.8}\right) + 0.3\!\left(\frac{0.1}{0.7}\right) + 0.4\!\left(\frac{0.1}{0.6}\right) = \frac{197}{840} = 0.234524, \]

and likewise \(\pi_2 = 139/315 = 0.441270\), \(\pi_3 = 73/120 = 0.608333\), \(\pi_4 = 451/630 = 0.715873\); their sum is \(2 = n\). \(\checkmark\) The pairwise probabilities are \(\pi_{12} = 17/360\), \(\pi_{13} = 8/105\), \(\pi_{14} = 1/9\), \(\pi_{23} = 9/56\), \(\pi_{24} = 7/30\), \(\pi_{34} = 13/35\), summing to \(1 = n(n-1)/2\). \(\checkmark\)

\[ \hat Y_{HT} = \frac{6}{0.441270} + \frac{12}{0.715873} = 13.597122 + 16.762749 = 30.3599, \]

and with \(Y_i/\pi_i = 12.791878, 13.597122, 14.794521, 16.762749\) the Yates–Grundy variance is \(2.428334\); computing it directly over the six possible samples gives the same value, with \(E(\hat Y_{HT}) = 30\) exactly.

5. Result

(a) The PPS sample with replacement is units 3, 1, 5. (b) PPS sampling makes the estimator of the total 25.6 times as precise as equal-probability sampling, because \(Y_i\) is nearly proportional to \(X_i\); if \(Y_i\) were inversely related to \(X_i\), PPS would be worse. (c) \(\hat Y_{HT} = 30.36\) against the true 30, with standard error \(\sqrt{2.428334} = 1.558\), 5.2% of the total. Note that \(\pi_1 = 0.2345\) is not \(np_1 = 0.2\): successive sampling does not give inclusion probabilities exactly proportional to size (worked also in Unit 1, Examples 1.1, 1.2 and 1.3).

Practical 2 — Ratio Estimator in Simple Random Sampling, Compared with SRS

1. Problem

A population has \(N = 50\) plots. The auxiliary variable \(x\) is last season's yield, known for every plot with \(\bar X = 12\); \(y\) is this season's yield, measured on a simple random sample of \(n = 8\) plots without replacement.

Plot12345678Total
\(x\)8101113141591292
\(y\)2026273435392330234

Estimate \(\bar Y\) by the ratio estimator, find its estimated variance, and compare it with the sample mean.

2. Aim

To estimate a population mean by the ratio estimator, which uses a known auxiliary mean, and to measure its gain in precision over the sample mean.

3. Formula

\[ \hat{\bar Y} = \bar y, \quad v(\bar y) = \frac{1-f}{n}s_y^{2}; \qquad \hat R = \frac{\bar y}{\bar x}, \quad \hat{\bar Y}_R = \hat R\bar X, \quad v(\hat{\bar Y}_R) = \frac{1-f}{n}\cdot\frac{1}{n-1}\sum\left(y_i - \hat R x_i\right)^{2} \]

The ratio estimator fits one constant, the slope through the origin, so \(n - 1\) residual degrees of freedom remain. It is worth using when \(\rho > C_x/(2C_y)\), \(C\) being a coefficient of variation.

Applying it:

  1. Compute \(\bar x, \bar y, f\), the sums of squares and products, and \(r\).
  2. Mean per unit: \(\bar y\) and \(v(\bar y)\).
  3. Ratio: \(\hat R\) (keep it as a fraction while multiplying), \(\hat{\bar Y}_R\), the residuals \(e_i = y_i - \hat Rx_i\), and the variance.
  4. Efficiency \(= v(\bar y)/v(\hat{\bar Y}_R)\); check the condition \(\rho > C_x/(2C_y)\).

4. Calculation

\[ \bar x = \frac{92}{8} = 11.5, \quad \bar y = \frac{234}{8} = 29.25, \quad f = \frac{8}{50} = 0.16, \quad 1 - f = 0.84, \] \[ S_{xx} = 42, \quad S_{yy} = 291.5, \quad S_{xy} = 110; \qquad s_x^{2} = 6.000000, \quad s_y^{2} = 41.642857, \] \[ r^{2} = \frac{110^{2}}{42 \times 291.5} = \frac{12100}{12243} = 0.988320, \qquad r = 0.994143. \]

Mean per unit. \(v(\bar y) = \frac{0.84}{8}\times 41.642857 = 4.372500\), s.e. \(= 2.091052\).

Ratio. \(\hat R = 29.25/11.5 = 117/46 = 2.543478\), and \(\hat{\bar Y}_R = \frac{117}{46}\times 12 = \frac{702}{23} = 30.521739\) (rounding \(\hat R\) first gives 30.521736, wrong in the last digit). The residuals are

\(i\)12345678
\(e_i\)−0.3478260.565217−0.9782610.934783−0.6086960.8478260.108696−0.521739
\[ \sum e_i^{2} = 3.644612, \quad s_e^{2} = \frac{3.644612}{7} = 0.520659, \quad v(\hat{\bar Y}_R) = \frac{0.84}{8}\times 0.520659 = 0.054669, \quad \text{s.e.} = 0.233814. \]

Efficiency \(= 4.372500/0.054669 = 80.0\). \(C_x = \sqrt{6}/11.5 = 0.212999\), \(C_y = \sqrt{41.642857}/29.25 = 0.220620\), so \(C_x/(2C_y) = 0.4827\).

5. Result

\(\hat{\bar Y}_R = 30.52\) with standard error 0.234, against \(\bar y = 29.25\) with 2.091: the ratio estimator is 80 times as efficient as the sample mean here, because \(r = 0.994\) is far above \(C_x/(2C_y) = 0.48\). Report the estimate with its variance and say which formula produced it.

Practical 3 — Separate and Combined Ratio Estimators

1. Problem

A population has two strata, \(N_1 = 100\) and \(N_2 = 200\), with known stratum means of the auxiliary variable \(\bar X_1 = 5.5\) and \(\bar X_2 = 17\). A stratified sample gives:

Stratum\(n_h\)sample \(x\)sample \(y\)\(\bar x_h\)\(\bar y_h\)
144, 6, 8, 210, 14, 18, 6512
2512, 16, 20, 14, 1820, 26, 32, 23, 291626

Estimate \(\bar Y\) by the separate and by the combined ratio estimator, and compare them.

2. Aim

To estimate a population mean from a stratified sample by the separate and the combined ratio estimators, and to choose between them.

3. Formula

\[ \hat{\bar Y}_{RS} = \sum_h W_h\frac{\bar y_h}{\bar x_h}\bar X_h, \qquad \hat{\bar Y}_{RC} = \frac{\bar y_{st}}{\bar x_{st}}\bar X, \qquad \bar y_{st} = \sum_h W_h\bar y_h, \quad \bar x_{st} = \sum_h W_h\bar x_h \]

The separate estimator allows the ratio to differ between strata, but carries a bias of order \(\sum W_h/n_h\); the combined one has a bias of order \(1/n\) but is wrong in every stratum if the ratios genuinely differ (Unit 2, section 3).

Applying it:

  1. Compute \(W_h\) and \(\bar X = \sum W_h\bar X_h\).
  2. Separate: one ratio per stratum, applied to its own known mean, then weighted.
  3. Combined: the stratified means, one ratio, applied to \(\bar X\).
  4. Compare; choose by the stratum sample sizes.

4. Calculation

\[ W_1 = \tfrac13, \quad W_2 = \tfrac23, \qquad \bar X = \tfrac13(5.5) + \tfrac23(17) = \frac{79}{6} = 13.166667. \] \[ \hat R_1 = \frac{12}{5} = 2.4, \quad \hat R_2 = \frac{26}{16} = 1.625; \qquad \hat{\bar Y}_{RS} = \tfrac13(2.4 \times 5.5) + \tfrac23(1.625 \times 17) = \tfrac13(13.2) + \tfrac23(27.625) = \frac{1369}{60} = 22.816667. \] \[ \bar y_{st} = \frac{64}{3} = 21.333333, \quad \bar x_{st} = \frac{37}{3} = 12.333333, \quad \hat R_c = \frac{64}{37} = 1.729730, \quad \hat{\bar Y}_{RC} = \frac{64}{37}\times\frac{79}{6} = \frac{2528}{111} = 22.774775. \] \[ 22.816667 - 22.774775 = 0.041892. \]

5. Result

\(\hat{\bar Y}_{RS} = 22.817\) and \(\hat{\bar Y}_{RC} = 22.775\): they differ by about 0.2%, and they differ because the stratum ratios are not equal (2.4 against 1.625). The separate estimator is fitting something real, but with only four and five units per stratum its two \(O(1/n_h)\) biases are the larger worry: on this sample the combined estimator is the safer report, and the choice would reverse if each stratum held fifty units (worked also in Unit 2, Example 2.2).

Practical 4 — Regression Estimator, Compared with SRS and the Ratio Estimator

1. Problem

For the sample of Practical 2 (\(N = 50\), \(n = 8\), \(\bar X = 12\)):

Plot12345678Total
\(x\)8101113141591292
\(y\)2026273435392330234

estimate \(\bar Y\) by the linear regression estimator, find its estimated variance, and compare it with the sample mean and the ratio estimator.

2. Aim

To estimate a population mean by the regression estimator, and to compare its precision with the sample mean and the ratio estimator on the same sample.

3. Formula

\[ b = \frac{S_{xy}}{S_{xx}}, \qquad \hat{\bar Y}_{lr} = \bar y + b\left(\bar X - \bar x\right), \qquad v(\hat{\bar Y}_{lr}) = \frac{1-f}{n}\cdot\frac{1}{n-2}\left(S_{yy} - bS_{xy}\right) \]

The regression estimator fits an intercept as well as a slope, so only \(n - 2\) residual degrees of freedom remain; the residual sum of squares is \(S_{yy} - bS_{xy}\), without listing the residuals. Its large-sample variance is \(\frac{1-f}{n}s_y^{2}(1 - r^{2})\).

Applying it:

  1. Compute \(b\) and \(\bar X - \bar x\); form the estimate.
  2. Residual sum of squares by the identity; divide by \(n - 2\); the variance.
  3. Tabulate the three estimators with their variances and efficiencies over SRS.

4. Calculation

\[ b = \frac{110}{42} = \frac{55}{21} = 2.619048, \quad \bar X - \bar x = 0.5, \quad \hat{\bar Y}_{lr} = 29.250000 + 2.619048 \times 0.5 = 30.559524. \] \[ S_{yy} - bS_{xy} = 291.5 - 288.095238 = 3.404762, \quad s_{y\cdot x}^{2} = \frac{3.404762}{6} = 0.567460, \quad v(\hat{\bar Y}_{lr}) = \frac{0.84}{8}\times 0.567460 = 0.059583, \quad \text{s.e.} = 0.244097. \]
EstimatorEstimate of \(\bar Y\)Estimated variances.e.Efficiency over SRS
Mean per unit29.2500004.3725002.0910521.00
Ratio (Practical 2)30.5217390.0546690.23381480.0
Regression30.5595240.0595830.24409773.4

Each efficiency is \(v(\bar y)\) divided by the variance in its row; one decimal place is all the six-figure variances support. The fitted line is \(\hat y = -0.869048 + 2.619048x\). With the large-sample formula, \(0.105 \times 41.642857 \times 0.011680 = 0.051071\).

5. Result

\(\hat{\bar Y}_{lr} = 30.56\), s.e. 0.244: 73 times as efficient as the sample mean. Both auxiliary estimators cut the standard error by a factor near nine, because \(x\) and \(y\) are almost collinear — and here the ratio estimator wins, which surprises students told that the regression estimator is never worse. The fitted intercept is very nearly zero, so the line through the origin is almost the same line; and the regression estimator still pays for its intercept: its residual sum of squares, 3.404762, is smaller than the ratio's 3.644612, but it is divided by 6 instead of 7. "At least as efficient" is a statement about large-sample variances (0.051071 here); at \(n = 8\) the degree of freedom spent on the intercept costs more than the intercept is worth.

Practical 5 — Separate and Combined Regression Estimators

1. Problem

For the two strata of Practical 3 (\(N_1 = 100\), \(N_2 = 200\), \(\bar X_1 = 5.5\), \(\bar X_2 = 17\)), a sample with the same stratum means but with the values scattered about the line is:

Stratum\(n_h\)sample \(x\)sample \(y\)\(\bar x_h\)\(\bar y_h\)
144, 6, 8, 211, 13, 19, 5512
2512, 16, 20, 14, 1821, 26, 33, 22, 281626

Estimate \(\bar Y\) by the separate and by the combined regression estimator, and compare them with the ratio estimators.

2. Aim

To estimate a population mean from a stratified sample by the separate and the combined regression estimators, and to choose between them.

3. Formula

\[ \hat{\bar Y}_{lrs} = \sum_{h=1}^{L} W_h\left[\bar y_h + b_h\left(\bar X_h - \bar x_h\right)\right], \qquad \hat{\bar Y}_{lrc} = \bar y_{st} + b_c\left(\bar X - \bar x_{st}\right), \qquad b_c = \frac{\sum_h W_h^{2}\,\frac{S_{xy,h}}{n_h(n_h-1)}}{\sum_h W_h^{2}\,\frac{S_{xx,h}}{n_h(n_h-1)}} \]

\(b_c\) is the variance-weighted pooled slope, ignoring finite population corrections, which are negligible here.

Applying it:

  1. The within-stratum slopes \(b_h = S_{xy,h}/S_{xx,h}\).
  2. The separate estimator.
  3. The stratified means, the pooled slope, the combined estimator (form products from fractions).
  4. Tabulate with the ratio estimators.

4. Calculation

Stratum 1: \(x - \bar x = -1, 1, 3, -3\), \(y - \bar y = -1, 1, 7, -7\); \(S_{xx,1} = 20\), \(S_{xy,1} = 44\), \(b_1 = 2.2\). Stratum 2: \(x - \bar x = -4, 0, 4, -2, 2\), \(y - \bar y = -5, 0, 7, -4, 2\); \(S_{xx,2} = 40\), \(S_{xy,2} = 60\), \(b_2 = 1.5\).

\[ \hat{\bar Y}_{lrs} = \tfrac13(12 + 2.2 \times 0.5) + \tfrac23(26 + 1.5 \times 1) = \tfrac13(13.1) + \tfrac23(27.5) = 4.366667 + 18.333333 = 22.700000. \] \[ \bar y_{st} = \frac{64}{3}, \quad \bar x_{st} = \frac{37}{3}; \qquad b_c = \frac{\left(\tfrac13\right)^{2}\frac{44}{12} + \left(\tfrac23\right)^{2}\frac{60}{20}}{\left(\tfrac13\right)^{2}\frac{20}{12} + \left(\tfrac23\right)^{2}\frac{40}{20}} = \frac{47/27}{29/27} = \frac{47}{29} = 1.620690. \] \[ \bar X - \bar x_{st} = \frac{79}{6} - \frac{37}{3} = \frac56, \qquad \hat{\bar Y}_{lrc} = 21.333333 + \frac{47}{29}\times\frac56 = 21.333333 + 1.350575 = 22.683908. \]
EstimatorEstimate of \(\bar Y\)
Stratified mean, \(\bar y_{st}\) (no auxiliary information)21.333333
Separate ratio (Practical 3)22.816667
Combined ratio (Practical 3)22.774775
Separate regression22.700000
Combined regression22.683908

5. Result

The separate and combined regression estimates differ by \(22.700000 - 22.683908 = 0.016092\), about 0.07%, even though the slopes 2.2 and 1.5 are not close: the choice moves the estimate very little and the bias a great deal, so it is decided by \(n_h\). With four and five observations per stratum, the combined estimator is the defensible one here. All four auxiliary estimators sit near 22.7, while the stratified mean alone gives 21.33: the gain comes from using \(x\) at all; the choice among the four ways of using it is second order.

Practical 6 — Cluster Sampling with Equal Cluster Sizes

1. Problem

A population has \(N = 4\) clusters of \(M = 3\) units; \(n = 2\) clusters are to be selected by simple random sampling.

ClusterValues\(\bar y_i\)
14, 6, 55
210, 12, 1111
37, 9, 88
413, 15, 1414

Find the intra-cluster correlation, the variance of the cluster-sample mean, and its efficiency against a simple random sample of the same number of units. Then, for a budget of 1000 with \(c_1 = 40\) per cluster and \(c_2 = 5\) per unit, find the best cluster size at \(\rho = 0.916\) and at \(\rho = 0.1\).

2. Aim

To estimate a population mean by cluster sampling, measure the loss of precision through the intra-cluster correlation, and choose the cluster size for a budget.

3. Formula

\[ S_b^{2} = \frac{M}{N-1}\sum_i(\bar y_i - \bar Y)^{2}, \qquad \operatorname{Var}(\bar y_{cl}) = \frac{1-f}{nM}S_b^{2}, \qquad \operatorname{Var}(\bar y_{SRS}) = \left(1 - \frac{nM}{NM}\right)\frac{S^{2}}{nM} \] \[ \frac{\operatorname{Var}(\bar y_{cl})}{\operatorname{Var}(\bar y_{SRS})} = \left[1 + (M-1)\rho\right]\frac{NM-1}{M(N-1)}, \qquad M_{\text{opt}} = \sqrt{\frac{c_1}{c_2}\cdot\frac{1-\rho}{\rho}} \]

Applying it:

  1. Compute \(S_b^{2}\), the within-cluster mean square \(S_w^{2}\), and the overall \(S^{2}\).
  2. Compute \(\rho\) from the within-cluster cross-products.
  3. The design effect, from \(\rho\) and from the two variances: they must agree.
  4. For the budget, the optimum cluster size at each \(\rho\).

4. Calculation

The grand mean is 9.5.

\[ S_b^{2} = \frac{3}{3}\left[(5-9.5)^{2} + (11-9.5)^{2} + (8-9.5)^{2} + (14-9.5)^{2}\right] = 20.25 + 2.25 + 2.25 + 20.25 = 45, \] \[ S_w^{2} = \frac{4 \times 2}{4 \times 2} = 1, \qquad S^{2} = \frac{1}{11}\sum_{ij}(y_{ij} - 9.5)^{2} = 13, \qquad \rho = \frac{131}{143} = 0.916084. \] \[ \left[1 + 2(0.916084)\right]\times\frac{11}{9} = 2.832168 \times 1.222222 = 3.461538. \] \[ \operatorname{Var}(\bar y_{cl}) = \frac{0.5}{6}\times 45 = 3.75, \qquad \operatorname{Var}(\bar y_{SRS}) = 0.5 \times \frac{13}{6} = 1.083333, \qquad \frac{3.75}{1.083333} = \frac{45}{13} = 3.461538. \checkmark \] \[ \rho = 0.916084: \ M_{\text{opt}} = \sqrt{8 \times 0.091603} = 0.856; \qquad \rho = 0.1: \ M_{\text{opt}} = \sqrt{8 \times 9} = 8.485. \]

5. Result

\(\rho = 0.92\) and the design effect is 3.46: the cluster sample of six units is as precise as a simple random sample of \(6/3.46 = 1.73\) units. (The large-\(N\) design effect alone, 2.83, would understate the loss by 18%, because \(N = 4\) is nowhere near large.) Measuring three units in a cluster is close to measuring the same unit three times. For the budget, \(M_{\text{opt}} < 1\) at \(\rho = 0.916\), so take one unit from each of many clusters; at \(\rho = 0.1\), \(M = 8\) or 9. A cluster size chosen without an estimate of \(\rho\) is chosen arbitrarily (worked also in Unit 3, Examples 3.1 and 3.2).

Practical 7 — Sub-Sampling (Two-Stage) with Equal First-Stage Units

1. Problem

The population of Practical 6 (\(N = 4\) first-stage units of \(M = 3\)) is sampled in two stages: \(n = 2\) first-stage units, and \(m = 2\) units from each. Find the variance of the sample mean, compare it with measuring the whole first-stage unit, and find the optimum \(m\) for \(c_1 = 30\) per first-stage unit and \(c_2 = 10\) per unit. Repeat the optimum for \(S_1^{2} = 10\), \(S_2^{2} = 20\), \(M = 3\), \(c_1 = 20\), \(c_2 = 30\).

2. Aim

To find the variance of the two-stage sample mean and the number of second-stage units that minimises it for a given cost.

3. Formula

\[ \operatorname{Var}(\bar y) = \left(\frac1n - \frac1N\right)S_1^{2} + \left(\frac{1}{nm} - \frac{1}{nM}\right)S_2^{2}, \qquad m_{\text{opt}} = \sqrt{\frac{c_1}{c_2}\cdot\frac{S_2^{2}}{S_1^{2} - S_2^{2}/M}} \]

\(S_1^{2}\) is the mean square among first-stage unit means; the per-unit between-cluster mean square of Practical 6 is \(M\) times it (Unit 4, section 1).

Applying it:

  1. Compute \(S_1^{2}\) from the first-stage means; \(S_2^{2}\) is the pooled within mean square.
  2. Compute the two terms of the variance.
  3. Set \(m = M\) as a check against Practical 6.
  4. Compute \(m_{\text{opt}}\).

4. Calculation

\[ S_1^{2} = \frac{20.25 + 2.25 + 2.25 + 20.25}{4 - 1} = 15 = \frac{45}{3}, \qquad S_2^{2} = 1. \] \[ \left(\frac12 - \frac14\right)15 = 3.750000, \qquad \left(\frac14 - \frac16\right)1 = 0.083333, \qquad \operatorname{Var}(\bar y) = 3.833333. \]

With \(m = M = 3\) the second term vanishes, leaving 3.750000 — exactly the single-stage cluster variance of Practical 6, the check that the scale of \(S_1^{2}\) is right.

\[ S_1^{2} - \frac{S_2^{2}}{M} = 15 - \frac13 = 14.666667, \qquad m_{\text{opt}} = \sqrt{\frac{30}{10}\cdot\frac{1}{14.666667}} = \sqrt{0.204545} = 0.452267. \] \[ \text{Second population: } 10 - 6.666667 = 3.333333, \qquad m_{\text{opt}} = \sqrt{\frac{20}{30}\cdot\frac{20}{3.333333}} = \sqrt 4 = 2. \]

5. Result

\(\operatorname{Var}(\bar y) = 3.8333\): sub-sampling two units instead of three cost 0.083333 of the variance (2.17%) and saved a third of the measurement. In this population \(m_{\text{opt}} = 0.45 < 1\), so measure one unit in each of as many first-stage units as the budget reaches. In the second, \(m_{\text{opt}} = 2\) exactly. When first-stage units differ from one another and their contents do not, visit many and measure few; when they are alike and their contents vary, visit few and measure many (worked also in Unit 4, Example 4.1).

How Marks Are Lost

THE RECURRING ERRORS

What the Practical Record Should Contain

FOR EACH EXPERIMENT
  1. 1. Problem — the sampling design named (with or without replacement, equal or unequal probabilities, stratified or clustered), the population quantities assumed known and how they are known, and what is to be found.
  2. 2. Aim — in one line.
  3. 3. Formula — the estimator and its variance, and the steps that apply them, including the selection procedure.
  4. 4. Calculation — the random numbers actually used, and the full working: every summary total, every intermediate quantity, no step omitted; the estimate, its estimated variance and its standard error; the comparison with the simple random sampling estimator, as an efficiency, with both variances shown.
  5. 5. Result — the conclusion in words, naming which estimator is recommended and why.