All seven are worked below, each set out as 1. Problem, 2. Aim, 3. Formula, 4. Calculation, 5. Result. Practicals 2 and 4 share one sample, and 3 and 5 share the same two strata, so the estimators can be compared directly.
| # | Practical | Worked here | Also in |
|---|---|---|---|
| 1 | PPS sampling with and without replacement | Practical 1 | Unit 1, Examples 1.1, 1.2 and 1.3 |
| 2 | Ratio estimators in SRS, comparison with SRS | Practical 2 | — |
| 3 | Separate and combined ratio estimators, comparison | Practical 3 | Unit 2, Example 2.2 |
| 4 | Regression estimators in SRS, comparison with SRS and ratio estimators | Practical 4 | — |
| 5 | Separate and combined regression estimators, comparison | Practical 5 | — |
| 6 | Cluster sampling with equal cluster sizes | Practical 6 | Unit 3, Examples 3.1 and 3.2 |
| 7 | Sub-sampling (two-stage) with equal first-stage units | Practical 7 | Unit 4, Example 4.1 |
(a) Five units have sizes \(x_i = 12, 8, 20, 5, 15\). Draw a sample of three with probability proportional to size, with replacement, by the cumulative total method, using the random numbers \(37, 09, 52\); and show how Lahiri's method treats the pairs \((4, 9)\) and \((2, 6)\).
(b) A population of \(N = 5\) units has sizes \(X_i = 10, 20, 30, 25, 15\) and (for this demonstration only) known values \(Y_i = 12, 26, 33, 30, 19\). For a PPS sample of \(n = 2\) with replacement, find the variance of the Hansen–Hurwitz estimator of the total, compare it with equal-probability sampling, and estimate the total if the draws fall on units 2 and 3.
(c) \(N = 4\) units have \(p = (0.1, 0.2, 0.3, 0.4)\) and \(Y = (3, 6, 9, 12)\), total 30. Two units are drawn successively without replacement. Find the inclusion probabilities, the Horvitz–Thompson estimate from the sample \(\{2, 4\}\), and its Yates–Grundy variance.
To draw samples with probability proportional to size, with and without replacement, and to estimate a population total with the estimator that belongs to each draw.
Which estimator goes with which draw. This is the part most often got wrong in the examination, so state it in the first line of the answer: Hansen–Hurwitz for PPS with replacement (it uses \(p_i\)); Horvitz–Thompson for PPS without replacement (it uses the inclusion probabilities \(\pi_i\), with the Yates–Grundy variance; Unit 1, section 4).
Applying it:
(a) \(X = 60\):
| Unit | \(x_i\) | Cumulative | Range | \(p_i = x_i/X\) |
|---|---|---|---|---|
| 1 | 12 | 12 | 1–12 | 0.200000 |
| 2 | 8 | 20 | 13–20 | 0.133333 |
| 3 | 20 | 40 | 21–40 | 0.333333 |
| 4 | 5 | 45 | 41–45 | 0.083333 |
| 5 | 15 | 60 | 46–60 | 0.250000 |
37 falls in 21–40, 09 in 1–12 and 52 in 46–60: the sample is units \(3, 1, 5\). By Lahiri's method, \(M = 20\): the pair \((4, 9)\) is rejected because \(9 > x_4 = 5\), and the pair \((2, 6)\) is accepted because \(6 \le x_2 = 8\).
(b) \(p_i = X_i/100 = 0.10, 0.20, 0.30, 0.25, 0.15\), and \(Y_i/p_i = 120, 130, 110, 120, 126.666667\), whose departures from \(Y = 120\) are \(0, 10, -10, 0, 6.666667\). So
\[ \sum_i p_i\left(\frac{Y_i}{p_i} - Y\right)^{2} = 0.20(100) + 0.30(100) + 0.15(44.444444) = 56.666667, \qquad \operatorname{Var}(\hat Y_{HH}) = \frac{1}{2}\cdot\frac{170}{3} = \frac{85}{3} = 28.333333. \]Under simple random sampling with replacement, \(\bar Y = 24\), \(\sigma^{2} = (144 + 4 + 81 + 36 + 25)/5 = 58\), and \(\operatorname{Var}(N\bar y) = 25 \times 58/2 = 725\); the ratio is \(725/(85/3) = 25.588235\). From units 2 and 3, \(\hat Y_{HH} = \tfrac12(130 + 110) = 120\).
(c) For unit 1,
\[ \pi_1 = 0.1 + 0.2\!\left(\frac{0.1}{0.8}\right) + 0.3\!\left(\frac{0.1}{0.7}\right) + 0.4\!\left(\frac{0.1}{0.6}\right) = \frac{197}{840} = 0.234524, \]and likewise \(\pi_2 = 139/315 = 0.441270\), \(\pi_3 = 73/120 = 0.608333\), \(\pi_4 = 451/630 = 0.715873\); their sum is \(2 = n\). \(\checkmark\) The pairwise probabilities are \(\pi_{12} = 17/360\), \(\pi_{13} = 8/105\), \(\pi_{14} = 1/9\), \(\pi_{23} = 9/56\), \(\pi_{24} = 7/30\), \(\pi_{34} = 13/35\), summing to \(1 = n(n-1)/2\). \(\checkmark\)
\[ \hat Y_{HT} = \frac{6}{0.441270} + \frac{12}{0.715873} = 13.597122 + 16.762749 = 30.3599, \]and with \(Y_i/\pi_i = 12.791878, 13.597122, 14.794521, 16.762749\) the Yates–Grundy variance is \(2.428334\); computing it directly over the six possible samples gives the same value, with \(E(\hat Y_{HT}) = 30\) exactly.
(a) The PPS sample with replacement is units 3, 1, 5. (b) PPS sampling makes the estimator of the total 25.6 times as precise as equal-probability sampling, because \(Y_i\) is nearly proportional to \(X_i\); if \(Y_i\) were inversely related to \(X_i\), PPS would be worse. (c) \(\hat Y_{HT} = 30.36\) against the true 30, with standard error \(\sqrt{2.428334} = 1.558\), 5.2% of the total. Note that \(\pi_1 = 0.2345\) is not \(np_1 = 0.2\): successive sampling does not give inclusion probabilities exactly proportional to size (worked also in Unit 1, Examples 1.1, 1.2 and 1.3).
A population has \(N = 50\) plots. The auxiliary variable \(x\) is last season's yield, known for every plot with \(\bar X = 12\); \(y\) is this season's yield, measured on a simple random sample of \(n = 8\) plots without replacement.
| Plot | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | Total |
|---|---|---|---|---|---|---|---|---|---|
| \(x\) | 8 | 10 | 11 | 13 | 14 | 15 | 9 | 12 | 92 |
| \(y\) | 20 | 26 | 27 | 34 | 35 | 39 | 23 | 30 | 234 |
Estimate \(\bar Y\) by the ratio estimator, find its estimated variance, and compare it with the sample mean.
To estimate a population mean by the ratio estimator, which uses a known auxiliary mean, and to measure its gain in precision over the sample mean.
The ratio estimator fits one constant, the slope through the origin, so \(n - 1\) residual degrees of freedom remain. It is worth using when \(\rho > C_x/(2C_y)\), \(C\) being a coefficient of variation.
Applying it:
Mean per unit. \(v(\bar y) = \frac{0.84}{8}\times 41.642857 = 4.372500\), s.e. \(= 2.091052\).
Ratio. \(\hat R = 29.25/11.5 = 117/46 = 2.543478\), and \(\hat{\bar Y}_R = \frac{117}{46}\times 12 = \frac{702}{23} = 30.521739\) (rounding \(\hat R\) first gives 30.521736, wrong in the last digit). The residuals are
| \(i\) | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| \(e_i\) | −0.347826 | 0.565217 | −0.978261 | 0.934783 | −0.608696 | 0.847826 | 0.108696 | −0.521739 |
Efficiency \(= 4.372500/0.054669 = 80.0\). \(C_x = \sqrt{6}/11.5 = 0.212999\), \(C_y = \sqrt{41.642857}/29.25 = 0.220620\), so \(C_x/(2C_y) = 0.4827\).
\(\hat{\bar Y}_R = 30.52\) with standard error 0.234, against \(\bar y = 29.25\) with 2.091: the ratio estimator is 80 times as efficient as the sample mean here, because \(r = 0.994\) is far above \(C_x/(2C_y) = 0.48\). Report the estimate with its variance and say which formula produced it.
A population has two strata, \(N_1 = 100\) and \(N_2 = 200\), with known stratum means of the auxiliary variable \(\bar X_1 = 5.5\) and \(\bar X_2 = 17\). A stratified sample gives:
| Stratum | \(n_h\) | sample \(x\) | sample \(y\) | \(\bar x_h\) | \(\bar y_h\) |
|---|---|---|---|---|---|
| 1 | 4 | 4, 6, 8, 2 | 10, 14, 18, 6 | 5 | 12 |
| 2 | 5 | 12, 16, 20, 14, 18 | 20, 26, 32, 23, 29 | 16 | 26 |
Estimate \(\bar Y\) by the separate and by the combined ratio estimator, and compare them.
To estimate a population mean from a stratified sample by the separate and the combined ratio estimators, and to choose between them.
The separate estimator allows the ratio to differ between strata, but carries a bias of order \(\sum W_h/n_h\); the combined one has a bias of order \(1/n\) but is wrong in every stratum if the ratios genuinely differ (Unit 2, section 3).
Applying it:
\(\hat{\bar Y}_{RS} = 22.817\) and \(\hat{\bar Y}_{RC} = 22.775\): they differ by about 0.2%, and they differ because the stratum ratios are not equal (2.4 against 1.625). The separate estimator is fitting something real, but with only four and five units per stratum its two \(O(1/n_h)\) biases are the larger worry: on this sample the combined estimator is the safer report, and the choice would reverse if each stratum held fifty units (worked also in Unit 2, Example 2.2).
For the sample of Practical 2 (\(N = 50\), \(n = 8\), \(\bar X = 12\)):
| Plot | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | Total |
|---|---|---|---|---|---|---|---|---|---|
| \(x\) | 8 | 10 | 11 | 13 | 14 | 15 | 9 | 12 | 92 |
| \(y\) | 20 | 26 | 27 | 34 | 35 | 39 | 23 | 30 | 234 |
estimate \(\bar Y\) by the linear regression estimator, find its estimated variance, and compare it with the sample mean and the ratio estimator.
To estimate a population mean by the regression estimator, and to compare its precision with the sample mean and the ratio estimator on the same sample.
The regression estimator fits an intercept as well as a slope, so only \(n - 2\) residual degrees of freedom remain; the residual sum of squares is \(S_{yy} - bS_{xy}\), without listing the residuals. Its large-sample variance is \(\frac{1-f}{n}s_y^{2}(1 - r^{2})\).
Applying it:
| Estimator | Estimate of \(\bar Y\) | Estimated variance | s.e. | Efficiency over SRS |
|---|---|---|---|---|
| Mean per unit | 29.250000 | 4.372500 | 2.091052 | 1.00 |
| Ratio (Practical 2) | 30.521739 | 0.054669 | 0.233814 | 80.0 |
| Regression | 30.559524 | 0.059583 | 0.244097 | 73.4 |
Each efficiency is \(v(\bar y)\) divided by the variance in its row; one decimal place is all the six-figure variances support. The fitted line is \(\hat y = -0.869048 + 2.619048x\). With the large-sample formula, \(0.105 \times 41.642857 \times 0.011680 = 0.051071\).
\(\hat{\bar Y}_{lr} = 30.56\), s.e. 0.244: 73 times as efficient as the sample mean. Both auxiliary estimators cut the standard error by a factor near nine, because \(x\) and \(y\) are almost collinear — and here the ratio estimator wins, which surprises students told that the regression estimator is never worse. The fitted intercept is very nearly zero, so the line through the origin is almost the same line; and the regression estimator still pays for its intercept: its residual sum of squares, 3.404762, is smaller than the ratio's 3.644612, but it is divided by 6 instead of 7. "At least as efficient" is a statement about large-sample variances (0.051071 here); at \(n = 8\) the degree of freedom spent on the intercept costs more than the intercept is worth.
For the two strata of Practical 3 (\(N_1 = 100\), \(N_2 = 200\), \(\bar X_1 = 5.5\), \(\bar X_2 = 17\)), a sample with the same stratum means but with the values scattered about the line is:
| Stratum | \(n_h\) | sample \(x\) | sample \(y\) | \(\bar x_h\) | \(\bar y_h\) |
|---|---|---|---|---|---|
| 1 | 4 | 4, 6, 8, 2 | 11, 13, 19, 5 | 5 | 12 |
| 2 | 5 | 12, 16, 20, 14, 18 | 21, 26, 33, 22, 28 | 16 | 26 |
Estimate \(\bar Y\) by the separate and by the combined regression estimator, and compare them with the ratio estimators.
To estimate a population mean from a stratified sample by the separate and the combined regression estimators, and to choose between them.
\(b_c\) is the variance-weighted pooled slope, ignoring finite population corrections, which are negligible here.
Applying it:
Stratum 1: \(x - \bar x = -1, 1, 3, -3\), \(y - \bar y = -1, 1, 7, -7\); \(S_{xx,1} = 20\), \(S_{xy,1} = 44\), \(b_1 = 2.2\). Stratum 2: \(x - \bar x = -4, 0, 4, -2, 2\), \(y - \bar y = -5, 0, 7, -4, 2\); \(S_{xx,2} = 40\), \(S_{xy,2} = 60\), \(b_2 = 1.5\).
\[ \hat{\bar Y}_{lrs} = \tfrac13(12 + 2.2 \times 0.5) + \tfrac23(26 + 1.5 \times 1) = \tfrac13(13.1) + \tfrac23(27.5) = 4.366667 + 18.333333 = 22.700000. \] \[ \bar y_{st} = \frac{64}{3}, \quad \bar x_{st} = \frac{37}{3}; \qquad b_c = \frac{\left(\tfrac13\right)^{2}\frac{44}{12} + \left(\tfrac23\right)^{2}\frac{60}{20}}{\left(\tfrac13\right)^{2}\frac{20}{12} + \left(\tfrac23\right)^{2}\frac{40}{20}} = \frac{47/27}{29/27} = \frac{47}{29} = 1.620690. \] \[ \bar X - \bar x_{st} = \frac{79}{6} - \frac{37}{3} = \frac56, \qquad \hat{\bar Y}_{lrc} = 21.333333 + \frac{47}{29}\times\frac56 = 21.333333 + 1.350575 = 22.683908. \]| Estimator | Estimate of \(\bar Y\) |
|---|---|
| Stratified mean, \(\bar y_{st}\) (no auxiliary information) | 21.333333 |
| Separate ratio (Practical 3) | 22.816667 |
| Combined ratio (Practical 3) | 22.774775 |
| Separate regression | 22.700000 |
| Combined regression | 22.683908 |
The separate and combined regression estimates differ by \(22.700000 - 22.683908 = 0.016092\), about 0.07%, even though the slopes 2.2 and 1.5 are not close: the choice moves the estimate very little and the bias a great deal, so it is decided by \(n_h\). With four and five observations per stratum, the combined estimator is the defensible one here. All four auxiliary estimators sit near 22.7, while the stratified mean alone gives 21.33: the gain comes from using \(x\) at all; the choice among the four ways of using it is second order.
A population has \(N = 4\) clusters of \(M = 3\) units; \(n = 2\) clusters are to be selected by simple random sampling.
| Cluster | Values | \(\bar y_i\) |
|---|---|---|
| 1 | 4, 6, 5 | 5 |
| 2 | 10, 12, 11 | 11 |
| 3 | 7, 9, 8 | 8 |
| 4 | 13, 15, 14 | 14 |
Find the intra-cluster correlation, the variance of the cluster-sample mean, and its efficiency against a simple random sample of the same number of units. Then, for a budget of 1000 with \(c_1 = 40\) per cluster and \(c_2 = 5\) per unit, find the best cluster size at \(\rho = 0.916\) and at \(\rho = 0.1\).
To estimate a population mean by cluster sampling, measure the loss of precision through the intra-cluster correlation, and choose the cluster size for a budget.
Applying it:
The grand mean is 9.5.
\[ S_b^{2} = \frac{3}{3}\left[(5-9.5)^{2} + (11-9.5)^{2} + (8-9.5)^{2} + (14-9.5)^{2}\right] = 20.25 + 2.25 + 2.25 + 20.25 = 45, \] \[ S_w^{2} = \frac{4 \times 2}{4 \times 2} = 1, \qquad S^{2} = \frac{1}{11}\sum_{ij}(y_{ij} - 9.5)^{2} = 13, \qquad \rho = \frac{131}{143} = 0.916084. \] \[ \left[1 + 2(0.916084)\right]\times\frac{11}{9} = 2.832168 \times 1.222222 = 3.461538. \] \[ \operatorname{Var}(\bar y_{cl}) = \frac{0.5}{6}\times 45 = 3.75, \qquad \operatorname{Var}(\bar y_{SRS}) = 0.5 \times \frac{13}{6} = 1.083333, \qquad \frac{3.75}{1.083333} = \frac{45}{13} = 3.461538. \checkmark \] \[ \rho = 0.916084: \ M_{\text{opt}} = \sqrt{8 \times 0.091603} = 0.856; \qquad \rho = 0.1: \ M_{\text{opt}} = \sqrt{8 \times 9} = 8.485. \]\(\rho = 0.92\) and the design effect is 3.46: the cluster sample of six units is as precise as a simple random sample of \(6/3.46 = 1.73\) units. (The large-\(N\) design effect alone, 2.83, would understate the loss by 18%, because \(N = 4\) is nowhere near large.) Measuring three units in a cluster is close to measuring the same unit three times. For the budget, \(M_{\text{opt}} < 1\) at \(\rho = 0.916\), so take one unit from each of many clusters; at \(\rho = 0.1\), \(M = 8\) or 9. A cluster size chosen without an estimate of \(\rho\) is chosen arbitrarily (worked also in Unit 3, Examples 3.1 and 3.2).
The population of Practical 6 (\(N = 4\) first-stage units of \(M = 3\)) is sampled in two stages: \(n = 2\) first-stage units, and \(m = 2\) units from each. Find the variance of the sample mean, compare it with measuring the whole first-stage unit, and find the optimum \(m\) for \(c_1 = 30\) per first-stage unit and \(c_2 = 10\) per unit. Repeat the optimum for \(S_1^{2} = 10\), \(S_2^{2} = 20\), \(M = 3\), \(c_1 = 20\), \(c_2 = 30\).
To find the variance of the two-stage sample mean and the number of second-stage units that minimises it for a given cost.
\(S_1^{2}\) is the mean square among first-stage unit means; the per-unit between-cluster mean square of Practical 6 is \(M\) times it (Unit 4, section 1).
Applying it:
With \(m = M = 3\) the second term vanishes, leaving 3.750000 — exactly the single-stage cluster variance of Practical 6, the check that the scale of \(S_1^{2}\) is right.
\[ S_1^{2} - \frac{S_2^{2}}{M} = 15 - \frac13 = 14.666667, \qquad m_{\text{opt}} = \sqrt{\frac{30}{10}\cdot\frac{1}{14.666667}} = \sqrt{0.204545} = 0.452267. \] \[ \text{Second population: } 10 - 6.666667 = 3.333333, \qquad m_{\text{opt}} = \sqrt{\frac{20}{30}\cdot\frac{20}{3.333333}} = \sqrt 4 = 2. \]\(\operatorname{Var}(\bar y) = 3.8333\): sub-sampling two units instead of three cost 0.083333 of the variance (2.17%) and saved a third of the measurement. In this population \(m_{\text{opt}} = 0.45 < 1\), so measure one unit in each of as many first-stage units as the budget reaches. In the second, \(m_{\text{opt}} = 2\) exactly. When first-stage units differ from one another and their contents do not, visit many and measure few; when they are alike and their contents vary, visit few and measure many (worked also in Unit 4, Example 4.1).