Skip to the content

Topics Covered

Two-Stage Sampling Optimum Sub-Sample Response Bias Non-Response Randomized Response Warner's Model Unrelated Question Small Area Estimation
On this page
  1. 1. Two-Stage Sampling with Equal First-Stage Units
  2. 2. Non-Sampling Errors, Quantitatively
  3. 3. Randomized Response
  4. 4. Small Area Estimation
  5. Key Take-aways
Where this unit starts. Multistage sampling is already introduced at Foundation level: Sampling Techniques, Unit 4, section 3 defines primary and secondary sampling units, lists the advantages and gives two examples. Likewise Unit 1, section 5 distinguishes sampling from non-sampling error and lists the sources of the latter — frame, questionnaire, investigator, respondent, non-response, coverage, processing. Neither is repeated here.

Again, no formulae appear in either. This unit supplies the variance of a two-stage estimator and shows how it splits into a between-PSU and a within-PSU part, which is what makes the optimum sub-sample size computable; then treats non-sampling error quantitatively, as a bias and a variance rather than a list; then covers two topics with no Foundation counterpart at all — randomized response, which buys truthful answers to sensitive questions at a measurable price in precision, and small area estimation, which is what to do when the sample in a domain of interest is too small to support a direct estimate.

1. Two-Stage Sampling with Equal First-Stage Units

THE ESTIMATOR AND THE TWO-PART VARIANCE

Let there be \(N\) first-stage units (PSUs) of \(M\) second-stage units each. Select \(n\) PSUs by simple random sampling without replacement, and from each selected PSU select \(m\) units, again without replacement. The estimator of the population mean per unit is

\[ \bar y = \frac{1}{n}\sum_{i \in s}\bar y_i, \qquad \bar y_i = \frac{1}{m}\sum_{j \in s_i} y_{ij}, \]

and its variance is

\[ \operatorname{Var}(\bar y) = \left(\frac{1}{n} - \frac{1}{N}\right)S_1^{2} + \left(\frac{1}{nm} - \frac{1}{nM}\right)S_2^{2}, \]

where

\[ S_1^{2} = \frac{1}{N-1}\sum_{i=1}^{N}\left(\bar Y_i - \bar Y\right)^{2}, \qquad S_2^{2} = \frac{1}{N(M-1)}\sum_{i=1}^{N}\sum_{j=1}^{M}\left(y_{ij} - \bar Y_i\right)^{2}. \]

Mind the scale of \(S_1^{2}\). It is the mean square among the PSU means, not the per-unit between-cluster mean square \(S_b^{2}\) of Unit 3; the two differ by a factor \(M\), since \(S_b^{2} = M S_1^{2}\). Substituting \(S_b^{2}\) here inflates the first term \(M\)-fold, and it is the commonest slip in the formula.

Read the two terms. The first is the price of not visiting every PSU; it disappears when \(n = N\). The second is the price of not measuring every unit in the PSUs visited; it disappears when \(m = M\), and setting \(m = M\) recovers the single-stage cluster formula of Unit 3 exactly. Two-stage sampling is the general case of which cluster sampling and simple random sampling are the two extremes: \(m = M\) gives clusters, \(n = N\) gives stratified-like complete coverage of PSUs.

An unbiased variance estimator from the sample alone is

\[ \widehat{\operatorname{Var}}(\bar y) = \left(\frac{1}{n} - \frac{1}{N}\right)s_1^{2} + \left(\frac{1}{nm} - \frac{1}{nM}\right)s_2^{2}, \]

with \(s_1^{2} = \frac{1}{n-1}\sum_{i \in s}(\bar y_i - \bar y)^{2}\cdot m\) adjusted for the within-PSU contribution it contains, and \(s_2^{2}\) the pooled within-PSU sample mean square. The adjustment matters: the observed scatter of the \(\bar y_i\) contains both between-PSU variation and the sampling error of each \(\bar y_i\), so using it raw over-estimates \(S_1^{2}\).

THE OPTIMUM SUB-SAMPLE SIZE

With the cost function \(C = c_1 n + c_2 nm\) — \(c_1\) to reach a PSU, \(c_2\) per unit measured — minimising the variance for fixed cost gives

\[ m_{\text{opt}} = \sqrt{\frac{c_1}{c_2}\cdot \frac{S_2^{2}}{S_1^{2} - S_2^{2}/M}}. \]

Three readings, and the corner cases are as important as the formula:

EXAMPLE 4.1 — A TWO-STAGE DESIGN ON THE CLUSTER POPULATION

Given. The population of Example 3.1: \(N = 4\) PSUs of \(M = 3\) units. Its four PSU means were \(5, 11, 8, 14\) about a grand mean of \(9.5\), so

\[ S_1^{2} = \frac{20.25 + 2.25 + 2.25 + 20.25}{4 - 1} = \frac{45}{3} = 15, \]

which is the \(S_b^{2} = 45\) of that example divided by \(M = 3\), as it must be. The pooled within-PSU mean square is \(S_2^{2} = 1\). Take \(n = 2\) PSUs and \(m = 2\) units from each.

Step 1 — the two terms.

\[ \left(\frac{1}{2} - \frac{1}{4}\right)15 = 0.25 \times 15 = 3.750000, \] \[ \left(\frac{1}{4} - \frac{1}{6}\right)1 = 0.083333 \times 1 = 0.083333. \]

Step 2 — the variance.

\[ \operatorname{Var}(\bar y) = 3.750000 + 0.083333 = 3.833333. \]

Step 3 — compare with measuring the whole PSU. Setting \(m = M = 3\) kills the second term and leaves \(3.750000\) — exactly the single-stage cluster variance computed in Example 3.1 for \(n = 2\). That agreement is the check that the scale of \(S_1^{2}\) is right, and it is worth doing every time. So sub-sampling two units instead of three cost \(0.083333\) out of \(3.833333\), which is \(2.17\%\) of the variance, and saved a third of the measurement.

Step 4 — the optimum \(m\) here. With \(c_1 = 30\) and \(c_2 = 10\),

\[ S_1^{2} - \frac{S_2^{2}}{M} = 15 - \frac{1}{3} = 14.666667, \qquad m_{\text{opt}} = \sqrt{\frac{30}{10}\cdot\frac{1}{14.666667}} = \sqrt{0.204545} = 0.452267, \]

far below 1, so \(m = 1\). Measure one unit in each of as many PSUs as the budget reaches. With \(\rho = 0.916084\) in this population that is exactly right: the three units in a PSU are near copies of one another, so the second and third add almost nothing.

Step 5 — a population where the optimum is interior. Suppose instead \(S_1^{2} = 10\), \(S_2^{2} = 20\), \(M = 3\), \(c_1 = 20\), \(c_2 = 30\):

\[ S_1^{2} - \frac{S_2^{2}}{M} = 10 - 6.666667 = 3.333333, \qquad m_{\text{opt}} = \sqrt{\frac{20}{30}\cdot\frac{20}{3.333333}} = \sqrt{4} = 2. \]

Two units per PSU exactly, and with a budget of \(600\) that gives \(n = 600/(20 + 30 \times 2) = 7.5\), so eight PSUs and a slightly trimmed budget, or seven and some in reserve.

Interpretation. The two cases differ only in where the variation sits. When PSUs differ from one another and their contents do not, visit many and measure few; when PSUs are alike and their contents vary, visit few and measure many. The formula is the quantitative form of that sentence, and the corner solutions are not failures of it but its correct answer.

2. Non-Sampling Errors, Quantitatively

A MODEL, RATHER THAN A LIST

The Foundation section lists the sources. To do anything with them, write a model. Let the recorded value for unit \(i\) on occasion \(t\) be

\[ y_{it} = Y_i + b_i + e_{it}, \]

where \(Y_i\) is the true value, \(b_i\) a response bias that does not change on re-measurement, and \(e_{it}\) a response variance term with mean zero. Then for the sample mean

\[ E(\bar y) = \bar Y + \bar b, \qquad \operatorname{MSE}(\bar y) = \underbrace{\operatorname{Var}_{\text{sampling}}} _{\text{falls like } 1/n} + \underbrace{\frac{\sigma_e^{2}}{n}}_{\text{response variance}} + \underbrace{\bar b^{2}}_{\text{does not fall at all}}. \]

That last term is the whole problem. Sampling error and response variance both shrink as the survey grows; response bias does not. A census has \(n = N\) and therefore no sampling error whatever, and is still wrong by \(\bar b\). That is the precise sense in which the Foundation table's remark — that non-sampling error is present in a census and may increase with \(n\) — is true: a larger survey uses more and less well-trained investigators, and \(\bar b\) grows.

Non-response, the commonest source, has the same structure. If a proportion \(W_2\) of the population would not respond and has mean \(\bar Y_2\) against the respondents' \(\bar Y_1\), then reporting the respondent mean carries a bias of

\[ \bar Y_1 - \bar Y = W_2\left(\bar Y_1 - \bar Y_2\right), \]

the product of how many are missing and how different they are. A \(30\%\) non-response rate is harmless if the missing are like the present, and fatal if they are not — and the data cannot tell you which. Hansen and Hurwitz's sub-sampling of non-respondents is the standard repair: take a sub-sample of size \(h_2 = m_2/k\) from the \(m_2\) non-respondents, pursue those intensively, and estimate

\[ \bar y' = \frac{m_1\bar y_1 + m_2\bar y_{2h}}{m_1 + m_2}, \]

which is unbiased because the sub-sample represents the whole non-respondent stratum. The cost is the intensive follow-up; the benefit is that the bias is replaced by a variance, and a variance can be reported.

3. Randomized Response

WARNER'S MODEL

Some questions will not be answered truthfully if asked directly. Warner's device removes the respondent's reason to lie by ensuring that the investigator never learns which question was answered.

The respondent spins a device that selects, with probability \(P\), the statement "I belong to group A" and with probability \(1-P\) its negation "I do not belong to group A". The investigator sees only "yes" or "no", and never which statement was addressed. If \(\pi\) is the proportion in group A, the probability of a "yes" is

\[ \lambda = P\pi + (1-P)(1-\pi), \]

which solves to give the estimator

\[ \hat\pi = \frac{\hat\lambda - (1-P)}{2P - 1}, \qquad \hat\lambda = \frac{\text{number of yes answers}}{n}, \]

unbiased because \(\hat\lambda\) is. Its variance follows from that of a binomial proportion divided by the squared coefficient:

\[ \operatorname{Var}(\hat\pi) = \frac{\lambda(1-\lambda)}{n\left(2P-1\right)^{2}} = \underbrace{\frac{\pi(1-\pi)}{n}}_{\text{if } \pi \text{ were observed}} + \underbrace{\frac{P(1-P)}{n(2P-1)^{2}}}_{\text{the price of the device}}. \]

The design tension is visible in the denominator. \(P\) near \(\tfrac12\) gives the best protection — a "yes" then says almost nothing — and \((2P-1)^{2} \to 0\), so the variance explodes. \(P\) near 1 is efficient and offers no protection, so respondents lie again and the whole exercise fails. \(P = \tfrac12\) exactly is inadmissible: the estimator does not exist.

EXAMPLE 4.2 — WHAT THE PROTECTION COSTS

Given. \(n = 400\) respondents, \(P = 0.7\), and \(148\) "yes" answers.

Step 1 — the observed proportion.

\[ \hat\lambda = \frac{148}{400} = 0.370000. \]

Step 2 — the estimate.

\[ \hat\pi = \frac{0.370000 - 0.3}{2(0.7) - 1} = \frac{0.070000}{0.4} = 0.175000. \]

Step 3 — check it forwards. \(\lambda = 0.7(0.175) + 0.3(0.825) = 0.1225 + 0.2475 = 0.37\). \(\checkmark\)

Step 4 — the standard error.

\[ \operatorname{Var}(\hat\pi) = \frac{0.370000 \times 0.630000}{400 \times 0.16} = \frac{0.233100}{64} = 0.00364219, \qquad SE = 0.060351. \]

Step 5 — what a direct question would have cost. If \(\pi\) could have been observed honestly,

\[ SE = \sqrt{\frac{0.175 \times 0.825}{400}} = 0.018998, \]

so the device has multiplied the standard error by \(3.176619\) — equivalent to throwing away about \(90\%\) of the sample.

Step 6 — the trade-off, across \(P\).

\(P\)\((2P-1)^{2}\)\(SE(\hat\pi)\)Protection
0.60.040.120701very good
0.70.160.060351good
0.80.360.040234moderate
0.90.640.030175poor

Interpretation. Moving \(P\) from \(0.9\) to \(0.6\) quadruples the standard error. The choice is not statistical: it depends on how sensitive the question is, and the right way to make it is to ask what \(P\) respondents will believe protects them. An efficient design that nobody trusts produces precise estimates of a lie.

EXAMPLE 4.3 — THE UNRELATED QUESTION MODEL

The improvement. Warner's second statement is the negation of the first, which some respondents find no more comfortable than the direct question. The unrelated question model replaces it with something harmless — "were you born in the first half of the year?" — whose population proportion \(\pi_Y\) may be known or may itself be estimated. With \(\pi_Y\) unknown, two independent samples with different selection probabilities \(P_1 \ne P_2\) are used:

\[ \lambda_1 = P_1\pi + (1-P_1)\pi_Y, \qquad \lambda_2 = P_2\pi + (1-P_2)\pi_Y, \]

two linear equations in two unknowns, solved by

\[ \hat\pi = \frac{(1-P_2)\hat\lambda_1 - (1-P_1)\hat\lambda_2}{P_1 - P_2}. \]

Given. \(P_1 = 0.7\) with \(n_1 = 300\) and \(135\) yes answers; \(P_2 = 0.3\) with \(n_2 = 300\) and \(105\) yes answers.

Step 1 — the two proportions. \(\hat\lambda_1 = 135/300 = 0.450000\), \(\hat\lambda_2 = 105/300 = 0.350000\).

Step 2 — solve. Here \(1 - P_2 = 0.7\) and \(1 - P_1 = 0.3\), so

\[ \hat\pi = \frac{(1-P_2)\hat\lambda_1 - (1-P_1)\hat\lambda_2}{P_1 - P_2} = \frac{0.7(0.450000) - 0.3(0.350000)}{0.7 - 0.3} \\ = \frac{0.315000 - 0.105000}{0.4} = \frac{0.210000}{0.4} = 0.525000. \]

Note which complement goes with which proportion: \(\hat\lambda_1\) is multiplied by \(1 - P_2\), not by \(1 - P_1\). Pairing them the other way round is the commonest error in this calculation, and the check in Step 3 catches it.

Step 3 — recover the unrelated proportion and check.

\[ \hat\pi_Y = \frac{P_1\hat\lambda_2 - P_2\hat\lambda_1}{P_1 - P_2} = \frac{0.7(0.350000) - 0.3(0.450000)}{0.4} = \frac{0.245000 - 0.135000}{0.4} = 0.275000, \]

and substituting both back,

\[ 0.7(0.525) + 0.3(0.275) = 0.3675 + 0.0825 = 0.450000 = \hat\lambda_1, \checkmark \] \[ 0.3(0.525) + 0.7(0.275) = 0.1575 + 0.1925 = 0.350000 = \hat\lambda_2. \checkmark \]

Interpretation. Both equations reproduce exactly, so the solution is right. Note that \(\hat\pi_Y = 0.275\) is an estimate of a quantity that was never asked about directly and is of no interest — it is the price of not having to know it in advance, paid in one extra parameter and a second sample.

4. Small Area Estimation

THREE ESTIMATORS, AND WHY THE THIRD IS USED

A national survey is designed to give precise national figures. Asked for a figure for one district, it may have eight observations there, and the direct estimate is useless. Three approaches:

\[ w^{*} = \frac{\operatorname{MSE}\!\left(\hat{\bar Y}_{syn}\right)} {\operatorname{Var}\!\left(\hat{\bar Y}_d\right) + \operatorname{MSE}\!\left(\hat{\bar Y}_{syn}\right)}, \]

which leans on the direct estimate where it is reliable and on the synthetic one where it is not, automatically. In practice \(w\) is often set by a rule of thumb such as \(w = n_d/(n_d + \bar n)\), because the synthetic estimator's bias cannot be measured from the data that produced it.

EXAMPLE 4.4 — THE COMPOSITE ESTIMATOR BEATS BOTH

Given. For one small area: a direct estimate \(6.2\) from \(n_d = 8\) observations with \(S^{2} = 9\); a synthetic estimate \(5.8\) with variance \(0.04\) and a bias believed to be about \(0.3\).

Step 1 — the two mean squared errors.

\[ \operatorname{Var}\!\left(\hat{\bar Y}_d\right) = \frac{9}{8} = 1.125000, \qquad SE = 1.060660, \] \[ \operatorname{MSE}\!\left(\hat{\bar Y}_{syn}\right) = 0.04 + (0.3)^{2} = 0.04 + 0.09 = 0.130000. \]

Step 2 — the optimal weight.

\[ w^{*} = \frac{0.130000}{1.125000 + 0.130000} = \frac{0.130000}{1.255000} = 0.103586. \]

Step 3 — the composite estimate.

It is least error-prone written as the synthetic estimate plus a correction, since the two estimates differ by only \(0.4\):

\[ \hat{\bar Y}_C = 5.8 + w^{*}\left(6.2 - 5.8\right) = 5.8 + 0.103586(0.4) = 5.8 + 0.041434 = 5.841434. \]

Step 4 — its mean squared error.

\[ \operatorname{MSE} = (0.103586)^{2}(1.125000) + (0.896414)^{2}(0.130000) = 0.012071 + 0.104463 = 0.116534. \]

There is a shortcut worth knowing: at the optimal weight the mean squared error collapses to \(w^{*}\operatorname{MSE}(\hat{\bar Y}_{syn})\), the harmonic-mean form \(\frac{1.125000 \times 0.130000}{1.255000} = 0.116534\), which is a useful check on the longer route.

Step 5 — compare all three.

EstimatorValueMSERoot MSE
Direct6.2000001.1250001.060660
Synthetic5.8000000.1300000.360555
Composite5.8414340.1165340.341371

Interpretation. The composite estimator beats both its ingredients, which is the whole reason for forming it: a weighted average of two estimators has a mean squared error below the smaller of the two whenever the weight is chosen well. Note how little weight the direct estimate receives — about a tenth — because its variance is nearly ten times the synthetic estimator's mean squared error. The honest caveat is that the \(0.3\) bias was assumed: it cannot be estimated from the same data, and every published small-area figure rests on an assumption of that kind.

Key Take-aways