This unit begins with the three questions that section leaves open. How biased, exactly — the word "slightly" is not a quantity, and section 1 computes the bias over all ten possible samples and compares it with the theoretical expression. What the difference estimator is, and why using a pre-assigned slope makes the estimator exactly unbiased where estimating one does not. And what happens in a stratified design, where the estimator can be formed separately in each stratum or once over the combined sample — two different estimators with different biases and different variances.
The ratio estimator is a ratio of two random variables, \(\hat{\bar Y}_R = (\bar y/\bar x)\bar X\), and the expectation of a ratio is not the ratio of the expectations. That single fact is the whole source of the bias.
The expression, to order \(1/n\). Write \(\bar y = \bar Y(1+\varepsilon_0)\) and \(\bar x = \bar X(1+\varepsilon_1)\), so that \(E(\varepsilon_0) = E(\varepsilon_1) = 0\). Then
\[ \hat{\bar Y}_R = \bar Y\,\frac{1+\varepsilon_0}{1+\varepsilon_1} = \bar Y(1+\varepsilon_0)\left(1 - \varepsilon_1 + \varepsilon_1^{2} - \cdots\right), \]expanding the geometric series, which is legitimate for \(|\varepsilon_1| < 1\). Keeping terms to second order and taking expectations, the first-order terms vanish and
\[ E\!\left(\hat{\bar Y}_R\right) - \bar Y \approx \bar Y\left[E(\varepsilon_1^{2}) - E(\varepsilon_0\varepsilon_1)\right] = \frac{1-f}{n\,\bar X}\left(R\,S_x^{2} - S_{xy}\right), \]with \(R = \bar Y/\bar X\) and \(f = n/N\). Read it: the bias vanishes when \(S_{xy} = R\,S_x^{2}\), that is when the regression of \(y\) on \(x\) has slope exactly \(R\) — a line through the origin, which is precisely the situation the ratio estimator was designed for.
The bias is \(O(1/n)\) while the standard error is \(O(1/\sqrt n)\), so their ratio is \(O(1/\sqrt n)\) and the bias becomes negligible as the sample grows. That is the honest form of the Foundation course’s word "slightly", and Example 2.1 puts a number on it.
Given. The population from which the Foundation example drew its sample, now written out in full so that the bias can be computed rather than described: \(N = 5\), \(n = 2\), sampling without replacement.
| Unit | 1 | 2 | 3 | 4 | 5 | Mean |
|---|---|---|---|---|---|---|
| \(x_i\) | 10 | 12 | 15 | 9 | 14 | \(\bar X = 12\) |
| \(y_i\) | 9 | 11 | 14 | 8 | 13 | \(\bar Y = 11\) |
so \(R = 11/12 = 0.916667\). Note \(y_i = x_i - 1\) exactly, which makes \(\rho = 1\): a perfectly straight relationship that does not pass through the origin. That is the sharpest possible test of the ratio estimator, and section 2 uses the same data to show what the regression estimator does with it.
Step 1 — enumerate. There are \(\binom{5}{2} = 10\) samples. For each, compute \(\hat{\bar Y}_R = (\bar y/\bar x)\times 12\). The first three are
\[ \{1,2\}: \frac{10}{11}\times 12 = 10.909091, \qquad \{1,3\}: \frac{11.5}{12.5}\times 12 = 11.040000, \] \[ \{1,4\}: \frac{8.5}{9.5}\times 12 = 10.736842, \]and so on for all ten.
Step 2 — average them. Every sample has probability \(1/10\), so
\[ E\!\left(\hat{\bar Y}_R\right) = \frac{1}{10}\sum_{s}\hat{\bar Y}_R(s) = 10.986005. \]Step 3 — the exact bias.
\[ \text{Bias} = 10.986005 - 11 = -0.013995, \]which is \(-0.1272\%\) of \(\bar Y\).
Step 4 — the theoretical expression, for comparison. With \(f = 2/5\),
\[ S_x^{2} = S_{xy} = S_y^{2} = 6.5, \qquad R\,S_x^{2} = 0.916667 \times 6.5 = 5.958333, \] \[ R\,S_x^{2} - S_{xy} = 5.958333 - 6.5 = -0.541667, \] \[ \text{Bias} \approx \frac{1 - 0.4}{2 \times 12}\times(-0.541667) = \frac{0.6}{24}\times(-0.541667) = -0.013542. \]Step 5 — compare with the standard error. The approximate variance of the ratio estimator of the mean is \(\frac{1-f}{n}\left(S_y^{2} + R^{2}S_x^{2} - 2R\,S_{xy}\right)\), giving a standard error of \(0.116369\). So
\[ \frac{\text{bias}}{\text{SE}} = \frac{-0.013995}{0.116369} = -0.1203. \]Interpretation. The theoretical bias \(-0.013542\) recovers the exact \(-0.013995\) to within \(3.2\%\) of itself, at a sample size of two — which is as unfavourable as the approximation will ever face. And the bias is twelve per cent of one standard error: even here, an interval built ignoring the bias is barely affected, and at realistic sample sizes the ratio falls like \(1/\sqrt n\) and disappears. That is what justifies using a biased estimator, and it is a quantitative justification rather than a reassurance.
The regression estimator of the Foundation section estimates the slope from the sample. Suppose instead that a value \(b_0\) is fixed in advance, from previous surveys or from theory. The difference estimator is
\[ \hat{\bar Y}_D = \bar y + b_0\left(\bar X - \bar x\right). \]It is exactly unbiased, for any \(b_0\) whatever:
\[ E\!\left(\hat{\bar Y}_D\right) = E(\bar y) + b_0\left(\bar X - E(\bar x)\right) = \bar Y + b_0\left(\bar X - \bar X\right) = \bar Y, \]because \(b_0\) is a constant and both sample means are unbiased. No expansion, no approximation, no condition on \(n\). Its variance is exact too:
\[ \operatorname{Var}\!\left(\hat{\bar Y}_D\right) = \frac{1-f}{n}\left(S_y^{2} + b_0^{2}S_x^{2} - 2b_0S_{xy}\right), \]a quadratic in \(b_0\) minimised at \(b_0 = S_{xy}/S_x^{2} = \beta\), the population regression coefficient, where it takes the value \(\frac{1-f}{n}S_y^{2}(1-\rho^{2})\).
The three estimators, in one line each.
| Estimator | Slope used | Bias | Variance |
|---|---|---|---|
| Mean per unit | \(0\) | none | \(\frac{1-f}{n}S_y^{2}\) |
| Ratio | \(R = \bar Y/\bar X\) | \(O(1/n)\) | \(\frac{1-f}{n}\left(S_y^{2} + R^{2}S_x^{2} - 2RS_{xy}\right)\) |
| Difference | \(b_0\), fixed | none | \(\frac{1-f}{n}\left(S_y^{2} + b_0^{2}S_x^{2} - 2b_0S_{xy}\right)\) |
| Regression | \(b\), estimated | \(O(1/n)\) | \(\approx\frac{1-f}{n}S_y^{2}(1-\rho^{2})\) |
The ratio estimator is the difference estimator with \(b_0\) replaced by the estimated \(\hat R = \bar y/\bar x\) — and that substitution is exactly where its bias comes from. The regression estimator is the same story with \(b\) in place of \(\beta\).
On the data of Example 2.1 the point is stark. There \(\rho = 1\) exactly, so the regression estimator has approximate variance \(S_y^{2}(1-\rho^{2}) = 0\) — it recovers \(\bar Y\) perfectly, because \(y = x - 1\) and the fitted line reproduces that relation. The ratio estimator, forced through the origin, still has a standard error of \(0.116369\). A correlation of 1 is not enough for the ratio estimator; it needs proportionality, which is strictly more.
With \(L\) strata there is a choice. Either form a ratio within each stratum and combine, or combine first and form one ratio.
\[ \hat{\bar Y}_{RS} = \sum_{h=1}^{L}W_h\,\frac{\bar y_h}{\bar x_h}\,\bar X_h \qquad\text{(separate)}, \] \[ \hat{\bar Y}_{RC} = \frac{\bar y_{st}}{\bar x_{st}}\,\bar X \qquad\text{(combined)}, \qquad \bar y_{st} = \sum_h W_h\bar y_h, \;\; \bar x_{st} = \sum_h W_h\bar x_h, \]with \(W_h = N_h/N\). The separate estimator needs \(\bar X_h\) known for every stratum; the combined one needs only the overall \(\bar X\), which is often all that a census provides.
Which is better, and when.
The same dichotomy, with the same reasoning, applies to the regression estimator: \(\hat{\bar Y}_{lrs} = \sum_h W_h\left[\bar y_h + b_h(\bar X_h - \bar x_h)\right]\) separately, or \(\bar y_{st} + b_c(\bar X - \bar x_{st})\) combined.
Given. Two strata, \(N_1 = 100\) and \(N_2 = 200\), so \(N = 300\) and \(W_1 = 1/3\), \(W_2 = 2/3\). The known stratum means of the auxiliary variable are \(\bar X_1 = 5.5\) and \(\bar X_2 = 17\), so
\[ \bar X = \tfrac13(5.5) + \tfrac23(17) = \frac{79}{6} = 13.166667. \]| Stratum | \(n_h\) | sample \(x\) | sample \(y\) | \(\bar x_h\) | \(\bar y_h\) |
|---|---|---|---|---|---|
| 1 | 4 | 4, 6, 8, 2 | 10, 14, 18, 6 | 5 | 12 |
| 2 | 5 | 12, 16, 20, 14, 18 | 20, 26, 32, 23, 29 | 16 | 26 |
Step 1 — the separate estimator. The stratum ratios are
\[ \hat R_1 = \frac{12}{5} = 2.400000, \qquad \hat R_2 = \frac{26}{16} = 1.625000, \]and applying each to its own known mean,
\[ \hat R_1\bar X_1 = 2.4 \times 5.5 = 13.200000, \qquad \hat R_2\bar X_2 = 1.625 \times 17 = 27.625000, \] \[ \hat{\bar Y}_{RS} = \tfrac13(13.200000) + \tfrac23(27.625000) = \frac{1369}{60} = 22.816667. \]Step 2 — the combined estimator. First the stratified means,
\[ \bar y_{st} = \tfrac13(12) + \tfrac23(26) = \frac{64}{3} = 21.333333, \qquad \bar x_{st} = \tfrac13(5) + \tfrac23(16) = \frac{37}{3} = 12.333333, \]then one ratio for the whole sample,
\[ \hat R_c = \frac{64/3}{37/3} = \frac{64}{37} = 1.729730, \qquad \hat{\bar Y}_{RC} = 1.729730 \times 13.166667 = \frac{2528}{111} = 22.774775. \]Step 3 — the difference.
\[ 22.816667 - 22.774775 = 0.041892. \]Interpretation. The two differ by about \(0.2\%\), and they differ because the stratum ratios are not equal: \(2.4\) against \(1.625\). Had they been equal the two estimators would have coincided exactly, which is worth checking as an algebraic exercise. Here the ratios differ by nearly \(50\%\), so the separate estimator is fitting something real — but with only four and five units per stratum, its two \(O(1/n_h)\) biases are the larger worry. On this sample the combined estimator is the safer report, and the choice would reverse if each stratum held fifty units instead of five.