A finite population has four units \(Y_1 = 2,\; Y_2 = 4,\; Y_3 = 6,\; Y_4 = 8\) (\(N = 4\), population mean \(\bar Y = 5\)). Draw all possible simple random samples without replacement (SRSWOR) of size \(n = 2\) and verify that the sample mean is unbiased for \(\bar Y\); find its sampling variance.
To verify empirically that \(E(\bar y) = \bar Y\) under SRSWOR and that \(\text{Var}(\bar y) = \dfrac{N-n}{Nn}\,S^2\).
Applying it:
Blank working table (fill the mean of each of the \(\binom{4}{2}=6\) samples):
| Sample | Units | \(\bar y\) | \((\bar y - \bar Y)^2\) |
|---|---|---|---|
| 1 | |||
| 2 | |||
| 3 | |||
| 4 | |||
| 5 | |||
| 6 | |||
| Total | |||
| Sample | Units | \(\bar y\) | \((\bar y - \bar Y)^2\) |
|---|---|---|---|
| 1 | {2, 4} | 3 | 4 |
| 2 | {2, 6} | 4 | 1 |
| 3 | {2, 8} | 5 | 0 |
| 4 | {4, 6} | 5 | 0 |
| 5 | {4, 8} | 6 | 1 |
| 6 | {6, 8} | 7 | 4 |
| Total | 30 | 10 | |
\(E(\bar y) = 30/6 = 5 = \bar Y\).
Direct variance \(= 10/6 = 1.667\).
Formula check: \(S^2 = \dfrac{(2-5)^2+(4-5)^2+(6-5)^2+(8-5)^2}{4-1} = \dfrac{20}{3} = 6.667\), so \(\text{Var}(\bar y) = \dfrac{4-2}{4\cdot 2}\cdot\dfrac{20}{3} = \dfrac{40}{24} = 1.667\).
\(E(\bar y) = 5 = \bar Y\), so the sample mean is an unbiased estimator of the population mean under SRSWOR, and its variance is \(1.667\), matching the theoretical formula.
For the same population (\(2, 4, 6, 8\)) and all six SRSWOR samples of size \(n = 2\), verify that the sample mean square \(s^2\) is an unbiased estimator of the population mean square \(S^2 = 6.667\).
To verify that \(E(s^2) = S^2\) under SRSWOR.
Applying it:
Blank working table:
| Sample | \(\bar y\) | \(s^2\) |
|---|---|---|
| {2, 4} | ||
| {2, 6} | ||
| {2, 8} | ||
| {4, 6} | ||
| {4, 8} | ||
| {6, 8} | ||
| Total |
| Sample | \(\bar y\) | \(s^2\) |
|---|---|---|
| {2, 4} | 3 | 2 |
| {2, 6} | 4 | 8 |
| {2, 8} | 5 | 18 |
| {4, 6} | 5 | 2 |
| {4, 8} | 6 | 8 |
| {6, 8} | 7 | 2 |
| Total | — | 40 |
\(E(s^2) = 40/6 = 6.667\), and \(S^2 = 20/3 = 6.667\).
\(E(s^2) = 6.667 = S^2\), so the sample mean square is an unbiased estimator of the population mean square under SRSWOR.
For the population \(2, 4, 6, 8\), draw all ordered samples of size \(n = 2\) with replacement (SRSWR) and verify that \(\bar y\) is unbiased for \(\bar Y = 5\); find its variance.
To verify that \(E(\bar y) = \bar Y\) under SRSWR and that \(\text{Var}(\bar y) = \sigma^2/n\).
Applying it:
Blank working table:
| Quantity | Value |
|---|---|
| Number of ordered samples \(N^n\) | |
| \(E(\bar y)\) | |
| \(\sigma^2\) | |
| \(\text{Var}_{WR}(\bar y) = \sigma^2/n\) |
| Quantity | Value |
|---|---|
| Number of ordered samples \(N^n\) | 16 |
| \(E(\bar y)\) | 5 |
| \(\sigma^2 = 20/4\) | 5 |
| \(\text{Var}_{WR}(\bar y) = 5/2\) | 2.5 |
By the symmetry of replacement each unit appears with equal frequency across the 16 samples, so \(E(\bar y) = \bar Y = 5\).
\(E(\bar y) = 5 = \bar Y\) (unbiased) and \(\text{Var}_{WR}(\bar y) = 2.5\). Since \(2.5 > 1.667\), SRSWOR is more precise than SRSWR for a fixed \(n\).
Using the results of Experiments 1 and 3, compare the mean and variance of \(\bar y\) under SRSWR and SRSWOR and compute the relative efficiency of SRSWOR.
To show that both designs give unbiased means but SRSWOR has the smaller variance, by the factor \((N-n)/(N-1)\).
Applying it:
Blank working table:
| SRSWR | SRSWOR | |
|---|---|---|
| \(E(\bar y)\) | ||
| \(\text{Var}(\bar y)\) | ||
| FPC factor | ||
| Relative efficiency |
| SRSWR | SRSWOR | |
|---|---|---|
| \(E(\bar y)\) | 5 | 5 |
| \(\text{Var}(\bar y)\) | 2.500 | 1.667 |
| FPC factor | 1 | \((N-n)/N = 0.5\) |
| Relative efficiency | 1.00 | 1.50 |
\(\text{RE} = 2.5/1.667 = 1.50\).
Both designs are unbiased. SRSWOR has the smaller variance and is 50 % more efficient than SRSWR for this population.
A population is divided into three strata with \(N_1 = 200,\; N_2 = 300,\; N_3 = 500\) and within-stratum standard deviations \(S_1 = 4,\; S_2 = 6,\; S_3 = 8\). A total sample of \(n = 50\) is to be drawn. Allocate the sample among the strata by proportional and by optimum (Neyman) allocation.
To compute stratum sample sizes under proportional allocation \((n_h = nN_h/N)\) and Neyman allocation \((n_h = nN_hS_h/\sum N_kS_k)\).
Applying it:
Blank working table:
| Stratum | \(N_h\) | \(S_h\) | \(N_h S_h\) | \(n_h\) (prop.) | \(n_h\) (Neyman) |
|---|---|---|---|---|---|
| 1 | |||||
| 2 | |||||
| 3 | |||||
| Total |
| Stratum | \(N_h\) | \(S_h\) | \(N_h S_h\) | \(n_h\) (prop.) | \(n_h\) (Neyman) |
|---|---|---|---|---|---|
| 1 | 200 | 4 | 800 | 10 | 6 |
| 2 | 300 | 6 | 1800 | 15 | 14 |
| 3 | 500 | 8 | 4000 | 25 | 30 |
| Total | 1000 | — | 6600 | 50 | 50 |
Proportional: \(n_1 = 50\cdot 200/1000 = 10,\; n_2 = 15,\; n_3 = 25\).
Neyman: \(n_1 = 50\cdot 800/6600 = 6.06 \approx 6,\; n_2 = 50\cdot 1800/6600 = 13.64 \approx 14,\; n_3 = 50\cdot 4000/6600 = 30.30 \approx 30\).
Under proportional allocation the sample is \((10, 15, 25)\); under Neyman allocation it is \((6, 14, 30)\). The most variable stratum (stratum 3, largest \(S_h\)) receives a larger share under Neyman allocation.
For the strata of Experiment 5 with weights \(W_1 = 0.2,\; W_2 = 0.3,\; W_3 = 0.5\) and variances \(S_1^2 = 16,\; S_2^2 = 36,\; S_3^2 = 64\), compare the variance of the estimated mean under proportional and optimum allocation with SRSWOR (\(n = 50\)) and find the gain in efficiency.
To compute \(\text{Var}_{prop}\), \(\text{Var}_{opt}\) and \(\text{Var}_{SRS}\), and the percentage gain of stratified over SRSWOR.
Applying it:
Blank working table:
| Stratum | \(W_h\) | \(S_h\) | \(S_h^2\) | \(W_h S_h\) | \(W_h S_h^2\) |
|---|---|---|---|---|---|
| 1 | |||||
| 2 | |||||
| 3 | |||||
| Total | — | — |
| Stratum | \(W_h\) | \(S_h\) | \(S_h^2\) | \(W_h S_h\) | \(W_h S_h^2\) |
|---|---|---|---|---|---|
| 1 | 0.2 | 4 | 16 | 0.8 | 3.2 |
| 2 | 0.3 | 6 | 36 | 1.8 | 10.8 |
| 3 | 0.5 | 8 | 64 | 4.0 | 32.0 |
| Total | 1.0 | — | — | 6.6 | 46.0 |
\(\text{Var}_{prop} = 46/50 = 0.920\).
\(\text{Var}_{opt} = 6.6^2/50 = 43.56/50 = 0.871\).
\(\text{Var}_{SRS} = 50/50 = 1.000\) (taking \(S^2 \approx 50\)).
Gain: proportional \(= 1.000/0.920 - 1 = 8.7\,\%\); optimum \(= 1.000/0.871 - 1 = 14.8\,\%\).
The precision ordering is \(\text{Var}_{opt}(0.871) < \text{Var}_{prop}(0.920) < \text{Var}_{SRS}(1.000)\). Stratification gives a gain of about 8.7 % (proportional) and 14.8 % (optimum) over SRSWOR.
A population of \(N = 16\) units (arranged in order) has the values 2, 3, 5, 7, 9, 10, 12, 13, 15, 17, 18, 20, 22, 24, 25, 27. Draw a linear systematic sample of size \(n = 4\) (so \(k = N/n = 4\)) and compare its precision with stratified sampling (one unit per group of \(k\)) and SRSWOR.
To obtain all \(k\) possible systematic samples, verify unbiasedness, and compare \(\text{Var}_{sys}\), \(\text{Var}_{stratified}\) and \(\text{Var}_{SRSWOR}\).
Applying it:
Blank working table (systematic samples):
| \(r\) | Units \((r, r+k, r+2k, r+3k)\) | \(\bar y_r\) | \((\bar y_r - \bar Y)^2\) |
|---|---|---|---|
| 1 | |||
| 2 | |||
| 3 | |||
| 4 | |||
| Total | |||
Population mean \(\bar Y = 229/16 = 14.3125\).
| \(r\) | Units | \(\bar y_r\) | \((\bar y_r - \bar Y)^2\) |
|---|---|---|---|
| 1 | {2, 9, 15, 22} | 12.00 | 5.348 |
| 2 | {3, 10, 17, 24} | 13.50 | 0.660 |
| 3 | {5, 12, 18, 25} | 15.00 | 0.473 |
| 4 | {7, 13, 20, 27} | 16.75 | 5.941 |
| Total | 57.25 | 12.422 | |
\(E(\bar y_{sys}) = 57.25/4 = 14.3125 = \bar Y\) (unbiased).
\(\text{Var}_{sys} = 12.422/4 = 3.11\).
SRSWOR: \(\sum(Y_i-\bar Y)^2 = 955.44\), so \(S^2 = 955.44/15 = 63.70\) and \(\text{Var}_{SRSWOR} = \dfrac{16-4}{16\cdot 4}\cdot 63.70 = \dfrac{12}{64}\cdot 63.70 = 11.94\).
Stratified (strata {2,3,5,7}, {9,10,12,13}, {15,17,18,20}, {22,24,25,27}; \(S_h^2 = 4.917, 3.333, 4.333, 4.333\)): \(\text{Var}_{st} = (0.25)^2\cdot\tfrac{3}{4}\,(4.917+3.333+4.333+4.333) = 0.0625\cdot 0.75\cdot 16.917 = 0.79\).
All three designs are unbiased. The precision ordering is \(\text{Var}_{stratified}(0.79) < \text{Var}_{sys}(3.11) < \text{Var}_{SRSWOR}(11.94)\). For this ordered, near-linear population, stratified sampling is the most precise and systematic sampling is far better than SRSWOR.