Population: \(U_1, U_2, \ldots, U_N\); values \(Y_1, Y_2, \ldots, Y_N\). Sample of size \(n\) observes \(y_1, y_2, \ldots, y_n\).
Note: \(\sigma^2 = \dfrac{N - 1}{N} S^2\).
At each draw, every unit has equal probability \(1/N\) of selection. Selected unit is replaced before next draw, so the same unit can appear multiple times.
Probability of any particular sample of size \(n\): \(1/N^n\).
Each unit can be selected at most once. Probability of any particular ordered sample is \(\dfrac{1}{N(N-1)\cdots(N-n+1)}\). Probability of any particular unordered sample of size \(n\) is \(\dfrac{1}{\binom{N}{n}}\).
| SRSWR | SRSWOR | |
|---|---|---|
| P(any unit drawn at draw r) | 1/N | 1/N |
| P(unit included in sample of size n) | 1 − (1 − 1/N)n | n/N |
| P(any specific sample) | 1/Nn | 1 / C(N, n) |
The textbook says that in SRSWOR “all draws are independent but not identical”. It is the other way round.
In SRSWR the draws are both independent and identically distributed: the unit is replaced, so every draw starts from the same population.
| SRSWOR | SRSWR | |
|---|---|---|
| Selected unit replaced? | No | Yes, before the next draw |
| Draws | dependent, identically distributed | independent and identically distributed |
| Number of possible samples | \(\binom{N}{n}\) (unordered) | \(N^n\) (ordered) |
| Probability of one sample | \(1/\binom{N}{n}\) | \(1/N^n\) |
| \(P(\text{unit } i \text{ at draw } r)\) | \(1/N\) (\(1/(N-r+1)\) given it is not yet drawn) | \(1/N\) |
| \(\bar y\) | unbiased for \(\bar Y\) | unbiased for \(\bar Y\) |
| \(s^2\) | unbiased for \(S^2\) | unbiased for \(\sigma^2\) |
| \(\text{Var}(\bar y)\) | \(\dfrac{N-n}{N}\cdot\dfrac{S^2}{n}\) | \(\dfrac{N-1}{N}\cdot\dfrac{S^2}{n} = \dfrac{\sigma^2}{n}\) |
Merits. (1) Personal bias is eliminated: every unit has the same chance, so the sample is more representative than a judgement sample. (2) The theory is complete: \(\bar y\) is unbiased, and its variance and standard error can be estimated from the sample itself.
Limitations. (1) It needs an up-to-date frame (a list of every unit), which often does not exist. (2) If the selected units are scattered, collecting the data is costly and slow. (3) By chance, some groups can be over-represented and others missed altogether (a middle-income group, say). (4) For the same precision it usually needs a larger sample than stratified sampling (Unit 3).
Other practical methods: pseudo-random numbers from calculators or computers, systematic random selection.
From 50 students, draw a sample of 5. Number them 1–50, write on chits, draw 5 chits without replacement.
For \(N = 200\) (3-digit numbers 001–200): from random number table read 145, 287 (skip), 092, 388 (skip), 050, 156, 173. Sample = {145, 92, 50, 156, 173}.
A random number table is a long list of digits in which each of \(0, 1, \ldots, 9\) occurs with the same frequency, independently of the others. For \(N \le 99\) the digits are read in pairs (00–99), for \(N \le 999\) in threes, and so on; a number above \(N\) (or 00, 000) is skipped, and in SRSWOR so is a repeat. The classical tables are:
Notes on the textbook. It spells Tippett as “Tippet” and A. J. Thompson as “A.S. Thomson”. In its lottery method it speaks of selecting “\(r\) items out of \(n\)”; in the notation of this page that is \(n\) units out of \(N\).
Properties:
\[ E(\bar y) = \bar Y, \qquad \text{Var}(\bar y) = \dfrac{\sigma^2}{n}. \](Sample mean is unbiased; variance ignores \(N\) because units are independent.)
The factor \((N - n)/N\) is the finite population correction (FPC).
So \(s^2\) is unbiased for \(S^2\).
\(\text{SE}(\bar y) = \sqrt{\text{Var}(\bar y)}\).
Let the population be \(Y_1, Y_2, \dots, Y_N\) with mean \(\bar Y = \dfrac{1}{N}\sum_{i=1}^{N} Y_i\), population variance \(\sigma^2 = \dfrac{1}{N}\sum_{i=1}^{N}(Y_i - \bar Y)^2\) and population mean square \(S^2 = \dfrac{1}{N-1}\sum_{i=1}^{N}(Y_i - \bar Y)^2\).
Draw an SRSWOR of size \(n\). For each population unit define the inclusion indicator
\[ a_i = \begin{cases} 1, & \text{if unit } i \text{ is in the sample},\\ 0, & \text{otherwise},\end{cases} \qquad \sum_{i=1}^{N} a_i = n, \qquad \bar y = \frac{1}{n}\sum_{i=1}^{N} a_i\, Y_i . \]Because every unit is equally likely to be included, and every pair equally likely to be included together,
\[ E(a_i) = \pi_i = \frac{n}{N}, \qquad E(a_i a_j) = \pi_{ij} = \frac{n(n-1)}{N(N-1)} \;\;(i \neq j). \](\(\pi_i\) follows since unit \(i\) lies in \(\binom{N-1}{n-1}\) of the \(\binom{N}{n}\) equally likely samples: \(\binom{N-1}{n-1}/\binom{N}{n} = n/N\); \(\pi_{ij}\) follows similarly.)
In simple random sampling (with or without replacement), the sample mean \(\bar y\) is an unbiased estimator of the population mean \(\bar Y\); that is, \(E(\bar y) = \bar Y\).
Proof (SRSWOR). Using \(\bar y = \dfrac{1}{n}\sum_{i=1}^{N} a_i Y_i\) and the linearity of expectation,
\[ E(\bar y) = \frac{1}{n}\sum_{i=1}^{N} Y_i\, E(a_i) = \frac{1}{n}\sum_{i=1}^{N} Y_i\cdot\frac{n}{N} = \frac{1}{N}\sum_{i=1}^{N} Y_i = \bar Y . \]Proof (SRSWR). Here \(y_1, \dots, y_n\) are independent and identically distributed, each taking any population value \(Y_i\) with probability \(1/N\), so \(E(y_j) = \dfrac{1}{N}\sum_{i=1}^{N} Y_i = \bar Y\). Hence \(E(\bar y) = \dfrac{1}{n}\sum_{j=1}^{n} E(y_j) = \bar Y\). \(\blacksquare\)
In SRSWOR of size \(n\) from \(N\) units,
\[ \text{Var}(\bar y) = \frac{N-n}{N\,n}\, S^2 = \frac{S^2}{n}\left(1 - \frac{n}{N}\right), \]where \(\dfrac{N-n}{N}\) is the finite population correction. In SRSWR, \(\text{Var}(\bar y) = \dfrac{\sigma^2}{n}\).
Proof (SRSWOR). From the indicators,
\[ \text{Var}(a_i) = \pi_i(1-\pi_i) = \frac{n}{N}\cdot\frac{N-n}{N}, \qquad \text{Cov}(a_i,a_j) = \pi_{ij} - \pi_i\pi_j = -\frac{n(N-n)}{N^{2}(N-1)}\;\;(i\neq j). \]Therefore
\[ \text{Var}(\bar y) = \frac{1}{n^{2}}\!\left[\sum_{i=1}^{N} Y_i^{2}\,\text{Var}(a_i) + \sum_{i\neq j} Y_i Y_j\,\text{Cov}(a_i,a_j)\right] = \frac{N-n}{n\,N^{2}}\!\left[\sum_i Y_i^{2} - \frac{1}{N-1}\sum_{i\neq j} Y_i Y_j\right]. \]Writing \(\sum_{i\neq j} Y_i Y_j = \big(\sum_i Y_i\big)^{2} - \sum_i Y_i^{2}\) and simplifying,
\[ \sum_i Y_i^{2} - \frac{1}{N-1}\!\left[\Big(\textstyle\sum_i Y_i\Big)^{2} - \sum_i Y_i^{2}\right] = \frac{N\sum_i Y_i^{2} - \big(\sum_i Y_i\big)^{2}}{N-1} = \frac{N\sum_i (Y_i-\bar Y)^{2}}{N-1} = N\,S^{2}, \]since \(N\sum Y_i^{2} - (\sum Y_i)^{2} = N\sum (Y_i-\bar Y)^{2} = N(N-1)S^{2}\). Substituting,
\[ \text{Var}(\bar y) = \frac{N-n}{n\,N^{2}}\cdot N\,S^{2} = \frac{N-n}{N\,n}\,S^{2}. \]Proof (SRSWR). The \(y_j\) are i.i.d. with common variance \(\text{Var}(y_j) = \dfrac{1}{N}\sum_{i=1}^{N}(Y_i-\bar Y)^2 = \sigma^{2}\); by independence \(\text{Var}(\bar y) = \dfrac{1}{n^{2}}\sum_{j=1}^{n}\text{Var}(y_j) = \dfrac{\sigma^{2}}{n}\). \(\blacksquare\)
For \(N = 6\), \(Y = 4, 6, 8, 10, 12, 14\), \(n = 2\): \(\bar Y = 9\), \(S^2 = 14\). Averaging \(\bar y\) over all \(\binom{6}{2} = 15\) samples gives \(E(\bar y) = 9 = \bar Y\) (Theorem 1), and the variance of those 15 sample means is \(14/3 = 4.67\), exactly \(\dfrac{N-n}{N\,n}S^2 = \dfrac{4}{12}\cdot 14 = \dfrac{14}{3}\) (Theorem 2).
For \(n > 1\): \(\text{Var}_{WOR}(\bar y) < \text{Var}_{WR}(\bar y)\) — SRSWOR is more efficient than SRSWR.
So when \(n/N\) is large (high sampling fraction), SRSWOR is markedly better.
\(N = 6\) units with values \(Y = 4, 6, 8, 10, 12, 14\). \(\bar Y = 9\); deviations \(\pm5, \pm3, \pm1\), so \(\sum(Y_i - \bar Y)^2 = 25 + 9 + 1 + 1 + 9 + 25 = 70\). Hence \(\sigma^2 = 70/6 = 11.67,\; S^2 = 70/5 = 14\).
SRSWR with \(n = 2\): \(\text{Var}(\bar y) = 11.67/2 = 5.83\). SRSWOR with \(n = 2\): \((6-2)/(6 \cdot 2) \cdot 14 = (4/12) \cdot 14 = 4.67\). WOR is more precise.
\(N = 100\) farms, sample of 10 yields \(\bar y = 250\) kg. \(\hat Y = 100 \cdot 250 = 25\,000\) kg. If sample SD = 30, \(\text{SE}(\bar y) = \sqrt{(90/100)(30^2/10)} = \sqrt{81} = 9\); SE(total) = 100 × 9 = 900 kg.
Expand the square in \(s^2\) and separate the cross products:
\[ s^2 = \frac{1}{n-1}\Big[\sum_{i=1}^n y_i^2 - n\bar y^2\Big] = \frac1n\sum_{i=1}^n y_i^2 - \frac{1}{n(n-1)}\sum_{i\ne j} y_i y_j . \]With the inclusion indicators, \(E\big(\sum y_i^2\big) = \sum_{i=1}^N E(a_i)Y_i^2 = \frac nN\sum Y_i^2\) and \(E\big(\sum_{i\ne j} y_iy_j\big) = \sum_{i\ne j}E(a_ia_j)Y_iY_j = \frac{n(n-1)}{N(N-1)}\sum_{i\ne j}Y_iY_j\). Substituting,
\[ E(s^2) = \frac1N\sum_{i=1}^N Y_i^2 - \frac{1}{N(N-1)}\sum_{i\ne j}Y_iY_j , \]which is the same expression as \(s^2\) with every small letter replaced by a capital and \(n\) by \(N\): it is \(S^2\).
Here \(E(y_i) = \bar Y\), \(\text{Var}(y_i) = \sigma^2\), so \(E(y_i^2) = \sigma^2 + \bar Y^2\), and \(E(\bar y^2) = \sigma^2/n + \bar Y^2\). Hence
\[ E(s^2) = \frac{1}{n-1}\Big[n(\sigma^2 + \bar Y^2) - n\Big(\frac{\sigma^2}{n} + \bar Y^2\Big)\Big] = \frac{1}{n-1}(n-1)\sigma^2 = \sigma^2 . \]Since \((N-1)S^2 = N\sigma^2\): in SRSWOR \(E(s^2) = S^2 = \frac{N}{N-1}\sigma^2\), so \(\hat\sigma^2 = \frac{N-1}{N}s^2\); in SRSWR \(\hat\sigma^2 = s^2\). The estimators of the mean (\(\bar y\)), the total (\(N\bar y\)) and the proportion (\(p\)) are the same under both schemes; only their variances, and the estimator of \(\sigma^2\), differ.
When is SRSWOR better? \(\text{Var}_{WOR}/\text{Var}_{WR} = (N-n)/(N-1)\), which is below 1 exactly when \(n > 1\); at \(n = 1\) the two schemes are the same thing.
To estimate the mean within a margin \(d\) with confidence \((1-\alpha)\):
For finite populations (apply FPC):
\[ n \;=\; \dfrac{n_0}{1 + n_0/N}, \quad n_0 = (z_{\alpha/2}\sigma/d)^2. \]For estimating a proportion within margin \(d\):
Take \(p = 0.5\) for the most conservative (largest) sample size.
Estimate population mean with σ ≈ 10 to within ±2 at 95 %: \(n = (1.96 \cdot 10/2)^2 = 96\) (round up to 97). With \(N = 1000\): adjusted \(n = 97/(1 + 97/1000) = 88.4\), rounded up to 89 — a sample size is always rounded up, or the precision asked for is not met.
Estimate proportion to within ±0.05 at 95 %: \(n = 1.96^2 \cdot 0.25/0.0025 = 384.16\) ⇒ 385.
Ask that \(\bar y\) fall within \(d\) of \(\bar Y\) with probability \(1-\alpha\): \(P(|\bar y - \bar Y| \ge d) = \alpha\). For large \(n\), \(\bar y\) is approximately normal with standard error \(S\sqrt{1/n - 1/N}\), so
\[ d = z_{\alpha/2}\,S\sqrt{\frac1n - \frac1N} \;\Longrightarrow\; \frac1n = \frac{d^2}{z_{\alpha/2}^2S^2} + \frac1N \;\Longrightarrow\; n = \frac{N z_{\alpha/2}^2 S^2}{N d^2 + z_{\alpha/2}^2S^2} . \]Dividing through by \(Nd^2\) gives the form above, \(n = n_0/(1 + n_0/N)\) with \(n_0 = z_{\alpha/2}^2S^2/d^2\). At 95%, \(z^2 = 1.96^2 = 3.8416\), which the textbook rounds to 3.84.
For small \(n\) the textbook replaces 1.96 by \(t_\alpha\) on \(n-1\) degrees of freedom and says that \((\bar y - \bar Y)/\big(S\sqrt{1/n - 1/N}\big)\) follows Student's \(t\). It does not: a \(t\) statistic has the sample \(s\) in its denominator (with the known \(S\) it is normal, not \(t\)), and the \(t\) distribution also needs a normal population. In practice \(S\) is replaced by \(s\) from a pilot survey, \(n = t_\alpha^2 s^2/(d^2 + t_\alpha^2 s^2/N)\); and since \(t_\alpha\) itself depends on \(n\), the formula is solved by trial: guess \(n\), look up \(t_\alpha\), recompute \(n\), and repeat until it settles.
Suppose \(M\) of the \(N\) population units possess an attribute (so population proportion \(P = M/N\)). Let \(m\) be the count in the sample and \(\hat p = m/n\).
where \(Q = 1 - P\).
Estimator of variance: \(\widehat{\text{Var}}(\hat p) = \dfrac{N - n}{N\, n}\cdot \dfrac{n \hat p \hat q}{n - 1} = \dfrac{(N - n)\hat p \hat q}{N(n - 1)}\).
From a population of 500 voters, an SRSWOR of 50 yields 30 in favour. \(\hat p = 0.6\). \(\widehat{\text{Var}} = (450/500)(0.6 \cdot 0.4/49) = 0.00441\). SE = 0.0664.
If \(N = 10\,000, n = 100, m = 18\): \(\hat p = 0.18\). With FPC ≈ 1, \(\widehat{\text{Var}} \approx 0.18(0.82)/99 = 0.00149\); SE = 0.0386.
Classify the units into \(A\) (possessing the attribute) and \(\alpha\) (not), and put \(Y_i = 1\) for \(A\), \(0\) for \(\alpha\). Then \(\sum Y_i = X\), the number in \(A\), and \(\bar Y = X/N = P\); likewise \(\bar y = x/n = p\). Because \(Y_i^2 = Y_i\),
\[ S^2 = \frac{1}{N-1}\Big[\sum Y_i^2 - N\bar Y^2\Big] = \frac{NP - NP^2}{N-1} = \frac{NPQ}{N-1}, \qquad s^2 = \frac{npq}{n-1} . \]So \(E(p) = P\) (under both schemes), \(Np\) is unbiased for the number \(X\), and \(\text{Var}_{WOR}(p) = \frac{N-n}{N}\cdot\frac{S^2}{n} = \frac{N-n}{N-1}\cdot\frac{PQ}{n}\).
When an auxiliary variable \(X\), correlated with the study variable \(Y\) and with known population mean \(\bar X\), is available, estimation can be sharpened using the ratio and regression methods.
The ratio estimator of the population mean is
It is (slightly) biased but often far more precise than \(\bar y\) when \(Y\) is roughly proportional to \(X\) (the regression line passes near the origin). Its approximate variance is
The ratio estimator beats the mean-per-unit estimator when the correlation \(\rho > \tfrac12\,(C_x/C_y)\) (ratio of coefficients of variation).
A sample of 5 gives \(y = 9,11,14,8,13\) and \(x = 10,12,15,9,14\), with known \(\bar X = 12.5\). Then \(\bar y = 11,\ \bar x = 12\), so \(\hat R = 11/12 = 0.917\) and \(\hat{\bar Y}_R = 0.917\times12.5 = 11.46\).
The linear regression estimator corrects \(\bar y\) using the deviation of the sample auxiliary mean from its known population value:
Its approximate variance is \(\text{Var}(\hat{\bar Y}_{lr}) \approx \dfrac{N-n}{Nn}\,S_y^2(1-\rho^2)\), so it is never worse than the mean-per-unit estimator and is at least as efficient as the ratio estimator. The ratio estimator is the special case \(b = \hat R\) (line through the origin).
For the data above, \(b = \dfrac{\sum(x-\bar x)(y-\bar y)}{\sum(x-\bar x)^2} = 1.0\), so \(\hat{\bar Y}_{lr} = 11 + 1.0\,(12.5 - 12) = 11.5\).
| Quantity | SRSWR | SRSWOR |
|---|---|---|
| \(E(\bar y)\) | \(\bar Y\) | \(\bar Y\) |
| \(\text{Var}(\bar y)\) | \(\sigma^2/n\) | \((N-n)S^2/(Nn)\) |
| \(E(s^2)\) | \(\sigma^2\) | \(S^2\) |
| \(\text{Var}(\hat p)\) | \(PQ/n\) | \((N-n)PQ/[(N-1)n]\) |
Three problems in the textbook's order: a sample drawn from Tippett's random numbers, and two small populations in which every possible sample is listed, so that unbiasedness and the variance formulas can be seen to hold exactly rather than taken on trust. The exercises follow, with their answers checked.
Source note. Every figure below was recomputed exactly, as a fraction, by listing all the samples. The textbook's answers agree, apart from small rounding differences in Worked Problem 3, which come from rounding \(\bar Y = 10/3\) to 3.33 before squaring; these are noted where they occur.
Draw a random sample of size 10 from a population of 400 units without replacement.
Number the units 001 to 400. Since \(N\) has three digits, read the random number table in groups of three, skipping any number above 400 (and 000) and, because sampling is without replacement, any number already drawn. Take the first 30 four-digit numbers of Tippett's table:
| 2952 | 6641 | 3992 | 9792 | 7969 | 5911 |
| 4167 | 9524 | 1545 | 1396 | 7203 | 5356 |
| 2370 | 7483 | 3408 | 2762 | 3563 | 1089 |
| 0560 | 5246 | 0112 | 6107 | 6008 | 8126 |
| 2754 | 9143 | 1405 | 9025 | 7002 | 6111 |
Run the digits together row by row and cut them into threes: 295, 266, 413 (skip), 992 (skip), 979 (skip), 279, 695 (skip), 911 (skip), 416 (skip), 795 (skip), 241, 545 (skip), 139, 672 (skip), 035, 356, 237, 074, 833 (skip), 408 (skip), 276. The sample is
295, 266, 279, 241, 139, 35, 356, 237, 74, 276.
Any starting point and any direction (rows, columns, diagonals) is allowed, provided it is fixed before the numbers are read.
A population has the 5 values 1, 2, 3, 6, 8. Write all samples of size 2 drawn without replacement, and verify that (i) \(\bar y\) is unbiased for \(\bar Y\); (ii) the variance of the sample means equals the formula for \(\text{Var}(\bar y)\); (iii) \(\text{Var}_{WOR}(\bar y) < \text{Var}_{WR}(\bar y)\); (iv) \(s^2\) is unbiased for \(S^2\).
Population constants. \(\sum Y_i = 20\), \(\sum Y_i^2 = 1 + 4 + 9 + 36 + 64 = 114\), so
\[ \bar Y = \frac{20}{5} = 4, \qquad S^2 = \frac{114 - 5(4^2)}{4} = \frac{34}{4} = 8.5, \qquad \sigma^2 = \frac{114}{5} - 4^2 = 6.8 . \]All \(\binom52 = 10\) samples:
| Sample | Values | \(\bar y\) | \(s^2\) | \((\bar y - \bar Y)^2\) |
|---|---|---|---|---|
| 1 | (1, 2) | 1.5 | 0.5 | 6.25 |
| 2 | (1, 3) | 2 | 2 | 4 |
| 3 | (1, 6) | 3.5 | 12.5 | 0.25 |
| 4 | (1, 8) | 4.5 | 24.5 | 0.25 |
| 5 | (2, 3) | 2.5 | 0.5 | 2.25 |
| 6 | (2, 6) | 4 | 8 | 0 |
| 7 | (2, 8) | 5 | 18 | 1 |
| 8 | (3, 6) | 4.5 | 4.5 | 0.25 |
| 9 | (3, 8) | 5.5 | 12.5 | 2.25 |
| 10 | (6, 8) | 7 | 2 | 9 |
| Total | 40 | 85 | 25.5 |
(i) \(E(\bar y) = 40/10 = 4 = \bar Y\).
(ii) The variance of the 10 sample means is \(25.5/10 = 2.55\); the formula gives \(\dfrac{N-n}{N}\cdot\dfrac{S^2}{n} = \dfrac35\cdot\dfrac{8.5}{2} = 2.55\).
(iii) \(\text{Var}_{WR}(\bar y) = \dfrac{N-1}{N}\cdot\dfrac{S^2}{n} = \dfrac45\cdot\dfrac{8.5}{2} = 3.4 > 2.55\). (Listing all 25 ordered with-replacement samples gives exactly 3.4 too.) The ratio is \(3.4/2.55 = 4/3 = (N-1)/(N-n)\).
(iv) \(E(s^2) = 85/10 = 8.5 = S^2\).
A population has the values 2, 3, 5. Considering all samples of size 2 drawn with replacement, verify that (i) \(\bar y\) is unbiased for \(\bar Y\); (ii) \(s^2\) is unbiased for \(\sigma^2\); (iii) the variance of the sample means equals \(\text{Var}(\bar y)\).
\(\sum Y_i = 10\), \(\sum Y_i^2 = 38\), \(N = 3\):
\[ \bar Y = \frac{10}{3} = 3.33, \qquad \sigma^2 = \frac{38}{3} - \Big(\frac{10}{3}\Big)^2 = \frac{14}{9} = 1.56, \] \[ S^2 = \frac{38 - 3(10/3)^2}{2} = \frac73 = 2.33 . \]With replacement there are \(N^n = 3^2 = 9\) ordered samples:
| Sample | Values | \(\bar y\) | \(s^2\) | \((\bar y - \bar Y)^2\) |
|---|---|---|---|---|
| 1 | (2, 2) | 2 | 0 | 1.778 |
| 2 | (2, 3) | 2.5 | 0.5 | 0.694 |
| 3 | (2, 5) | 3.5 | 4.5 | 0.028 |
| 4 | (3, 2) | 2.5 | 0.5 | 0.694 |
| 5 | (3, 3) | 3 | 0 | 0.111 |
| 6 | (3, 5) | 4 | 2 | 0.444 |
| 7 | (5, 2) | 3.5 | 4.5 | 0.028 |
| 8 | (5, 3) | 4 | 2 | 0.444 |
| 9 | (5, 5) | 5 | 0 | 2.778 |
| Total | 30 | 14 | 7 |
(i) \(E(\bar y) = 30/9 = 10/3 = \bar Y\). (ii) \(E(s^2) = 14/9 = 1.56 = \sigma^2\). (iii) The variance of the sample means is \(7/9 = 0.78\), and \(\text{Var}(\bar y) = \dfrac{\sigma^2}{n} = \dfrac{14/9}{2} = \dfrac79\), or equally \(\dfrac{N-1}{N}\cdot\dfrac{S^2}{n} = \dfrac23\cdot\dfrac{7/3}{2} = \dfrac79\).
Rounding note. The textbook rounds \(\bar Y\) to 3.33 before squaring, which gives \(S^2 = 2.36\) (exactly \(7/3 = 2.33\)) and a total of 7.0001 in the last column (exactly 7). It also introduces the 9 samples with “in srswor”, where with-replacement sampling is meant.
Additional worked problems with step-by-step procedures to support self-study, matching this unit's topics.
Population of 200 students numbered 1–200; select 5 at random. Because the population size is a 3-digit number, read the first 3 digits of each 5-digit entry in the random-number table; discard any > 200 and any repeats. Starting at one entry and moving down the column gives 200, 023, 108, 070, 126 → students 23, 70, 108, 126, 200. Each unit had an equal, independent chance of selection.
Diseased plants in 24 areas: 1, 4, 1, 2, 5, 1, 1, 1, 7, 2, 3, 3, 2, 2, 3, 1, 2, 7, 2, 6, 3, 5, 3, 4. Select a sample of size 6 by SRSWR and SRSWOR; compare with the population mean.
Population mean \(= \dfrac{\sum y_i}{24} = \dfrac{71}{24} = 2.96\).
SRSWR (2-digit reading; repeats allowed) → areas 02, 10, 11, 12, 17, 17 with values 4, 2, 3, 3, 2, 2 → sample mean \(= \dfrac{16}{6} = 2.67\).
SRSWOR (repeats not allowed) → areas 01, 04, 12, 19, 20, 22 with values 1, 2, 3, 2, 6, 5 → sample mean \(= \dfrac{19}{6} = 3.17\).
Both sample means (2.67, 3.17) are close to the population mean (2.96) — illustrating that the sample mean estimates the population mean.
Population units 1, 2, 3, 4, 5; draw all samples of size \(n=3\) under SRSWOR.
Number of samples \(= \binom{5}{3} = 10\). Population mean \(\bar Y = 15/5 = 3\). The 10 sample means total 30.0, so \(E(\bar y) = \dfrac{\sum \bar y}{\binom{5}{3}} = \dfrac{30}{10} = 3 = \bar Y\).
Hence the sample mean is an unbiased estimate of the population mean under SRSWOR.
Population units 1, 2, 3, 4, 5; draw all samples of size \(n=2\) under SRSWR.
Number of samples \(= N^n = 5^2 = 25\). The 25 sample means total 75.0, so \(E(\bar y) = \dfrac{75}{25} = 3 = \bar Y\). The sample mean is unbiased under SRSWR as well.