Skip to the content

Topics Covered

SRSWR SRSWOR Notation Lottery Method Random Number Tables Estimators & Variance Sample Size SRS for Attributes Unbiasedness of s² Enumerating All Samples
On this page
  1. 1. Notation and Terminology
  2. 2. Two Schemes of SRS
  3. 3. Methods of Selecting an SRS
  4. 4. Estimators and their Variances
  5. 5. Determination of Sample Size
  6. 6. SRS for Attributes
  7. 7. Ratio Method of Estimation
  8. 8. Regression Method of Estimation
  9. Summary Formulas
  10. Worked Problems on Simple Random Sampling
  11. Key Take-aways
  12. Extra Practical Problems

1. Notation and Terminology

Population: \(U_1, U_2, \ldots, U_N\); values \(Y_1, Y_2, \ldots, Y_N\). Sample of size \(n\) observes \(y_1, y_2, \ldots, y_n\).

Note: \(\sigma^2 = \dfrac{N - 1}{N} S^2\).

2. Two Schemes of SRS

2.1 SRS with Replacement (SRSWR)

At each draw, every unit has equal probability \(1/N\) of selection. Selected unit is replaced before next draw, so the same unit can appear multiple times.

Probability of any particular sample of size \(n\): \(1/N^n\).

2.2 SRS without Replacement (SRSWOR)

Each unit can be selected at most once. Probability of any particular ordered sample is \(\dfrac{1}{N(N-1)\cdots(N-n+1)}\). Probability of any particular unordered sample of size \(n\) is \(\dfrac{1}{\binom{N}{n}}\).

2.3 Probabilities of Selection

SRSWRSRSWOR
P(any unit drawn at draw r)1/N1/N
P(unit included in sample of size n)1 − (1 − 1/N)nn/N
P(any specific sample)1/Nn1 / C(N, n)

Draws in SRSWOR: Dependent but Identically Distributed

CORRECTION — WHICH DRAWS ARE INDEPENDENT

The textbook says that in SRSWOR “all draws are independent but not identical”. It is the other way round.

In SRSWR the draws are both independent and identically distributed: the unit is replaced, so every draw starts from the same population.

SRSWORSRSWR
Selected unit replaced?NoYes, before the next draw
Drawsdependent, identically distributedindependent and identically distributed
Number of possible samples\(\binom{N}{n}\) (unordered)\(N^n\) (ordered)
Probability of one sample\(1/\binom{N}{n}\)\(1/N^n\)
\(P(\text{unit } i \text{ at draw } r)\)\(1/N\) (\(1/(N-r+1)\) given it is not yet drawn)\(1/N\)
\(\bar y\)unbiased for \(\bar Y\)unbiased for \(\bar Y\)
\(s^2\)unbiased for \(S^2\)unbiased for \(\sigma^2\)
\(\text{Var}(\bar y)\)\(\dfrac{N-n}{N}\cdot\dfrac{S^2}{n}\)\(\dfrac{N-1}{N}\cdot\dfrac{S^2}{n} = \dfrac{\sigma^2}{n}\)

Merits and Limitations of SRS

Merits. (1) Personal bias is eliminated: every unit has the same chance, so the sample is more representative than a judgement sample. (2) The theory is complete: \(\bar y\) is unbiased, and its variance and standard error can be estimated from the sample itself.

Limitations. (1) It needs an up-to-date frame (a list of every unit), which often does not exist. (2) If the selected units are scattered, collecting the data is costly and slow. (3) By chance, some groups can be over-represented and others missed altogether (a middle-income group, say). (4) For the same precision it usually needs a larger sample than stratified sampling (Unit 3).

3. Methods of Selecting an SRS

3.1 Lottery Method

  1. Number all units 1 to \(N\).
  2. Write each number on identical chits, place in a container, mix thoroughly.
  3. Draw \(n\) chits — with or without replacement as required.

3.2 Random Number Tables

  1. Number units 1 to \(N\). If \(N = 850\), use 3-digit groups.
  2. Pick any starting point in the table; read groups of digits in any direction.
  3. Accept numbers ≤ \(N\); discard the rest. Repeat in WOR if duplicate.

Other practical methods: pseudo-random numbers from calculators or computers, systematic random selection.

EXAMPLE 1 (Lottery)

From 50 students, draw a sample of 5. Number them 1–50, write on chits, draw 5 chits without replacement.

EXAMPLE 2 (Random number table)

For \(N = 200\) (3-digit numbers 001–200): from random number table read 145, 287 (skip), 092, 388 (skip), 050, 156, 173. Sample = {145, 92, 50, 156, 173}.

Random Number Tables in Print

A random number table is a long list of digits in which each of \(0, 1, \ldots, 9\) occurs with the same frequency, independently of the others. For \(N \le 99\) the digits are read in pairs (00–99), for \(N \le 999\) in threes, and so on; a number above \(N\) (or 00, 000) is skipped, and in SRSWOR so is a repeat. The classical tables are:

  1. Tippett (1927): 10,400 four-digit numbers, i.e. 41,600 digits, taken from British census reports.
  2. Fisher and Yates (1938), in Statistical Tables for Biological, Agricultural and Medical Research: 15,000 digits arranged in twos, drawn from the later digits of A. J. Thompson's 20-figure logarithm tables.
  3. Kendall and Babington Smith (1939): 1,00,000 digits in 25,000 sets of four.
  4. RAND Corporation (1955), A Million Random Digits: 10,00,000 digits, i.e. 2,00,000 five-digit numbers.

Notes on the textbook. It spells Tippett as “Tippet” and A. J. Thompson as “A.S. Thomson”. In its lottery method it speaks of selecting “\(r\) items out of \(n\)”; in the notation of this page that is \(n\) units out of \(N\).

4. Estimators and their Variances

4.1 SRSWR — Sample Mean

\[ \bar y \;=\; \dfrac{1}{n}\sum y_i. \]

Properties:

\[ E(\bar y) = \bar Y, \qquad \text{Var}(\bar y) = \dfrac{\sigma^2}{n}. \]

(Sample mean is unbiased; variance ignores \(N\) because units are independent.)

4.2 SRSWOR — Sample Mean

\[ E(\bar y) = \bar Y, \qquad \text{Var}(\bar y) = \dfrac{N - n}{N\, n}\, S^2 = \dfrac{S^2}{n}\left(1 - \dfrac{n}{N}\right). \]

The factor \((N - n)/N\) is the finite population correction (FPC).

4.3 Sample Mean Square (SRSWOR)

\[ E(s^2) \;=\; S^2. \]

So \(s^2\) is unbiased for \(S^2\).

4.4 Estimator of Population Total

\[ \hat Y = N \bar y, \qquad \text{Var}(\hat Y) = N^2\, \text{Var}(\bar y). \]

4.5 Estimator of Variance of \(\bar y\) (SRSWOR)

\[ \widehat{\text{Var}}(\bar y) \;=\; \dfrac{N - n}{N\, n}\, s^2. \]

4.6 Standard Error

\(\text{SE}(\bar y) = \sqrt{\text{Var}(\bar y)}\).

4.7 Theorems — Unbiasedness and Variance of the Sample Mean (with Proofs)

SET-UP (INCLUSION-INDICATOR METHOD)

Let the population be \(Y_1, Y_2, \dots, Y_N\) with mean \(\bar Y = \dfrac{1}{N}\sum_{i=1}^{N} Y_i\), population variance \(\sigma^2 = \dfrac{1}{N}\sum_{i=1}^{N}(Y_i - \bar Y)^2\) and population mean square \(S^2 = \dfrac{1}{N-1}\sum_{i=1}^{N}(Y_i - \bar Y)^2\).

Draw an SRSWOR of size \(n\). For each population unit define the inclusion indicator

\[ a_i = \begin{cases} 1, & \text{if unit } i \text{ is in the sample},\\ 0, & \text{otherwise},\end{cases} \qquad \sum_{i=1}^{N} a_i = n, \qquad \bar y = \frac{1}{n}\sum_{i=1}^{N} a_i\, Y_i . \]

Because every unit is equally likely to be included, and every pair equally likely to be included together,

\[ E(a_i) = \pi_i = \frac{n}{N}, \qquad E(a_i a_j) = \pi_{ij} = \frac{n(n-1)}{N(N-1)} \;\;(i \neq j). \]

(\(\pi_i\) follows since unit \(i\) lies in \(\binom{N-1}{n-1}\) of the \(\binom{N}{n}\) equally likely samples: \(\binom{N-1}{n-1}/\binom{N}{n} = n/N\); \(\pi_{ij}\) follows similarly.)

THEOREM 1 — UNBIASEDNESS

In simple random sampling (with or without replacement), the sample mean \(\bar y\) is an unbiased estimator of the population mean \(\bar Y\); that is, \(E(\bar y) = \bar Y\).

Proof (SRSWOR). Using \(\bar y = \dfrac{1}{n}\sum_{i=1}^{N} a_i Y_i\) and the linearity of expectation,

\[ E(\bar y) = \frac{1}{n}\sum_{i=1}^{N} Y_i\, E(a_i) = \frac{1}{n}\sum_{i=1}^{N} Y_i\cdot\frac{n}{N} = \frac{1}{N}\sum_{i=1}^{N} Y_i = \bar Y . \]

Proof (SRSWR). Here \(y_1, \dots, y_n\) are independent and identically distributed, each taking any population value \(Y_i\) with probability \(1/N\), so \(E(y_j) = \dfrac{1}{N}\sum_{i=1}^{N} Y_i = \bar Y\). Hence \(E(\bar y) = \dfrac{1}{n}\sum_{j=1}^{n} E(y_j) = \bar Y\). \(\blacksquare\)

THEOREM 2 — VARIANCE OF THE SAMPLE MEAN

In SRSWOR of size \(n\) from \(N\) units,

\[ \text{Var}(\bar y) = \frac{N-n}{N\,n}\, S^2 = \frac{S^2}{n}\left(1 - \frac{n}{N}\right), \]

where \(\dfrac{N-n}{N}\) is the finite population correction. In SRSWR, \(\text{Var}(\bar y) = \dfrac{\sigma^2}{n}\).

Proof (SRSWOR). From the indicators,

\[ \text{Var}(a_i) = \pi_i(1-\pi_i) = \frac{n}{N}\cdot\frac{N-n}{N}, \qquad \text{Cov}(a_i,a_j) = \pi_{ij} - \pi_i\pi_j = -\frac{n(N-n)}{N^{2}(N-1)}\;\;(i\neq j). \]

Therefore

\[ \text{Var}(\bar y) = \frac{1}{n^{2}}\!\left[\sum_{i=1}^{N} Y_i^{2}\,\text{Var}(a_i) + \sum_{i\neq j} Y_i Y_j\,\text{Cov}(a_i,a_j)\right] = \frac{N-n}{n\,N^{2}}\!\left[\sum_i Y_i^{2} - \frac{1}{N-1}\sum_{i\neq j} Y_i Y_j\right]. \]

Writing \(\sum_{i\neq j} Y_i Y_j = \big(\sum_i Y_i\big)^{2} - \sum_i Y_i^{2}\) and simplifying,

\[ \sum_i Y_i^{2} - \frac{1}{N-1}\!\left[\Big(\textstyle\sum_i Y_i\Big)^{2} - \sum_i Y_i^{2}\right] = \frac{N\sum_i Y_i^{2} - \big(\sum_i Y_i\big)^{2}}{N-1} = \frac{N\sum_i (Y_i-\bar Y)^{2}}{N-1} = N\,S^{2}, \]

since \(N\sum Y_i^{2} - (\sum Y_i)^{2} = N\sum (Y_i-\bar Y)^{2} = N(N-1)S^{2}\). Substituting,

\[ \text{Var}(\bar y) = \frac{N-n}{n\,N^{2}}\cdot N\,S^{2} = \frac{N-n}{N\,n}\,S^{2}. \]

Proof (SRSWR). The \(y_j\) are i.i.d. with common variance \(\text{Var}(y_j) = \dfrac{1}{N}\sum_{i=1}^{N}(Y_i-\bar Y)^2 = \sigma^{2}\); by independence \(\text{Var}(\bar y) = \dfrac{1}{n^{2}}\sum_{j=1}^{n}\text{Var}(y_j) = \dfrac{\sigma^{2}}{n}\). \(\blacksquare\)

NUMERICAL CHECK

For \(N = 6\), \(Y = 4, 6, 8, 10, 12, 14\), \(n = 2\): \(\bar Y = 9\), \(S^2 = 14\). Averaging \(\bar y\) over all \(\binom{6}{2} = 15\) samples gives \(E(\bar y) = 9 = \bar Y\) (Theorem 1), and the variance of those 15 sample means is \(14/3 = 4.67\), exactly \(\dfrac{N-n}{N\,n}S^2 = \dfrac{4}{12}\cdot 14 = \dfrac{14}{3}\) (Theorem 2).

SRSWR vs SRSWOR — Variance Comparison

For \(n > 1\): \(\text{Var}_{WOR}(\bar y) < \text{Var}_{WR}(\bar y)\) — SRSWOR is more efficient than SRSWR.

\[ \dfrac{\text{Var}_{WOR}}{\text{Var}_{WR}} = \dfrac{(N-n)S^2/(Nn)}{\sigma^2/n} = \dfrac{N - n}{N - 1} \approx \dfrac{N - n}{N} = 1 - f. \]

So when \(n/N\) is large (high sampling fraction), SRSWOR is markedly better.

EXAMPLE 1 (Population mean)

\(N = 6\) units with values \(Y = 4, 6, 8, 10, 12, 14\). \(\bar Y = 9\); deviations \(\pm5, \pm3, \pm1\), so \(\sum(Y_i - \bar Y)^2 = 25 + 9 + 1 + 1 + 9 + 25 = 70\). Hence \(\sigma^2 = 70/6 = 11.67,\; S^2 = 70/5 = 14\).

SRSWR with \(n = 2\): \(\text{Var}(\bar y) = 11.67/2 = 5.83\). SRSWOR with \(n = 2\): \((6-2)/(6 \cdot 2) \cdot 14 = (4/12) \cdot 14 = 4.67\). WOR is more precise.

EXAMPLE 2 (Total)

\(N = 100\) farms, sample of 10 yields \(\bar y = 250\) kg. \(\hat Y = 100 \cdot 250 = 25\,000\) kg. If sample SD = 30, \(\text{SE}(\bar y) = \sqrt{(90/100)(30^2/10)} = \sqrt{81} = 9\); SE(total) = 100 × 9 = 900 kg.

Proofs that \(s^2\) Is Unbiased

SRSWOR: \(E(s^2) = S^2\)

Expand the square in \(s^2\) and separate the cross products:

\[ s^2 = \frac{1}{n-1}\Big[\sum_{i=1}^n y_i^2 - n\bar y^2\Big] = \frac1n\sum_{i=1}^n y_i^2 - \frac{1}{n(n-1)}\sum_{i\ne j} y_i y_j . \]

With the inclusion indicators, \(E\big(\sum y_i^2\big) = \sum_{i=1}^N E(a_i)Y_i^2 = \frac nN\sum Y_i^2\) and \(E\big(\sum_{i\ne j} y_iy_j\big) = \sum_{i\ne j}E(a_ia_j)Y_iY_j = \frac{n(n-1)}{N(N-1)}\sum_{i\ne j}Y_iY_j\). Substituting,

\[ E(s^2) = \frac1N\sum_{i=1}^N Y_i^2 - \frac{1}{N(N-1)}\sum_{i\ne j}Y_iY_j , \]

which is the same expression as \(s^2\) with every small letter replaced by a capital and \(n\) by \(N\): it is \(S^2\).

SRSWR: \(E(s^2) = \sigma^2\)

Here \(E(y_i) = \bar Y\), \(\text{Var}(y_i) = \sigma^2\), so \(E(y_i^2) = \sigma^2 + \bar Y^2\), and \(E(\bar y^2) = \sigma^2/n + \bar Y^2\). Hence

\[ E(s^2) = \frac{1}{n-1}\Big[n(\sigma^2 + \bar Y^2) - n\Big(\frac{\sigma^2}{n} + \bar Y^2\Big)\Big] = \frac{1}{n-1}(n-1)\sigma^2 = \sigma^2 . \]
ESTIMATING THE POPULATION VARIANCE \(\sigma^2\)

Since \((N-1)S^2 = N\sigma^2\): in SRSWOR \(E(s^2) = S^2 = \frac{N}{N-1}\sigma^2\), so \(\hat\sigma^2 = \frac{N-1}{N}s^2\); in SRSWR \(\hat\sigma^2 = s^2\). The estimators of the mean (\(\bar y\)), the total (\(N\bar y\)) and the proportion (\(p\)) are the same under both schemes; only their variances, and the estimator of \(\sigma^2\), differ.

When is SRSWOR better? \(\text{Var}_{WOR}/\text{Var}_{WR} = (N-n)/(N-1)\), which is below 1 exactly when \(n > 1\); at \(n = 1\) the two schemes are the same thing.

5. Determination of Sample Size

To estimate the mean within a margin \(d\) with confidence \((1-\alpha)\):

\[ n \;\ge\; \left(\dfrac{z_{\alpha/2}\, \sigma}{d}\right)^2. \]

For finite populations (apply FPC):

\[ n \;=\; \dfrac{n_0}{1 + n_0/N}, \quad n_0 = (z_{\alpha/2}\sigma/d)^2. \]

For estimating a proportion within margin \(d\):

\[ n \;=\; \dfrac{z_{\alpha/2}^2\, p(1-p)}{d^2}. \]

Take \(p = 0.5\) for the most conservative (largest) sample size.

EXAMPLE 1

Estimate population mean with σ ≈ 10 to within ±2 at 95 %: \(n = (1.96 \cdot 10/2)^2 = 96\) (round up to 97). With \(N = 1000\): adjusted \(n = 97/(1 + 97/1000) = 88.4\), rounded up to 89 — a sample size is always rounded up, or the precision asked for is not met.

EXAMPLE 2

Estimate proportion to within ±0.05 at 95 %: \(n = 1.96^2 \cdot 0.25/0.0025 = 384.16\) ⇒ 385.

Deriving the Sample-Size Formula

LARGE \(n\), \(S\) KNOWN

Ask that \(\bar y\) fall within \(d\) of \(\bar Y\) with probability \(1-\alpha\): \(P(|\bar y - \bar Y| \ge d) = \alpha\). For large \(n\), \(\bar y\) is approximately normal with standard error \(S\sqrt{1/n - 1/N}\), so

\[ d = z_{\alpha/2}\,S\sqrt{\frac1n - \frac1N} \;\Longrightarrow\; \frac1n = \frac{d^2}{z_{\alpha/2}^2S^2} + \frac1N \;\Longrightarrow\; n = \frac{N z_{\alpha/2}^2 S^2}{N d^2 + z_{\alpha/2}^2S^2} . \]

Dividing through by \(Nd^2\) gives the form above, \(n = n_0/(1 + n_0/N)\) with \(n_0 = z_{\alpha/2}^2S^2/d^2\). At 95%, \(z^2 = 1.96^2 = 3.8416\), which the textbook rounds to 3.84.

CORRECTION — THE SMALL-SAMPLE VERSION

For small \(n\) the textbook replaces 1.96 by \(t_\alpha\) on \(n-1\) degrees of freedom and says that \((\bar y - \bar Y)/\big(S\sqrt{1/n - 1/N}\big)\) follows Student's \(t\). It does not: a \(t\) statistic has the sample \(s\) in its denominator (with the known \(S\) it is normal, not \(t\)), and the \(t\) distribution also needs a normal population. In practice \(S\) is replaced by \(s\) from a pilot survey, \(n = t_\alpha^2 s^2/(d^2 + t_\alpha^2 s^2/N)\); and since \(t_\alpha\) itself depends on \(n\), the formula is solved by trial: guess \(n\), look up \(t_\alpha\), recompute \(n\), and repeat until it settles.

6. SRS for Attributes

Suppose \(M\) of the \(N\) population units possess an attribute (so population proportion \(P = M/N\)). Let \(m\) be the count in the sample and \(\hat p = m/n\).

\[ E(\hat p) = P, \quad \text{Var}_{WR}(\hat p) = \dfrac{P Q}{n}, \quad \text{Var}_{WOR}(\hat p) = \dfrac{N - n}{N(n)} \cdot \dfrac{NPQ}{N - 1}, \]

where \(Q = 1 - P\).

Estimator of variance: \(\widehat{\text{Var}}(\hat p) = \dfrac{N - n}{N\, n}\cdot \dfrac{n \hat p \hat q}{n - 1} = \dfrac{(N - n)\hat p \hat q}{N(n - 1)}\).

EXAMPLE 1

From a population of 500 voters, an SRSWOR of 50 yields 30 in favour. \(\hat p = 0.6\). \(\widehat{\text{Var}} = (450/500)(0.6 \cdot 0.4/49) = 0.00441\). SE = 0.0664.

EXAMPLE 2

If \(N = 10\,000, n = 100, m = 18\): \(\hat p = 0.18\). With FPC ≈ 1, \(\widehat{\text{Var}} \approx 0.18(0.82)/99 = 0.00149\); SE = 0.0386.

Proportions as Means of an Indicator

EVERY RESULT FOR \(\bar y\) CARRIES OVER

Classify the units into \(A\) (possessing the attribute) and \(\alpha\) (not), and put \(Y_i = 1\) for \(A\), \(0\) for \(\alpha\). Then \(\sum Y_i = X\), the number in \(A\), and \(\bar Y = X/N = P\); likewise \(\bar y = x/n = p\). Because \(Y_i^2 = Y_i\),

\[ S^2 = \frac{1}{N-1}\Big[\sum Y_i^2 - N\bar Y^2\Big] = \frac{NP - NP^2}{N-1} = \frac{NPQ}{N-1}, \qquad s^2 = \frac{npq}{n-1} . \]

So \(E(p) = P\) (under both schemes), \(Np\) is unbiased for the number \(X\), and \(\text{Var}_{WOR}(p) = \frac{N-n}{N}\cdot\frac{S^2}{n} = \frac{N-n}{N-1}\cdot\frac{PQ}{n}\).

When an auxiliary variable \(X\), correlated with the study variable \(Y\) and with known population mean \(\bar X\), is available, estimation can be sharpened using the ratio and regression methods.

7. Ratio Method of Estimation

The ratio estimator of the population mean is

\[ \hat{\bar Y}_R = \dfrac{\bar y}{\bar x}\,\bar X = \hat R\,\bar X, \qquad \hat R = \dfrac{\bar y}{\bar x}. \]

It is (slightly) biased but often far more precise than \(\bar y\) when \(Y\) is roughly proportional to \(X\) (the regression line passes near the origin). Its approximate variance is

\[ \text{Var}(\hat{\bar Y}_R) \approx \dfrac{N-n}{Nn}\cdot\dfrac{1}{N-1} \sum_{i=1}^{N}\big(Y_i - R\,X_i\big)^2, \qquad R = \dfrac{\bar Y}{\bar X}. \]

The ratio estimator beats the mean-per-unit estimator when the correlation \(\rho > \tfrac12\,(C_x/C_y)\) (ratio of coefficients of variation).

EXAMPLE

A sample of 5 gives \(y = 9,11,14,8,13\) and \(x = 10,12,15,9,14\), with known \(\bar X = 12.5\). Then \(\bar y = 11,\ \bar x = 12\), so \(\hat R = 11/12 = 0.917\) and \(\hat{\bar Y}_R = 0.917\times12.5 = 11.46\).

8. Regression Method of Estimation

The linear regression estimator corrects \(\bar y\) using the deviation of the sample auxiliary mean from its known population value:

\[ \hat{\bar Y}_{lr} = \bar y + b\,(\bar X - \bar x), \qquad b = \dfrac{\sum (x_i-\bar x)(y_i-\bar y)}{\sum (x_i-\bar x)^2}. \]

Its approximate variance is \(\text{Var}(\hat{\bar Y}_{lr}) \approx \dfrac{N-n}{Nn}\,S_y^2(1-\rho^2)\), so it is never worse than the mean-per-unit estimator and is at least as efficient as the ratio estimator. The ratio estimator is the special case \(b = \hat R\) (line through the origin).

EXAMPLE

For the data above, \(b = \dfrac{\sum(x-\bar x)(y-\bar y)}{\sum(x-\bar x)^2} = 1.0\), so \(\hat{\bar Y}_{lr} = 11 + 1.0\,(12.5 - 12) = 11.5\).

Ratio vs Regression estimator 81012 1416 691215 Auxiliary variable x → Study variable y ratio (through origin) regression (best-fit)
Fig — The ratio estimator forces the line through the origin (slope \(\hat R = \bar y/\bar x\)), which is efficient only when \(y\) is roughly proportional to \(x\). The regression estimator fits the best line with a free intercept, so it is never less efficient — the two coincide only when the best-fit line already passes through the origin.

Summary Formulas

QuantitySRSWRSRSWOR
\(E(\bar y)\)\(\bar Y\)\(\bar Y\)
\(\text{Var}(\bar y)\)\(\sigma^2/n\)\((N-n)S^2/(Nn)\)
\(E(s^2)\)\(\sigma^2\)\(S^2\)
\(\text{Var}(\hat p)\)\(PQ/n\)\((N-n)PQ/[(N-1)n]\)

Worked Problems on Simple Random Sampling

Three problems in the textbook's order: a sample drawn from Tippett's random numbers, and two small populations in which every possible sample is listed, so that unbiasedness and the variance formulas can be seen to hold exactly rather than taken on trust. The exercises follow, with their answers checked.

Source note. Every figure below was recomputed exactly, as a fraction, by listing all the samples. The textbook's answers agree, apart from small rounding differences in Worked Problem 3, which come from rounding \(\bar Y = 10/3\) to 3.33 before squaring; these are noted where they occur.

A. Drawing a Sample

WORKED PROBLEM 1 — 10 units from 400, without replacement

Draw a random sample of size 10 from a population of 400 units without replacement.

Number the units 001 to 400. Since \(N\) has three digits, read the random number table in groups of three, skipping any number above 400 (and 000) and, because sampling is without replacement, any number already drawn. Take the first 30 four-digit numbers of Tippett's table:

295266413992979279695911
416795241545139672035356
237074833408276235631089
056052460112610760088126
275491431405902570026111

Run the digits together row by row and cut them into threes: 295, 266, 413 (skip), 992 (skip), 979 (skip), 279, 695 (skip), 911 (skip), 416 (skip), 795 (skip), 241, 545 (skip), 139, 672 (skip), 035, 356, 237, 074, 833 (skip), 408 (skip), 276. The sample is

295, 266, 279, 241, 139, 35, 356, 237, 74, 276.

Any starting point and any direction (rows, columns, diagonals) is allowed, provided it is fixed before the numbers are read.

B. Listing Every Sample

WORKED PROBLEM 2 — all SRSWOR samples of size 2 from 1, 2, 3, 6, 8

A population has the 5 values 1, 2, 3, 6, 8. Write all samples of size 2 drawn without replacement, and verify that (i) \(\bar y\) is unbiased for \(\bar Y\); (ii) the variance of the sample means equals the formula for \(\text{Var}(\bar y)\); (iii) \(\text{Var}_{WOR}(\bar y) < \text{Var}_{WR}(\bar y)\); (iv) \(s^2\) is unbiased for \(S^2\).

Population constants. \(\sum Y_i = 20\), \(\sum Y_i^2 = 1 + 4 + 9 + 36 + 64 = 114\), so

\[ \bar Y = \frac{20}{5} = 4, \qquad S^2 = \frac{114 - 5(4^2)}{4} = \frac{34}{4} = 8.5, \qquad \sigma^2 = \frac{114}{5} - 4^2 = 6.8 . \]

All \(\binom52 = 10\) samples:

SampleValues\(\bar y\)\(s^2\)\((\bar y - \bar Y)^2\)
1(1, 2)1.50.56.25
2(1, 3)224
3(1, 6)3.512.50.25
4(1, 8)4.524.50.25
5(2, 3)2.50.52.25
6(2, 6)480
7(2, 8)5181
8(3, 6)4.54.50.25
9(3, 8)5.512.52.25
10(6, 8)729
Total408525.5

(i) \(E(\bar y) = 40/10 = 4 = \bar Y\).

(ii) The variance of the 10 sample means is \(25.5/10 = 2.55\); the formula gives \(\dfrac{N-n}{N}\cdot\dfrac{S^2}{n} = \dfrac35\cdot\dfrac{8.5}{2} = 2.55\).

(iii) \(\text{Var}_{WR}(\bar y) = \dfrac{N-1}{N}\cdot\dfrac{S^2}{n} = \dfrac45\cdot\dfrac{8.5}{2} = 3.4 > 2.55\). (Listing all 25 ordered with-replacement samples gives exactly 3.4 too.) The ratio is \(3.4/2.55 = 4/3 = (N-1)/(N-n)\).

(iv) \(E(s^2) = 85/10 = 8.5 = S^2\).

1 2 3 4 5 6 7 8 WOR: 10 samples 1 2 3 4 5 6 7 8 WR: 25 samples mean of all sample means = 4 = Ṽ sample mean ȳ (n = 2, population 1, 2, 3, 6, 8)
Fig 2.1 — Worked Problem 2. Every possible sample mean, without replacement (top, variance \(2.55\)) and with replacement (bottom, variance \(3.4\)). Both distributions balance at \(\bar Y = 4\), which is what unbiasedness means; the with-replacement means reach further out, because a sample can repeat an extreme unit.
WORKED PROBLEM 3 — all SRSWR samples of size 2 from 2, 3, 5

A population has the values 2, 3, 5. Considering all samples of size 2 drawn with replacement, verify that (i) \(\bar y\) is unbiased for \(\bar Y\); (ii) \(s^2\) is unbiased for \(\sigma^2\); (iii) the variance of the sample means equals \(\text{Var}(\bar y)\).

\(\sum Y_i = 10\), \(\sum Y_i^2 = 38\), \(N = 3\):

\[ \bar Y = \frac{10}{3} = 3.33, \qquad \sigma^2 = \frac{38}{3} - \Big(\frac{10}{3}\Big)^2 = \frac{14}{9} = 1.56, \] \[ S^2 = \frac{38 - 3(10/3)^2}{2} = \frac73 = 2.33 . \]

With replacement there are \(N^n = 3^2 = 9\) ordered samples:

SampleValues\(\bar y\)\(s^2\)\((\bar y - \bar Y)^2\)
1(2, 2)201.778
2(2, 3)2.50.50.694
3(2, 5)3.54.50.028
4(3, 2)2.50.50.694
5(3, 3)300.111
6(3, 5)420.444
7(5, 2)3.54.50.028
8(5, 3)420.444
9(5, 5)502.778
Total30147

(i) \(E(\bar y) = 30/9 = 10/3 = \bar Y\). (ii) \(E(s^2) = 14/9 = 1.56 = \sigma^2\). (iii) The variance of the sample means is \(7/9 = 0.78\), and \(\text{Var}(\bar y) = \dfrac{\sigma^2}{n} = \dfrac{14/9}{2} = \dfrac79\), or equally \(\dfrac{N-1}{N}\cdot\dfrac{S^2}{n} = \dfrac23\cdot\dfrac{7/3}{2} = \dfrac79\).

Rounding note. The textbook rounds \(\bar Y\) to 3.33 before squaring, which gives \(S^2 = 2.36\) (exactly \(7/3 = 2.33\)) and a total of 7.0001 in the last column (exactly 7). It also introduces the 9 samples with “in srswor”, where with-replacement sampling is meant.

Exercises on Simple Random Sampling, with Answers Checked

PRACTICE
  1. Population 1, 2, 3, 4, 5, 6; all SRSWOR samples of size 2. Verify unbiasedness of \(\bar y\) and \(s^2\), that the sampling variance equals \(\text{Var}(\bar y)\), and that \(\text{Var}_{WOR} \le \text{Var}_{WR}\). Ans. \(\bar Y = 3.5\), \(S^2 = 3.5\), \(\sigma^2 = 2.917\), \(\text{Var}_{WOR}(\bar y) = 1.167\), \(\text{Var}_{WR}(\bar y) = 1.458\) (the 15 sample means have variance exactly \(7/6\)).
  2. Population 1, 2, 3, 4, 5; all SRSWR samples of size 2. Ans. \(\bar Y = 3\), \(S^2 = 2.5\), \(\sigma^2 = 2\), \(\text{Var}_{WR}(\bar y) = 1\) (over the 25 samples, \(E(s^2) = 2 = \sigma^2\)).
  3. Draw a random sample of 25 from a population of 800 without replacement. Ans. Read three-digit numbers from any random starting point, skipping those above 800, 000 and repeats, until 25 are obtained.

Key Take-aways

Extra Practical Problems

PRACTICE

Additional worked problems with step-by-step procedures to support self-study, matching this unit's topics.

STEP-BY-STEP PROCEDURE (selecting a simple random sample)

Lottery method

  1. Number or name every unit of the population on identical slips of paper.
  2. Fold the slips and mix them thoroughly in a container.
  3. Make a blindfold draw of as many slips as the required sample size — each unit has an equal chance of selection.

Random number table method

  1. Number the population units 1…\(N\). Decide how many digits to read (e.g. 2 digits if \(N\) is two-digit, 3 digits if three-digit).
  2. Pick a random starting point in the table and read numbers down a column.
  3. Accept a number if it lies in 1…\(N\); discard if it exceeds \(N\). For SRSWOR discard repeats; for SRSWR repeats are allowed.
  4. Continue until the required sample size is obtained; the corresponding units form the sample.

Counts of possible samples

Problem 1 — Selecting a Sample with a Random Number Table

METHOD

Population of 200 students numbered 1–200; select 5 at random. Because the population size is a 3-digit number, read the first 3 digits of each 5-digit entry in the random-number table; discard any > 200 and any repeats. Starting at one entry and moving down the column gives 200, 023, 108, 070, 126 → students 23, 70, 108, 126, 200. Each unit had an equal, independent chance of selection.

Problem 2 — SRS With and Without Replacement (diseased plants)

DATA

Diseased plants in 24 areas: 1, 4, 1, 2, 5, 1, 1, 1, 7, 2, 3, 3, 2, 2, 3, 1, 2, 7, 2, 6, 3, 5, 3, 4. Select a sample of size 6 by SRSWR and SRSWOR; compare with the population mean.

Population mean \(= \dfrac{\sum y_i}{24} = \dfrac{71}{24} = 2.96\).

SRSWR (2-digit reading; repeats allowed) → areas 02, 10, 11, 12, 17, 17 with values 4, 2, 3, 3, 2, 2 → sample mean \(= \dfrac{16}{6} = 2.67\).

SRSWOR (repeats not allowed) → areas 01, 04, 12, 19, 20, 22 with values 1, 2, 3, 2, 6, 5 → sample mean \(= \dfrac{19}{6} = 3.17\).

Both sample means (2.67, 3.17) are close to the population mean (2.96) — illustrating that the sample mean estimates the population mean.

Problem 3 — Sample Mean is Unbiased under SRSWOR

DATA

Population units 1, 2, 3, 4, 5; draw all samples of size \(n=3\) under SRSWOR.

Number of samples \(= \binom{5}{3} = 10\). Population mean \(\bar Y = 15/5 = 3\). The 10 sample means total 30.0, so \(E(\bar y) = \dfrac{\sum \bar y}{\binom{5}{3}} = \dfrac{30}{10} = 3 = \bar Y\).

Hence the sample mean is an unbiased estimate of the population mean under SRSWOR.

Problem 4 — Sample Mean is Unbiased under SRSWR

DATA

Population units 1, 2, 3, 4, 5; draw all samples of size \(n=2\) under SRSWR.

Number of samples \(= N^n = 5^2 = 25\). The 25 sample means total 75.0, so \(E(\bar y) = \dfrac{75}{25} = 3 = \bar Y\). The sample mean is unbiased under SRSWR as well.

Unsolved Exercises

PRACTICE
  1. Workers in 12 factories: 2145, 1547, 745, 215, 784, 3125, 126, 471, 841, 3215, 2496, 589. Draw an SRSWOR of size 4 with a random-number table; compare the sample and population averages.
  2. A class has 115 students. Select an SRSWR of size 15.
  3. Yields (q/ha) of 30 paddy varieties (49, 78, 57, …, 40). Draw an SRSWOR of size 8 and compare the sample mean with the population mean.
  4. Population 1–7. List all SRSWOR samples of size 2 and verify the sample mean estimates the population mean.
  5. How many SRSWR samples of size 5 can be drawn from a population of size 10? (Ans: \(10^5\))