Skip to the content

Topics Covered

Systematic Sampling Estimate of Mean Variance for Linear Trend Comparison with SRSWOR / StRS Intraclass Correlation Circular Systematic Sampling Cluster Sampling Multistage Sampling Quota Sampling
On this page
  1. 1. Systematic Sampling
  2. 2. Cluster Sampling
  3. 3. Multistage Sampling
  4. 4. Quota Sampling
  5. 5. Probability Proportional to Size (PPS) Sampling
  6. Key Take-aways

1. Systematic Sampling

DEFINITION

Suppose population size \(N\) is exactly divisible by sample size \(n\), so \(N = nk\) where \(k\) is the sampling interval. Choose a random integer \(r\) (called the random start) between 1 and \(k\), then select units numbered

\[ r,\; r + k,\; r + 2k,\; \ldots,\; r + (n-1)k. \]

This gives a sample of size \(n\). Notation: SYS\((N, n, k)\).

1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 N = 20, n = 5, k = 4 → random start r = 3 Units 3, 7, 11, 15, 19 are selected
Fig 4.1 — Systematic sampling — pick every k-th unit after a random start

The \(k\) Possible Systematic Samples

Example. A company with \(N = 5000\) customers wants a sample of \(n = 100\). The sampling ratio is \(n/N = 100/5000 = 1/50\), so the interval is \(k = 50\): one customer in every 50. A random start is drawn from 1 to 50, say 11, and the sample is customers 11, 61, 111, 161, …, up to \(11 + 99 \times 50 = 4961\). Every unit after the first is fixed by the interval.

Notation. Write \(y_{ij}\) for the \(j\)th unit of the \(i\)th systematic sample (\(i = 1, \ldots, k\) is the random start, \(j = 1, \ldots, n\)). Then

\[ \bar y_{i.} = \frac1n\sum_{j=1}^n y_{ij}, \qquad \bar y_{..} = \frac{1}{nk}\sum_{i=1}^k\sum_{j=1}^n y_{ij} = \frac1k\sum_{i=1}^k \bar y_{i.}, \] \[ S^2 = \frac{1}{nk-1}\sum_{i=1}^k\sum_{j=1}^n (y_{ij} - \bar y_{..})^2 . \]

The population mean \(\bar Y\) is written \(\bar y_{..}\) here, and \(\bar y_{sys}\) is the \(\bar y_{i.}\) of whichever start was drawn.

Random startUnits in the sampleProbabilityMean
1\(1,\; 1+k,\; \ldots,\; 1+(n-1)k\)\(1/k\)\(\bar y_{1.}\)
2\(2,\; 2+k,\; \ldots,\; 2+(n-1)k\)\(1/k\)\(\bar y_{2.}\)
\(\vdots\)\(\vdots\)\(\vdots\)\(\vdots\)
\(i\)\(i,\; i+k,\; \ldots,\; i+(n-1)k\)\(1/k\)\(\bar y_{i.}\)
\(\vdots\)\(\vdots\)\(\vdots\)\(\vdots\)
\(k\)\(k,\; 2k,\; \ldots,\; nk\)\(1/k\)\(\bar y_{k.}\)

Each row is one possible sample, and each of the \(N\) units lies in exactly one row, so every unit has the same chance \(1/k = n/N\) of selection. Each column is a block of \(k\) consecutive units, so the columns form \(n\) strata. Systematic sampling takes the unit in the same position from every stratum; stratified sampling with one unit per stratum would choose the position afresh in each stratum. This picture is behind every comparison below.

Merits

  1. Operationally very simple — only one random number needed.
  2. Spread evenly over the population — gives a representative sample.
  3. For populations with a slow trend, often more efficient than SRSWOR.
  4. No need to know population frame in advance — can be applied as units arrive.

Demerits

  1. If \(N\) is not exactly \(nk\), some units may have unequal selection probability.
  2. Periodic / cyclical population can cause severe bias if the period coincides with \(k\).
  3. Variance estimation is difficult from a single systematic sample.

Merits and Demerits, Looked at More Closely

CORRECTION — A RANDOM ORDER GIVES NO GAIN

The textbook says systematic sampling “may be more efficient than simple random sampling provided the sampling frame is arranged totally at random”, for example an alphabetical list. A list in random order makes systematic sampling equivalent to SRSWOR on average, not better: averaged over all \(6! = 720\) orderings of the values 1, 2, 4, 7, 11, 16 (\(n = 2\), \(k = 3\)), \(\text{Var}(\bar y_{sys})\) is exactly \(11.12\), the SRSWOR variance. The gain comes when the list is ordered by something related to \(y\) (a trend, §1.3). A periodic order can make it much worse: for the values 10, 20, 30, 20 repeated three times (\(N = 12\), \(k = 4\)), every systematic sample consists of one value repeated, and \(\text{Var}(\bar y_{sys}) = 50\) against \(13.64\) for SRSWOR.

TWO PRACTICAL DIFFICULTIES
  1. \(N\) not a multiple of \(n\). With \(k\) an integer near \(N/n\), different random starts can give samples of different sizes (for \(N = 23\), \(k = 5\): starts 1 to 3 give 5 units, starts 4 and 5 give 4), and \(\bar y_{sys}\) is then biased. Circular systematic sampling removes this: draw the start \(r\) from 1 to \(N\), and take units \(r, r+k, r+2k, \ldots\), counting on from unit 1 again after unit \(N\), until \(n\) are chosen. Every unit then has probability \(n/N\).
  2. No unbiased estimate of the variance. \(\text{Var}(\bar y_{sys}) = \frac1k\sum_{i=1}^k (\bar y_{i.} - \bar y_{..})^2\) is a variance between the \(k\) possible samples, and one sample shows only one of them. Using \(s^2/n\) as if the sample were SRS is not unbiased. The usual remedy is to draw several systematic samples with independent random starts and use the spread of their means.

Systematic sampling is also not a simple random sample in the strict sense: only the first unit is random, and most of the \(\binom Nn\) subsets can never be drawn.

1.1 Estimate of Mean

The systematic sample mean

\[ \bar y_{sys} \;=\; \dfrac{1}{n}\sum_{j=0}^{n-1} y_{r + jk}. \]

It is unbiased: \(E(\bar y_{sys}) = \bar Y\).

Proof of Unbiasedness

THEOREM 1

The random start \(i\) takes each value \(1, \ldots, k\) with probability \(1/k\), so

\[ E(\bar y_{sys}) = \frac1k\,\bar y_{1.} + \frac1k\,\bar y_{2.} + \cdots + \frac1k\,\bar y_{k.} = \frac1k\sum_{i=1}^k \bar y_{i.} = \bar y_{..} = \bar Y . \]

This needs \(N = nk\): then all \(k\) samples have \(n\) units and the average of their means is the population mean.

1.2 Variance — General Formula

Let \(S_{wsy}^2\) be the variance within systematic samples. Then

\[ \text{Var}(\bar y_{sys}) \;=\; \dfrac{N - 1}{N}\, S^2 - \dfrac{k(n - 1)}{N}\, S_{wsy}^2. \]

Equivalently, variance can be expressed via intra-class correlation:

\[ \text{Var}(\bar y_{sys}) \;=\; \dfrac{N-1}{N}\cdot\dfrac{S^2}{n}\bigl[1 + (n-1)\rho\bigr], \]

where \(\rho\) is the intra-class correlation among units of the same systematic sample.

Deriving the Variance

THEOREM 2

Since \(\bar y_{sys}\) takes the value \(\bar y_{i.}\) with probability \(1/k\),

\[ \text{Var}(\bar y_{sys}) = E(\bar y_{i.} - \bar y_{..})^2 = \frac1k\sum_{i=1}^k (\bar y_{i.} - \bar y_{..})^2 . \qquad (1) \]

Split each deviation as \(y_{ij} - \bar y_{..} = (y_{ij} - \bar y_{i.}) + (\bar y_{i.} - \bar y_{..})\). The cross term vanishes, because \(\sum_j (y_{ij} - \bar y_{i.}) = n\bar y_{i.} - n\bar y_{i.} = 0\) for each \(i\). Hence

\[ (N-1)S^2 = \sum_{i}\sum_{j}(y_{ij} - \bar y_{i.})^2 + n\sum_i (\bar y_{i.} - \bar y_{..})^2 = k(n-1)S_{wsy}^2 + nk\,\text{Var}(\bar y_{sys}), \]

where \(S_{wsy}^2 = \frac{1}{k(n-1)}\sum_i\sum_j (y_{ij} - \bar y_{i.})^2\) is the mean square within systematic samples. Dividing by \(N = nk\),

\[ \text{Var}(\bar y_{sys}) = \frac{N-1}{N}\,S^2 - \frac{k(n-1)}{N}\,S_{wsy}^2 . \]

When Systematic Beats SRSWOR

COMPARISON THROUGH \(S_{wsy}^2\) \[ \text{Var}_{SRS} - \text{Var}_{sys} = \frac{N-n}{N}\cdot\frac{S^2}{n} - \frac{N-1}{N}S^2 + \frac{k(n-1)}{N}S_{wsy}^2 \]

The \(S^2\) terms combine as \(\frac{N - n - Nn + n}{Nn}S^2 = -\frac{n-1}{n}S^2\), and \(\frac{k(n-1)}{N} = \frac{n-1}{n}\), so

\[ \text{Var}_{SRS} - \text{Var}_{sys} = \frac{n-1}{n}\big(S_{wsy}^2 - S^2\big) . \]

Systematic sampling is more precise than SRSWOR exactly when \(S_{wsy}^2 > S^2\): when the units within a systematic sample are more varied than the population as a whole. Spreading the sample across the list is what achieves this.

The Intraclass Correlation Form

THEOREM 3

Let \(\rho\) be the correlation between pairs of units in the same systematic sample,

\[ \rho = \frac{\sum_i \sum_{j \ne j'} (y_{ij} - \bar y_{..})(y_{ij'} - \bar y_{..})}{(n-1)(nk-1)S^2} . \]

From (1), \(n^2k\,\text{Var}(\bar y_{sys}) = \sum_i \big[\sum_j (y_{ij} - \bar y_{..})\big]^2\). Squaring the inner sum gives the squares, which add to \((nk-1)S^2\), and the cross products, which add to \((n-1)(nk-1)\rho S^2\). Hence

\[ \text{Var}(\bar y_{sys}) = \frac{nk-1}{nk}\cdot\frac{S^2}{n}\big[1 + (n-1)\rho\big] . \]

The short form \(\frac{S^2}{n}[1 + (n-1)\rho]\), often quoted, drops the factor \(\frac{N-1}{N} = \frac{nk-1}{nk}\); it is the large-\(N\) approximation.

EFFICIENCY OVER SRSWOR

With \(\text{Var}_{SRS} = \frac{N-n}{N}\cdot\frac{S^2}{n} = \frac{k-1}{k}\cdot\frac{S^2}{n}\),

\[ E = \frac{\text{Var}_{SRS}}{\text{Var}_{sys}} = \frac{n(k-1)}{(nk-1)\big[1 + (n-1)\rho\big]} . \]

Systematic sampling is the more efficient when \(E > 1\), that is, \(n(k-1) > (nk-1) + (n-1)(nk-1)\rho\), which simplifies to \((n-1)\big[1 + (nk-1)\rho\big] < 0\):

\[ E > 1 \iff \rho < -\frac{1}{nk-1}, \qquad E < 1 \iff \rho > -\frac{1}{nk-1} . \]

A negative correlation within samples is what helps: each sample then mixes high and low values. Two special cases: if \(\rho = 1\) (every sample constant), \(E = \frac{k-1}{nk-1}\), the worst possible; and since \(\text{Var}(\bar y_{sys}) \ge 0\), \(\rho\) can be no smaller than \(-\frac{1}{n-1}\), at which value \(\text{Var}(\bar y_{sys}) = 0\) and \(E\) is unbounded.

CORRECTION — THE BOUNDARY VALUE OF \(\rho\)

The textbook states that at \(\rho = -\frac{1}{nk-1}\), \(\text{Var}(\bar y_{sys}) = 0\) and \(E \to \infty\). That value is the boundary just derived, where \(E = 1\) exactly: the two designs are equally precise (with \(n = 3\), \(k = 5\): \(\rho = -1/14\), \(1 + 2\rho = 6/7\), and \(E = 1\)). The variance vanishes only at \(\rho = -\frac{1}{n-1}\), where \(1 + (n-1)\rho = 0\).

Systematic versus Stratified

THEOREM 4

Take the \(n\) columns of the table as strata of \(k\) units each, with means \(\bar y_{.j} = \frac1k\sum_i y_{ij}\), mean squares \(S_j^2 = \frac{1}{k-1}\sum_i (y_{ij} - \bar y_{.j})^2\), and pooled within-stratum mean square \(S_{wst}^2 = \frac{1}{n(k-1)}\sum_i\sum_j (y_{ij} - \bar y_{.j})^2\). Stratified sampling with one unit per stratum has \(N_j = k\), \(n_j = 1\), \(W_j = 1/n\), so

\[ \text{Var}(\bar y_{st}) = \sum_{j=1}^n \Big(1 - \frac1k\Big)\frac{S_j^2}{n^2} = \frac{k-1}{n^2k}\sum_j S_j^2 = \frac{k-1}{nk}\,S_{wst}^2 . \]

For the systematic sample, \(\bar y_{i.} - \bar y_{..} = \frac1n\sum_j (y_{ij} - \bar y_{.j})\), because \(\bar y_{..}\) is also the average of the column means. Squaring and summing as in Theorem 3, with \(\rho_{wst}\) the correlation between deviations from the stratum means within the same systematic sample,

\[ \rho_{wst} = \frac{\sum_i\sum_{j \ne j'} (y_{ij} - \bar y_{.j})(y_{ij'} - \bar y_{.j'})}{(n-1)\,n(k-1)\,S_{wst}^2}, \] \[ \text{Var}(\bar y_{sys}) = \frac{k-1}{nk}\,S_{wst}^2\big[1 + (n-1)\rho_{wst}\big], \qquad \frac{\text{Var}(\bar y_{st})}{\text{Var}(\bar y_{sys})} = \frac{1}{1 + (n-1)\rho_{wst}} . \]

So: \(\rho_{wst} > 0\) favours stratified sampling, \(\rho_{wst} = 0\) makes the two equal, and \(\rho_{wst} < 0\) favours systematic sampling. (The textbook states the first two cases only.)

1.3 Variance for a Linear Trend

If the population values follow a linear trend \(Y_i = a + bi\), the variances simplify to:

\[ \text{Var}_{SRSWOR}(\bar y) \;=\; \dfrac{(N-n)b^2 (N+1)}{12 n}, \] \[ \text{Var}_{StRS, prop}(\bar y) \;=\; \dfrac{b^2(k^2 - 1)}{12 n}, \qquad \text{Var}_{SYS}(\bar y) \;=\; \dfrac{b^2(k^2 - 1)}{12}. \]

Variance ordering for linear trend:

\[ \text{Var}_{StRS, prop} \;\le\; \text{Var}_{SYS} \;\le\; \text{Var}_{SRSWOR}. \]

Stratified (one unit per stratum) is best for a linear trend and SRSWOR is worst; systematic sits in between. In fact \(\text{Var}_{SYS} = n\,\text{Var}_{StRS,prop}\) and \(\text{Var}_{SRSWOR} = \tfrac{N+1}{k+1}\text{Var}_{SYS}\), so systematic sampling always beats SRSWOR here — spreading the sample evenly across the ordered trend captures its full range better than a purely random draw.

Since populations rarely follow an exact linear trend, in practice systematic ≈ stratified for many real data sets. (Systematic is only poor when the population is periodic with period equal to \(k\).)

Proof for a Linear Trend

THEOREM 5

Take \(Y_i = i\), \(i = 1, \ldots, N\) (any trend \(a + bi\) multiplies every variance by \(b^2\)). Then \(\bar Y = \frac{N+1}{2}\) and

\[ S^2 = \frac{1}{N-1}\Big[\frac{N(N+1)(2N+1)}{6} - N\Big(\frac{N+1}{2}\Big)^2\Big] \] \[ = \frac{N(N+1)}{N-1}\cdot\frac{N-1}{12} = \frac{N(N+1)}{12} . \]

SRSWOR. \(\text{Var}_{SRS} = \frac{N-n}{Nn}S^2 = \frac{(k-1)(nk+1)}{12}\).

Stratified. Each stratum is \(k\) consecutive integers, so \(S_j^2 = \frac{k(k+1)}{12}\), and \(\text{Var}_{st} = \frac{k-1}{n^2k}\cdot n\,\frac{k(k+1)}{12} = \frac{k^2-1}{12n}\).

Systematic. The \(i\)th sample is \(i, i+k, \ldots, i+(n-1)k\), so

\[ \bar y_{i.} = \frac1n\big[ni + (1 + 2 + \cdots + (n-1))k\big] = i + \frac{(n-1)k}{2}, \]

and \(\bar y_{..} = \frac{nk+1}{2}\), so \(\bar y_{i.} - \bar y_{..} = i - \frac{k+1}{2}\): the \(k\) sample means are equally spaced, one unit apart. Their variance is that of \(1, \ldots, k\):

\[ \text{Var}_{sys} = \frac1k\sum_{i=1}^k\Big(i - \frac{k+1}{2}\Big)^2 = \frac{k^2-1}{12} . \]

Dividing all three by \(\frac{k-1}{12}\):

\[ \text{Var}_{st} : \text{Var}_{sys} : \text{Var}_{SRS} = \frac{k+1}{n} : (k+1) : (nk+1), \]

which for large \(k\) is about \(\frac1n : 1 : n\). Stratified sampling removes the trend best, and systematic sampling comes next. (Check with \(N = 15\), \(n = 3\), \(k = 5\): the three variances are \(\frac23\), \(2\) and \(\frac{16}{3}\), in the ratio \(2 : 6 : 16 = \frac63 : 6 : 16\).)

1.4 Comparison with SRSWOR & Stratified

AspectSRSWORStratifiedSystematic
Sample selectionRandomRandom in each stratumOne random start, every k-th
Variance for homogeneous pop.SameSameSame
Variance with linear trendWorstBestIntermediate (beats SRSWOR)
Variance with periodic pop.SameSameCan be very bad
Operational easeModerateHard (frame, allocation)Easiest
EXAMPLE 1

From \(N = 50\) employees, draw a sample of 10 by systematic sampling. \(k = 5\); random start \(r = 3\); sample = {3, 8, 13, 18, 23, 28, 33, 38, 43, 48}.

EXAMPLE 2 (Comparison)

For population values \(1, 2, 3, \ldots, 12\) (\(N = 12, n = 4, k = 3, b = 1\)):

\(\text{Var}_{SYS} = (3^2 - 1)/12 = 8/12 = 0.667\).

\(\text{Var}_{SRSWOR} = (12-4)(1)(13)/(12 \cdot 4) = 104/48 = 2.167\).

\(\text{Var}_{StRS,prop} = 8/(12 \cdot 4) = 0.167\).

So Stratified < Systematic < SRSWOR for this linear trend (note systematic is BETTER than SRSWOR here because the linear trend's variance reduction within systematic samples dominates).

Worked Problems on Systematic Sampling

The textbook's problem, worked by both routes (the formulas of Theorems 2 and 4, and directly from the \(k\) possible samples), then its exercise with the answers checked.

Source note. Every figure was recomputed exactly. Worked Problem 1 agrees with the textbook up to rounding; in the exercise, the textbook's systematic variance is wrong and is corrected below.

WORKED PROBLEM 1 — 15 units, 5 systematic samples of 3

A population of \(N = 15\) units forms \(k = 5\) systematic samples of \(n = 3\) units (rows); the columns are the 3 strata:

Sample \(i\)Stratum 1Stratum 2Stratum 3Total\(\bar y_{i.}\)\(\sum_j y_{ij}^2\)
1148134.3381
2612203812.67580
34816289.33336
4310183110.33433
5412244013.33736
Column total1846861502166
\(\bar y_{.j}\)3.69.217.2

Find the variance of the sample mean under SRSWOR, stratified sampling (one unit per stratum) and systematic sampling, and compare their efficiencies.

Population constants. \(\bar y_{..} = 150/15 = 10\), and

\[ S^2 = \frac{1}{nk-1}\Big[\sum\sum y_{ij}^2 - nk\,\bar y_{..}^2\Big] = \frac{2166 - 15(10)^2}{14} = \frac{666}{14} = 47.57 . \]

SRSWOR. \(\text{Var}_{SRS} = \frac{N-n}{N}\cdot\frac{S^2}{n}\) \(= \frac{12}{15 \times 3}(47.57) = 12.69\).

Stratified. \(\sum_j \bar y_{.j}^2 = 12.96 + 84.64 + 295.84 = 393.44\), so

\[ S_{wst}^2 = \frac{1}{n(k-1)}\Big[\sum\sum y_{ij}^2 - k\sum_j \bar y_{.j}^2\Big] = \frac{2166 - 5(393.44)}{12} = \frac{198.8}{12} = 16.57, \]

and \(\text{Var}_{st} = \frac{k-1}{nk}S_{wst}^2 = \frac{4}{15}(16.57) = 4.42\).

Systematic. \(\sum_i \bar y_{i.}^2 = (13^2 + 38^2 + 28^2 + 31^2 + 40^2)/9\) \(= 4958/9 = 550.89\), so

\[ S_{wsy}^2 = \frac{1}{k(n-1)}\Big[\sum\sum y_{ij}^2 - n\sum_i \bar y_{i.}^2\Big] = \frac{2166 - 3(550.89)}{10} = 51.33, \] \[ \text{Var}_{sys} = \frac{nk-1}{nk}S^2 - \frac{n-1}{n}S_{wsy}^2 = \frac{14}{15}(47.57) - \frac23(51.33) = 10.18 . \]

Directly. The five possible means deviate from 10 by \(-5.67, 2.67, -0.67, 0.33, 3.33\); their squares add to 50.89, and \(\text{Var}_{sys} = 50.89/5 = 10.18\) (exactly \(458/45\)), as Theorem 2 says.

Comparison. \(\text{Var}_{st} = 4.42\) < \(\text{Var}_{sys} = 10.18\) < \(\text{Var}_{SRS} = 12.69\). Stratified sampling is the most efficient (2.87 times SRSWOR), then systematic (1.25 times SRSWOR). The correlations explain it: within systematic samples \(\rho = -0.156\), below the boundary \(-1/(nk-1) = -0.071\), so systematic beats SRSWOR; but \(\rho_{wst} = 0.652 > 0\), so stratified beats systematic.

Rounding note. The textbook rounds the sample means before squaring and gets \(\sum \bar y_{i.}^2 = 550.95\), \(S_{wsy}^2 = 51.315\) and \(\text{Var}_{sys} = 10.19\); at full precision they are 550.89, 51.33 and 10.18.

5 systematic sample means i=1: 4.33 i=2: 12.67 i=3: 9.33 i=4: 10.33 i=5: 13.33 SRSWOR ±3.56 systematic ±3.19 stratified ±2.10 2 4 6 8 10 12 14 16 18 ȳ․․ = 10 sample mean (N = 15, n = 3, k = 5)
Fig 4.2 — Worked Problem 1. The systematic estimator can take only the five values at the top, one for each random start; their spread about \(\bar y_{..} = 10\) is \(\text{Var}(\bar y_{sys})\). The bars compare one standard error under each design: stratified is tightest because the strata (columns) differ sharply, and systematic beats SRSWOR because the units within a systematic sample are negatively correlated (\(\rho = -0.156\)).
PRACTICE
  1. In the table below the 10 columns are systematic samples and the 4 rows are strata (\(N = 40\), \(n = 4\), \(k = 10\)). Find the variance of the sample mean under SRSWOR, stratified and systematic sampling, and comment on the efficiencies.
    Stratum12345678910
    I0112547786
    II68910131215161617
    III18192020242325282927
    IV26303131333235373838
    Ans. \(\bar y_{..} = 18.175\), \(S^2 = 136.25\), \(\text{Var}_{SRS} = 30.656\); \(S_{wst}^2 = 13.49\), \(\text{Var}_{st} = 3.034\); \(S_{wsy}^2 = 161.625\), \(\text{Var}_{sys} = 11.626\). Stratified is the most efficient, then systematic, then SRSWOR. The textbook prints \(\text{Var}_{sys} = 11.807\); the ten sample means (12.5, 14.5, 15.25, 15.75, 18.75, 17.75, 20.5, 22, 22.75, 22) have variance 11.626 about 18.175, and Theorem 2 gives the same. Its closing sentence should read “systematic sampling is more efficient than simple random sampling”, not “than stratified random sampling”.

2. Cluster Sampling

DEFINITION

The population is divided into groups called clusters; an SRS of clusters is selected and every unit within the chosen cluster is observed (single-stage cluster sampling).

Clusters are usually natural groupings like villages, schools, blocks, or households.

When is Cluster Sampling Appropriate?

Cluster vs Stratified — important contrast

StratifiedCluster
Internal variationStrata homogeneous within (low \(S_h^2\))Clusters heterogeneous within
Between groupsStrata differ (high between)Clusters similar to each other
Sample fromEvery stratumSome clusters only
GoalReduce varianceReduce cost

A cluster sample is generally less precise than SRS of the same total sample size (when clusters are formed by proximity) but is much cheaper to administer.

EXAMPLE 1

To study primary schools in a district: list the schools, randomly select 20 schools, and survey all teachers and students in those 20 schools.

EXAMPLE 2

National Sample Survey rural blocks: villages are clusters; a random sample of villages is selected and all households in the chosen villages are surveyed.

3. Multistage Sampling

DEFINITION

Sampling done in stages. At the first stage, primary sampling units (PSUs, e.g., districts) are selected. Within each selected PSU, secondary sampling units (SSUs, e.g., villages) are selected. The process can continue further — three-stage, four-stage etc.

Advantages

  1. Frame at lower levels needed only for selected higher units — reduces effort.
  2. Allows flexible cost control at each stage.
  3. Used in large national surveys.
EXAMPLE 1

Indian Census-NSSO style: select states (1st stage), districts within states (2nd stage), villages within districts (3rd stage), households within villages (4th stage).

EXAMPLE 2

Education survey: select schools (PSU), then classes within schools (SSU), then students within classes.

4. Quota Sampling

DEFINITION

The investigator is given quotas from each subgroup (gender, age, etc.) to fill, but is free to choose which individuals fill those quotas. It is a non-probability method.

Advantages

Disadvantages

EXAMPLE 1

Market researcher told: "interview 30 working women aged 25–34". Picks any 30 willing women — no random selection.

EXAMPLE 2

Exit poll instructions: "interview 50 men and 50 women voters from each constituency". No random scheme — interviewer picks who is convenient.

5. Probability Proportional to Size (PPS) Sampling

In PPS sampling, units are selected with probabilities proportional to a known size measure \(X_i\) (e.g. villages chosen with probability proportional to population, or shops to floor area). Large units — which usually contribute most to the total — get a higher chance of selection, improving efficiency when \(Y_i\) is roughly proportional to \(X_i\).

Selection probability of unit \(i\): \( \pi_i = X_i / \sum_{k} X_k. \)

HANSEN–HURWITZ ESTIMATOR (PPS with replacement) \[ \hat Y_{HH} = \dfrac{1}{n}\sum_{i=1}^{n}\dfrac{y_i}{p_i}, \qquad p_i = \dfrac{X_i}{\sum_k X_k}, \]

which is unbiased for the population total \(Y\); dividing the summand by the size restores an unbiased scale.

Selection methods: (i) Cumulative-total method — cumulate the sizes, draw a random number in \((0, \sum X_k]\), and pick the unit whose cumulative interval contains it; (ii) Lahiri's method — a rejection scheme that avoids cumulation.

EXAMPLE

Four villages with populations 200, 300, 100, 400 (total 1000) have selection probabilities \(0.20, 0.30, 0.10, 0.40\). Cumulative totals are 200, 500, 600, 1000; a random number 640 falls in the interval \((600, 1000]\), so the fourth village is selected.

Key Take-aways