The population of size \(N\) is divided into \(L\) non-overlapping strata (homogeneous groups) of sizes \(N_1, N_2, \ldots, N_L\) with \(N = \sum N_h\). From each stratum, an independent SRSWOR of size \(n_h\) is drawn, with \(n = \sum n_h\). This is stratified random sampling.
Stratification means division into layers. Auxiliary information (past data, or another variable related to the one under study) is used to divide the population so that
A survey of the cost of living across a state is the typical case: with SRS one cannot be sure that the high-, middle- and low-income groups are all represented; stratifying by income guarantees each its share. The strata must be formed properly and given suitable sample sizes, or the stratified sample may be no better than a simple random one (§6). Administrative convenience comes chiefly when the strata are geographical (districts, branches): then the field work of each stratum is compact and can be supervised locally.
Notation. This page writes \(L\) strata indexed by \(h\), with weights \(W_h = N_h/N\); the textbook writes \(k\) strata indexed by \(i\), with weights \(p_i\).
| Symbol | Meaning |
|---|---|
| \(N_h\) | Size of stratum \(h\) |
| \(n_h\) | Sample size in stratum \(h\) |
| \(W_h = N_h/N\) | Stratum weight (proportion) |
| \(\bar Y_h\) | Population mean of stratum \(h\) |
| \(S_h^2\) | Population mean square of stratum \(h\) |
| \(\bar y_h\) | Sample mean from stratum \(h\) |
| \(s_h^2\) | Sample mean square from stratum \(h\) |
The overall population mean: \(\bar Y = \sum_h W_h \bar Y_h\).
This is unbiased: \(E(\bar y_{st}) = \sum W_h E(\bar y_h) = \sum W_h \bar Y_h = \bar Y\).
Each \(\bar y_h\) comes from an SRSWOR within stratum \(h\), so \(E(\bar y_h) = \bar Y_h\). Then
\[ E(\bar y_{st}) = E\Big(\frac1N\sum_h N_h\bar y_h\Big) = \frac1N\sum_h N_h\bar Y_h = \bar Y , \]whatever the sample sizes \(n_h\) are.
The textbook first defines the mean of the stratified sample as the plain average of all \(n\) observations, \(\frac1n\sum_h n_h\bar y_h\) (its inner sum is printed to \(N_i\) where \(n_i\) is meant), and then says that “if \(n_i/n = N_i/N\), then [it is] considered as \(\bar y_{st}\)”. The estimator \(\bar y_{st} = \sum_h W_h\bar y_h\) is defined for every allocation; it coincides with the plain average only under proportional allocation. With any other allocation the plain average is biased, because it over-weights the strata that were sampled heavily.
Example. Strata \(\{1, 2, 4\}\) and \(\{6, 9, 10, 13\}\), so \(\bar Y = 45/7 = 6.43\), with \(n_1 = n_2 = 2\). Over all \(3 \times 6 = 18\) stratified samples, \(\bar y_{st}\) averages exactly \(45/7\), but the plain average of the four observations averages \(71/12 = 5.92\): the small stratum, with 3 of the 7 units, supplies half the sample.
Estimator of variance:
The samples in different strata are drawn independently, so the covariances between the \(\bar y_h\) vanish and
\[ \text{Var}(\bar y_{st}) = \frac{1}{N^2}\sum_h N_h^2\,\text{Var}(\bar y_h) = \frac{1}{N^2}\sum_h N_h^2\,\frac{N_h - n_h}{N_h}\cdot\frac{S_h^2}{n_h} \] \[ = \frac{1}{N^2}\sum_h N_h(N_h - n_h)\frac{S_h^2}{n_h} , \]which is \(\sum_h W_h^2 S_h^2(1/n_h - 1/N_h)\). Only the within-stratum mean squares \(S_h^2\) appear: the differences between strata have been designed out.
Check by listing. For the strata \(\{1, 2, 4\}\), \(\{6, 9, 10, 13\}\) with \(n_1 = n_2 = 2\), the 18 values of \(\bar y_{st}\) have variance exactly \(221/294 = 0.752\), which is what the formula gives.
Given total sample size \(n\), we must decide \(n_h\) for each stratum. Two important schemes:
Sample size proportional to stratum size; ignores within-stratum variance.
Variance under proportional allocation:
Allocates more to strata that are larger and more variable. Minimises variance for a given total sample size.
Variance under optimum allocation:
If sampling cost in stratum \(h\) is \(c_h\): \(n_h \propto N_h S_h / \sqrt{c_h}\) — the cost-optimal version.
\(n_h = n/L\) — used when no prior information on \(S_h\); rarely optimal.
\(N_1 = 600,\; N_2 = 400\); \(n = 100\). Then \(n_1 = 60, n_2 = 40\).
Two strata: \(N_1 = 600, S_1 = 4;\; N_2 = 400, S_2 = 9\). Total sample 100.
\(N_1 S_1 = 2400,\; N_2 S_2 = 3600\), sum = 6000.
\(n_1 = 100 \cdot 2400/6000 = 40,\; n_2 = 100 \cdot 3600/6000 = 60\). Note: stratum 2 has fewer units but more variability — gets a bigger sample share.
Minimise \(V = \frac{1}{N^2}\sum_h N_h(N_h/n_h - 1)S_h^2\) subject to \(\sum_h n_h = n\). With a Lagrange multiplier \(\lambda\), \(\phi = V + \lambda(\sum n_h - n)\), and only the \(h\)th term depends on \(n_h\):
\[ \frac{\partial\phi}{\partial n_h} = -\frac{N_h^2S_h^2}{N^2n_h^2} + \lambda = 0 \;\Longrightarrow\; n_h = \frac{N_hS_h}{N\sqrt\lambda} . \]Summing over \(h\), \(n = \sum N_hS_h/(N\sqrt\lambda)\), so \(\sqrt\lambda = \sum N_hS_h/(Nn)\) and
\[ n_h = n\,\frac{N_hS_h}{\sum_k N_kS_k} . \]It is a minimum because \(\partial^2\phi/\partial n_h^2 = 2N_h^2S_h^2/(N^2n_h^3) > 0\). (The textbook prints this second derivative without the factor 2; the sign, which is what matters, is unaffected.)
Substituting back gives the optimum variance of §5.2, \(\frac1n\big(\sum W_hS_h\big)^2 - \frac1N\sum W_hS_h^2\).
With cost function \(C = a + \sum_h c_hn_h\) (overhead \(a\), cost \(c_h\) per unit in stratum \(h\)), the same Lagrange argument with the constraint \(\sum c_hn_h = C - a\) gives \(n_h = N_hS_h/(N\sqrt{\lambda c_h})\), i.e.
\[ n_h \;\propto\; \frac{N_hS_h}{\sqrt{c_h}} . \]Take a larger sample in a stratum that is larger, more variable, or cheaper to survey.
The total \(n\) is fixed by the budget. Putting \(n_h = n\,(N_hS_h/\sqrt{c_h})/\sum_k(N_kS_k/\sqrt{c_k})\) into the cost equation,
\[ n = \frac{(C - a)\sum_h N_hS_h/\sqrt{c_h}}{\sum_h N_hS_h\sqrt{c_h}} . \](The textbook's proof writes \(n_h\) in terms of \(n\) before \(n\) is known; this equation is what determines it.) If every \(c_h = c_0\), then \(n = (C - a)/c_0\) and the allocation is Neyman's; if moreover every \(S_h\) is equal, it is proportional.
Setting \(\text{Var}(\bar y_{st}) = V_0\) with the same allocation and solving for \(n\):
\[ n = \frac{\big(\sum_h N_hS_h\sqrt{c_h}\big)\big(\sum_h N_hS_h/\sqrt{c_h}\big)}{N^2V_0 + \sum_h N_hS_h^2}, \]which for equal costs is \(n = \big(\sum N_hS_h\big)^2/\big(N^2V_0 + \sum N_hS_h^2\big)\).
The formulas rarely give whole numbers. Round each \(n_h\) down, then give the units still needed to the strata with the largest fractional parts, so that the \(n_h\) add up to \(n\). Every stratum keeps at least one unit (two, if its variance is to be estimated).
Define population variance decomposition:
Within-stratum variance \(\sum W_h S_h^2\) plus between-stratum variance.
Variance ordering (for the same total \(n\)):
Equality holds only when:
Relative efficiency of stratified over SRSWOR:
From Example 2 of Section 5: \(W_1 = 0.6,\; W_2 = 0.4;\; S_1^2 = 16, S_2^2 = 81\).
\(\sum W_h S_h^2 = 0.6(16) + 0.4(81) = 9.6 + 32.4 = 42.0\). With \(n = 100\) and large \(N\) (\(f \to 0\)):
\(\text{Var}_{prop} = 42.0/100 = 0.42\). To compute SRSWOR variance we need \(S^2\) (population). If between-stratum mean square is, say, 10: \(S^2 \approx 42 + 10 = 52\) ⇒ \(\text{Var}_{SRS} = 52/100 = 0.52\). Gain ≈ 24 %.
Continuing: \(\sum W_h S_h = 0.6(4) + 0.4(9) = 2.4 + 3.6 = 6\). \(\text{Var}_{opt} = 6^2/100 - (\text{small term}) = 0.36\) (ignoring the FPC small term).
So \(\text{Var}_{opt} = 0.36 < \text{Var}_{prop} = 0.42 < \text{Var}_{SRS} = 0.52\). Optimum allocation reduces variance further when stratum SDs differ.
With \(\bar S = \sum W_hS_h\),
\[ \text{Var}_{prop} - \text{Var}_{opt} = \Big(\frac1n - \frac1N\Big)\sum W_hS_h^2 - \frac1n\Big(\sum W_hS_h\Big)^2 + \frac1N\sum W_hS_h^2 \] \[ = \frac1n\Big[\sum W_hS_h^2 - \bar S^2\Big] = \frac1n\sum_h W_h(S_h - \bar S)^2 \;\ge\; 0 , \]with equality only when all the \(S_h\) are equal.
The exact analysis of variance of the population is
\[ (N-1)S^2 = \sum_h (N_h - 1)S_h^2 + \sum_h N_h(\bar Y_h - \bar Y)^2 . \]If every \(N_h\) is large, \(N_h - 1 \approx N_h\) and \(N - 1 \approx N\), so \(S^2 \approx \sum W_hS_h^2 + \sum W_h(\bar Y_h - \bar Y)^2\), and
\[ \text{Var}_{SRS} \approx \text{Var}_{prop} + \Big(\frac1n - \frac1N\Big)\sum_h W_h(\bar Y_h - \bar Y)^2 \;\ge\; \text{Var}_{prop} . \]Without the approximation the difference is
\[ \text{Var}_{SRS} - \text{Var}_{prop} = \frac{N-n}{nN(N-1)}\Big[\sum_h N_h(\bar Y_h - \bar Y)^2 - \frac1N\sum_h (N - N_h)S_h^2\Big], \]which can be negative when the stratum means hardly differ. For example, strata \(\{1, 5, 9\}\) and \(\{2, 5, 8\}\) have equal means, and with \(n = 2\): \(\text{Var}_{SRS} = 10/3\) but \(\text{Var}_{prop} = 25/6\). Stratification pays when the strata really differ.
The efficiency of a stratified design over SRS is the ratio \(E = \text{Var}_{SRS}/\text{Var}_{st}\); the gain in efficiency is \(E - 1 = (\text{Var}_{SRS} - \text{Var}_{st})/\text{Var}_{st}\), usually given as a percentage. The textbook defines the first and computes the second in its problems; the worked problems below give both.
| Allocation | \(n_h\) | Variance of \(\bar y_{st}\) |
|---|---|---|
| Equal | \(n/L\) | \(\dfrac{L}{n}\sum W_h^2 S_h^2 (1 - n/(LN_h))\) |
| Proportional | \(n W_h\) | \((1 - f)/n \sum W_h S_h^2\) |
| Optimum (Neyman) | \(n \dfrac{W_h S_h}{\sum W_k S_k}\) | \(\dfrac{1}{n}\left(\sum W_h S_h\right)^2 - \dfrac{1}{N}\sum W_h S_h^2\) |
Two problems in the textbook's order, the first with proportional allocation only, the second comparing SRS, proportional and optimum allocation, followed by the three exercises with their answers checked. Each needs the same four quantities: the stratum weights, the within-stratum mean squares \(S_h^2\), the overall \(S^2\) (from the analysis of variance of §6), and the variance formulas of §§4–5.
Source note. Every figure was recomputed exactly. Worked Problem 1 agrees with the textbook up to rounding. In Worked Problem 2 the textbook's optimum variance leaves out one of its two terms, and one product is mis-multiplied; both are corrected below and change the conclusion about how much optimum allocation gains. Two of the three exercise answers also need correcting.
A population of 400 students belongs to two institutions:
| Institution | Students \(N_h\) | Mean \(\bar Y_h\) | SD \(\sigma_h\) |
|---|---|---|---|
| I | 300 | 50 | 20 |
| II | 100 | 40 | 10 |
Draw a sample of 40 by proportional allocation, find the variance of the estimated mean, and compare with SRSWOR.
Allocation. \(n_h = nN_h/N\): \(n_1 = \tfrac{40}{400}\times 300 = 30\), \(n_2 = \tfrac{40}{400}\times 100 = 10\).
Mean squares. The SDs are population SDs (divisor \(N_h\)), so \(S_h^2 = \frac{N_h}{N_h - 1}\sigma_h^2\): \(S_1^2 = \tfrac{300}{299}(400) = 401.34\), \(S_2^2 = \tfrac{100}{99}(100) = 101.01\). Also \(\bar Y = (15000 + 4000)/400 = 47.5\), and since \((N_h - 1)S_h^2 = N_h\sigma_h^2\),
\[ S^2 = \frac{1}{N-1}\Big[\sum N_h\sigma_h^2 + \sum N_h\bar Y_h^2 - N\bar Y^2\Big] \] \[ = \frac{130000 + 910000 - 400(47.5)^2}{399} = \frac{137500}{399} = 344.61 . \]Variances.
\[ \text{Var}_{prop} = \frac{N-n}{N^2n}\sum N_hS_h^2 = \frac{360}{400^2 \times 40}(130502.35) = 7.34, \] \[ \text{Var}_{SRS} = \frac{N-n}{N}\cdot\frac{S^2}{n} = \frac{360}{400}\cdot\frac{344.61}{40} = 7.75 . \]Efficiency \(7.75/7.34 = 1.056\): a gain of 5.6%. It is small because the two means (50 and 40) differ little compared with the spread within each institution.
Rounding note. The textbook rounds \(N_hS_h^2\) to 130503 and computes the gain from the rounded variances, \((7.75 - 7.34)/7.34 = 5.59\%\); at full precision it is 5.63%.
Policies held in a city, stratified by the holder's age (amounts in lakh rupees):
| Stratum | Age group | Policies \(N_h\) | Mean \(\bar Y_h\) | \(S_h\) |
|---|---|---|---|---|
| 1 | 0–20 | 146 | 12 | 2.1 |
| 2 | 20–40 | 224 | 20 | 4.6 |
| 3 | 40–60 | 123 | 8 | 1.2 |
| 4 | 60 and above | 48 | 4 | 0.5 |
For a sample of 50 policies find (i) the sampling variance of the estimated total amount under (a) SRSWOR, (b) proportional and (c) optimum allocation; (ii) the stratum sample sizes; (iii) the gains in efficiency over SRS.
Working table.
| \(h\) | \(N_h\bar Y_h\) | \(N_h\bar Y_h^2\) | \(N_hS_h\) | \(N_hS_h^2\) | \((N_h-1)S_h^2\) |
|---|---|---|---|---|---|
| 1 | 1752 | 21024 | 306.6 | 643.86 | 639.45 |
| 2 | 4480 | 89600 | 1030.4 | 4739.84 | 4718.68 |
| 3 | 984 | 7872 | 147.6 | 177.12 | 175.68 |
| 4 | 192 | 768 | 24 | 12 | 11.75 |
| Total | 7408 | 119264 | 1508.6 | 5572.82 | 5545.56 |
\(N = 541\), \(\bar Y = 7408/541 = 13.6932\), and
\[ S^2 = \frac{5545.56 + 119264 - 541(13.6932)^2}{540} = 43.279 . \](i) Variances. The total is estimated by \(N\bar y\), so its variance is \(N^2 = 292681\) times that of the mean.
\[ \text{(a)}\;\; \text{Var}_{SRS}(\bar y) = \frac{491}{541}\cdot\frac{43.279}{50} = 0.7856, \qquad \text{Var}(\hat Y) = 229924.5 , \] \[ \text{(b)}\;\; \text{Var}_{prop} = \frac{491}{541^2 \times 50}(5572.82) = 0.1870, \qquad \text{Var}(\hat Y) = 54725.1 , \] \[ \text{(c)}\;\; \text{Var}_{opt} = \frac{(1508.6)^2}{50 \times 541^2} - \frac{5572.82}{541^2} = 0.1555 - 0.0190 = 0.1365, \] \[ \text{Var}(\hat Y) = 39944.7 . \](ii) Sample sizes. Proportional: \(n_h = 50N_h/541 = 13.49, 20.70, 11.37, 4.44\), rounded (largest fractions first) to 14, 21, 11, 4. Optimum: \(n_h = 50N_hS_h/1508.6 = 10.16, 34.15, 4.89, 0.80\), rounded to 10, 34, 5, 1.
(iii) Gains. Proportional: \((0.7856 - 0.1870)/0.1870 = 3.20\), i.e. 320% (efficiency 4.20). Optimum: \((0.7856 - 0.1365)/0.1365 = 4.76\), i.e. 476% (efficiency 5.76). Stratifying by age removes most of the variation, because the four mean amounts are far apart.
Corrections. (1) The textbook's optimum variance stops at the first term, 0.1555, leaving out \(-\sum N_hS_h^2/N^2 = -0.0190\); its total 45511.90 and its gain of 406% both follow from that omission (the correct gain is 476%). (2) Its proportional total, 52097.22, is not \(541^2 \times 0.1870\), which is 54731. (3) Its working table gives \(\bar Y_4^2\) as 36; it is \(4^2 = 16\), and its next column (\(48 \times 16 = 768\)) uses the right value. (4) It rounds \(\bar Y\) to 13.69 before squaring, which moves \(S^2\) from 43.28 to 43.37 and \(\text{Var}_{SRS}\) from 0.7856 to 0.7872 (total 230398.48). A difference of large, nearly equal numbers needs \(\bar Y\) to full precision.