Skip to the content

Topics Covered

Sample Mean as Estimator of Population Mean Estimator of Variance Stratified Mean and Variance Allocation Methods Ratio Estimator Regression Estimator Equal Clusters of size M Intracluster correlation ρ Unequal Clusters PPSWR — Hansen–Hurwitz Estimator PPSWOR — Horvitz–Thompson Yates–Grundy Variance Form

Topic Overview — What & Why

Unit III combines two pillars of applied statistics — how data are gathered (sampling) and how experiments are arranged (design). A flawed design or sample makes even the most sophisticated analysis worthless.

  • SRS, stratified, systematic sampling: different ways to pick observations from a population, each with its own variance formula and trade-offs (cost, simplicity, precision).
  • Ratio & regression estimation: use auxiliary information $X$ correlated with the variable of interest to reduce variance — standard tool in agricultural and economic surveys.
  • Cluster sampling: when individual selection is costly (e.g., visiting villages), pick groups; pay the price of higher variance for cheaper fieldwork.
  • Double sampling: two-phase scheme — cheap proxy first, expensive measurement second.
  • Sampling with varying probabilities (PPS): when units differ in size, give larger ones higher selection probability; Hansen-Hurwitz / Horvitz-Thompson estimators give unbiased totals.
  • One-way and two-way ANOVA: partition variation into "explained" and "residual"; basis of comparing means across groups.
  • Principles of design (Replication, Randomisation, Local Control): Fisher's three laws — ensure unbiased estimates of error and protect against systematic bias.
  • CRD, RBD, LSD: the canonical designs — from no blocking, to one blocking factor, to two simultaneously.
  • Missing plot techniques: patch-up methods (Yates) for analysis when one or more observations are lost.
  • Factorial & confounding: study several factors simultaneously; sacrifice high-order interactions to fit large designs into small blocks.
  • BIBD, connectedness, orthogonality: when block size is smaller than the number of treatments, BIBDs ensure every pair of treatments appears together equally often, preserving estimability.

1. Simple Random Sampling (SRS)

Why this section? SRS is the baseline against which every other sampling design is measured. Understanding the mean, variance, and unbiased variance estimator under SRS is the first thing every survey statistician must master.

SRSWOR: $n$ units chosen from $N$ such that each subset of size $n$ is equally likely. Each unit's inclusion probability $=n/N.$

Sample Mean as Estimator of Population Mean

For SRSWOR: $$\bar y=\tfrac{1}{n}\sum_{i=1}^n y_i,\quad E(\bar y)=\bar Y,\quad V(\bar y)=\frac{N-n}{Nn}S^2=\frac{1-f}{n}S^2,$$ where $f=n/N$ is sampling fraction and $S^2=\tfrac{1}{N-1}\sum_{i=1}^N(Y_i-\bar Y)^2.$

For SRSWR: $V(\bar y)=\sigma^2/n$ where $\sigma^2=\tfrac{1}{N}\sum(Y_i-\bar Y)^2.$ SRSWOR is more efficient.

Estimator of Variance

$\hat V(\bar y)=\tfrac{1-f}{n}s^2$ where $s^2=\tfrac{1}{n-1}\sum(y_i-\bar y)^2.$
EXAMPLE 1 $N=500$, $S^2=400$, $n=50$. $V(\bar y)=\frac{1-0.1}{50}(400)=7.2$, SD $\approx 2.68.$
EXAMPLE 2 For population $\{2,4,6,8\}$, $\bar Y=5,\,S^2=20/3.$ SRSWOR of size 2: 6 possible samples. $E(\bar y)=5$ (unbiased). $V(\bar y)=\frac{4-2}{4\cdot 2}(20/3)=5/3.$

🌍 Where it's used in real life

  1. Lottery-style choice of survey respondents.
  2. Auditing a random sample of invoices.
  3. Exit polls picking random voters.
  4. Quality check on random items from a batch.
  5. Randomly choosing patients for a trial.

2. Stratified Random Sampling

Divide population of $N$ into $L$ non-overlapping strata of sizes $N_1,\ldots,N_L$. Take SRS of $n_h$ from stratum $h$.

Stratified Mean and Variance

$$\bar y_{st}=\sum_{h=1}^L W_h\bar y_h,\quad W_h=N_h/N.$$ $$V(\bar y_{st})=\sum_h W_h^2\frac{1-f_h}{n_h}S_h^2,\quad f_h=n_h/N_h.$$

Allocation Methods

Efficiency

$V(\bar y_{Ney})\le V(\bar y_{prop})\le V(\bar y_{SRS}).$
EXAMPLE 1 Population: 60% rural ($S_1=10$), 40% urban ($S_2=20$), $n=100.$ Proportional: $n_1=60, n_2=40.$ Neyman: $n_1\propto 0.6\times10=6,\,n_2\propto 0.4\times20=8.$ So $n_1=42.86\approx 43, n_2=57.$
EXAMPLE 2 $N=1000,W_1=W_2=0.5, S_1=4,S_2=8,n=80.$ Proportional: $n_1=n_2=40$, $V=\sum W_h^2 S_h^2/n_h=0.25(16/40+64/40)=0.5.$ Neyman: $n_1=80\cdot 4/(4+8)=26.7$, $n_2=53.3$, smaller variance.

🌍 Where it's used in real life

  1. Surveying by age or income groups so all are represented.
  2. National health surveys split by state.
  3. Market research across customer segments.
  4. School exam sampling across big and small schools.
  5. Polls balanced by demographic groups.

3. Systematic Sampling

Choose random start $r\in\{1,\ldots,k\}$ where $k=N/n$, then take every $k$th unit. Yields one of $k$ possible samples.

Variance

$$V(\bar y_{sys})=\frac{S^2}{n}\big[1+(n-1)\rho_w\big]$$ where $\rho_w$ is intraclass correlation among units in the same systematic sample.
EXAMPLE 1 $N=20,n=4,k=5$. Random start $r=3\Rightarrow$ sample $\{3,8,13,18\}.$
EXAMPLE 2 With $\rho_w=-0.05,n=10,S^2=100$: $V_{sys}=\frac{100}{10}[1+9(-0.05)]=10(0.55)=5.5$, vs $V_{srs}=10$ — half the variance.

🌍 Where it's used in real life

  1. Picking every 10th item on an assembly line.
  2. Auditing every 50th transaction.
  3. Selecting every k-th house in a locality survey.
  4. Sampling records from a register.
  5. Environmental readings at fixed distances.

4. Ratio and Regression Estimators

Ratio Estimator

Use auxiliary $X$ correlated with $Y$, with known $\bar X$: $$\hat{\bar Y}_R=\frac{\bar y}{\bar x}\bar X.$$ Approximate variance: $$V(\hat{\bar Y}_R)\approx\frac{1-f}{n}\big(S_y^2+R^2 S_x^2-2R S_{xy}\big),\quad R=\bar Y/\bar X.$$ Efficient if $\rho_{xy}>\frac{1}{2}\frac{C_x}{C_y}$ (CV ratio condition).

Regression Estimator

$$\hat{\bar Y}_{lr}=\bar y+b(\bar X-\bar x),\quad b=S_{xy}/S_x^2.$$ $V(\hat{\bar Y}_{lr})\approx\frac{1-f}{n}S_y^2(1-\rho^2)$ — always at least as efficient as the ratio estimator.
EXAMPLE 1 For estimating crop yield using area as auxiliary: $\bar X=10$, $\bar y=50,\bar x=8\Rightarrow\hat Y_R=(50/8)(10)=62.5.$
EXAMPLE 2 $\bar y=100,\bar x=20,\bar X=22,b=4\Rightarrow\hat Y_{lr}=100+4(22-20)=108.$

🌍 Where it's used in real life

  1. Estimating crop yield using field area as a helper.
  2. Estimating this year's sales from last year's totals.
  3. Forest timber volume from tree counts.
  4. Household spending from known income.
  5. Fish-population estimates from tagged samples.

5. Cluster Sampling

Population grouped into $N$ clusters; SRS of $n$ clusters; all units within selected clusters surveyed.

Equal Clusters of size $M$

$$\bar y=\tfrac{1}{nM}\sum_{i\in s}\sum_j y_{ij},\quad V(\bar y)=\frac{1-f}{n}\frac{S_b^2}{M},$$ where $S_b^2=\frac{1}{N-1}\sum_i M(\bar y_i-\bar Y)^2$ measures between-cluster variance.

Intracluster correlation $\rho$

$V(\bar y_{cl})=\frac{1-f}{nM}S^2[1+(M-1)\rho].$ Cluster sampling efficient only if $\rho<0$ (clusters internally heterogeneous).

Unequal Clusters

Use ratio-type estimator $\hat{\bar Y}=\sum y_i/\sum M_i$ or PPS sampling.
EXAMPLE 1 Survey 100 villages out of 1000, all households per chosen village — typical cluster design; cheap travel cost.
EXAMPLE 2 $N=20$ clusters of $M=5$ each, $S_b^2=8,n=4.$ $V(\bar y)=\frac{0.8}{4}\cdot\frac{8}{5}=0.32.$

🌍 Where it's used in real life

  1. Surveying whole villages instead of scattered homes.
  2. Sampling entire schools or classrooms.
  3. Health surveys by chosen city blocks.
  4. Auditing every item in randomly chosen cartons.
  5. Market surveys through selected retail outlets.

6. Double (Two-Phase) Sampling

First sample of $n'$ units gathers cheap auxiliary $X$; second sample of $n\subset n'$ measures $Y$. Useful for ratio/regression when $\bar X$ is unknown.

EXAMPLE 1 First-phase $n'=500$ households measure income (cheap proxy); second-phase $n=100$ measure expenditure expensively. Use regression to combine.
EXAMPLE 2 Approximate variance: $V\approx\frac{1-f}{n}S_y^2(1-\rho^2)+\frac{1}{n'}S_y^2\rho^2.$

🌍 Where it's used in real life

  1. A cheap screening survey followed by detailed study.
  2. Quick income question, then a full expenditure survey.
  3. Soil test in the field, then lab analysis.
  4. Rapid disease screening, then confirmatory tests.
  5. Web survey, then in-depth phone interviews.

7. Sampling with Varying Probability (PPS)

PPSWR — Hansen–Hurwitz Estimator

Probability of selecting unit $i$ is $p_i$. Estimator of total: $$\hat Y_{HH}=\frac{1}{n}\sum_{i=1}^n\frac{y_i}{p_i},\quad V=\frac{1}{n}\sum_{i=1}^N p_i\Big(\frac{Y_i}{p_i}-Y\Big)^2.$$

PPSWOR — Horvitz–Thompson

$\pi_i$ = inclusion probability; joint $\pi_{ij}.$ $$\hat Y_{HT}=\sum_{i\in s}\frac{y_i}{\pi_i},\quad V=\sum_i\sum_j(\pi_{ij}-\pi_i\pi_j)\frac{Y_i Y_j}{\pi_i\pi_j}.$$ $\hat Y_{HT}$ is unbiased.

Yates–Grundy Variance Form

$$\hat V_{YG}=\sum_{i<j}\frac{\pi_i\pi_j-\pi_{ij}}{\pi_{ij}}\Big(\frac{y_i}{\pi_i}-\frac{y_j}{\pi_j}\Big)^2.$$ Non-negative iff $\pi_{ij}\le\pi_i\pi_j.$

Ordered vs Unordered Estimators

For PPSWOR with order, Murthy's unordered estimator improves efficiency by averaging over all orderings; non-negative variance estimator possible (Yates–Grundy form for fixed-size designs).
EXAMPLE 1 3 units $Y=(10,20,30),\,p=(0.2,0.3,0.5).$ HH: select 1 unit, $\hat Y=Y_i/p_i=$ 50,66.67,60 with prob 0.2,0.3,0.5. $E=10+20+30=60$ ✓.
EXAMPLE 2 Industrial firms — large factories chosen with higher probability proportional to employment; HT yields unbiased total output.

🌍 Where it's used in real life

  1. Sampling large factories with higher probability.
  2. Choosing cities in proportion to population.
  3. Auditing big accounts more often than small ones.
  4. Retail audits that weight large stores more.
  5. Forest plots chosen in proportion to size.

8. One-way ANOVA (Fixed Effects)

Model: $y_{ij}=\mu+\alpha_i+\varepsilon_{ij},\,\varepsilon_{ij}\overset{\text{iid}}{\sim}N(0,\sigma^2),\,\sum\alpha_i=0.$ $i=1,\ldots,k;\,j=1,\ldots,n_i.$

SourcedfSSMSF
Between$k-1$$\sum n_i(\bar y_{i\cdot}-\bar y_{\cdot\cdot})^2$SSB/(k-1)MSB/MSE
Within (Error)$N-k$$\sum\sum(y_{ij}-\bar y_{i\cdot})^2$SSE/(N-k)
Total$N-1$$\sum\sum(y_{ij}-\bar y_{\cdot\cdot})^2$

Test $H_0:\alpha_1=\cdots=\alpha_k=0$ using $F\sim F_{k-1,N-k}$ under $H_0.$

EXAMPLE 1 3 fertilizers, 5 plots each. SSB=120, SSE=80. F=(120/2)/(80/12)=60/6.67=9.0; if $F_{2,12,0.05}=3.89$, reject $H_0$.
EXAMPLE 2 For $k=2$ groups, ANOVA F equals $t^2$ from two-sample $t$-test (pooled variance).

🌍 Where it's used in real life

  1. Comparing average yields of several fertilisers.
  2. Finding which teaching method scores best.
  3. Comparing mean sales across store layouts.
  4. Drug vs placebo vs alternative recovery times.
  5. Comparing average call times across centres.

9. Two-way ANOVA

Without Interaction

$y_{ij}=\mu+\alpha_i+\beta_j+\varepsilon_{ij}$.
Sourcedf
Rows (A)$a-1$
Columns (B)$b-1$
Error$(a-1)(b-1)$
Total$ab-1$

With Interaction (replicated)

$y_{ijk}=\mu+\alpha_i+\beta_j+(\alpha\beta)_{ij}+\varepsilon_{ijk}$, df: A:$a-1$, B:$b-1$, AB:$(a-1)(b-1)$, Error:$ab(n-1)$, Total:$abn-1.$
EXAMPLE 1 3×4 design (no interaction): df rows=2, cols=3, error=6, total=11.
EXAMPLE 2 2×3 with 4 replicates: df A=1,B=2,AB=2,error=18,total=23.

🌍 Where it's used in real life

  1. Crop yield by fertiliser and irrigation together.
  2. Exam scores by method and school.
  3. Product ratings by design and price level.
  4. Reaction time by drug and age group.
  5. Fuel efficiency by engine and fuel type.

10. Principles of Design of Experiments

Three fundamental principles (Fisher):

EXAMPLE 1 A new drug tested on 30 patients: replication ensures we have ≥10 per group; randomization assigns patients; blocking by sex/age may control biological variation.
EXAMPLE 2 Field trial: replication of varieties across plots; randomized layout; blocking by soil fertility belts (RBD).

🌍 Where it's used in real life

  1. Designing fair clinical trials (randomisation).
  2. Field trials blocked by soil fertility.
  3. A/B/n testing on websites (replication).
  4. Industrial process-improvement experiments.
  5. Taste tests that control for tasting order.

11. CRD, RBD, Latin Square Design

Completely Randomized Design (CRD)

Randomized Block Design (RBD)

Latin Square Design (LSD)

CRDADCBADCABDCBCABDtreatments placed at random RBDCABDDACBBACDBCDAeach row (block) has all 4 LSDABCDBADCCDABDCBAeach row AND column has all 4
Three designs, increasing control. CRD imposes no structure; RBD blocks one nuisance factor so every treatment appears once per block (row); LSD blocks two, forcing each treatment once per row and once per column. More blocking removes more error variance — at the cost of more constraints.
EXAMPLE 1 LSD 4×4 with treatments A,B,C,D. Total observations 16; error df = 6. Standard layout:
A B C D
B A D C
C D A B
D C B A
EXAMPLE 2 RBD 5 blocks, 4 varieties: 20 plots; error df = 4×3 = 12.

🌍 Where it's used in real life

  1. Lab experiments on uniform samples (CRD).
  2. Field trials blocked into fertile strips (RBD).
  3. Trials controlling both rows and columns of a field (LSD).
  4. Machine trials blocking by operator and shift.
  5. Comparing tyres across the four car positions.

12. Missing Plot Techniques

Yates's method: estimate missing value by minimizing residual SS, then proceed with standard ANOVA but reduce error df by number of missing observations.

RBD: Single Missing Value

$$\hat y=\frac{bB+tT-G}{(b-1)(t-1)},$$ where $B,T,G$ are totals of the block, treatment, and grand total computed excluding the missing value.

LSD: Single Missing Value

$$\hat y=\frac{t(R+C+T)-2G}{(t-1)(t-2)}.$$
EXAMPLE 1 RBD with $b=4,t=5$, missing observation in block 2, treatment 3. Use formula above.
EXAMPLE 2 After estimating missing value, error df decreases by 1. Treatment SS is biased upward; corrected $T_{adj}=T_{obs}-\frac{(B+(t-1)T-G)^2}{t(t-1)^2(b-1)^2}.$

🌍 Where it's used in real life

  1. Estimating a lost yield reading in a field trial.
  2. Recovering a spoiled lab measurement.
  3. Handling an animal that drops out of a study.
  4. Filling a damaged plot in agriculture.
  5. Analysing data despite an equipment failure.

13. Factorial Experiments — $2^2$ and $2^3$

$2^2$ Design

Two factors A, B each at 2 levels. Treatments: $(1),a,b,ab.$ Main effects and interaction: $$A=\tfrac{1}{2r}[(a-1)(b+1)]=\tfrac{1}{2r}\big(a+ab-(1)-b\big).$$ $$B=\tfrac{1}{2r}\big(b+ab-(1)-a\big),\quad AB=\tfrac{1}{2r}\big((1)+ab-a-b\big).$$

$2^3$ Design

8 treatment combinations: $(1),a,b,c,ab,ac,bc,abc.$ Effects (main A, B, C; two-factor AB, AC, BC; three-factor ABC) computed using Yates's algorithm.

Yates's Algorithm

Sequence treatments in standard order; perform $k$ rounds of pairwise sum and difference to get factorial effect totals.
EXAMPLE 1 $2^2$: $(1)=20,a=30,b=25,ab=40,r=1.$ $A=\frac{1}{2}(30+40-20-25)=12.5.$ $B=\frac{1}{2}(25+40-20-30)=7.5.$ $AB=\frac{1}{2}(20+40-30-25)=2.5.$
EXAMPLE 2 $2^3$ requires 8 runs per replicate; with $r=2$ replicates, 16 runs, error df = 8.

🌍 Where it's used in real life

  1. Testing temperature and pressure effects together.
  2. Recipe tuning (sugar level × baking time).
  3. Website tests of headline × image × button.
  4. Chemical yield by catalyst × temperature.
  5. Crop trials of variety × spacing × fertiliser.

14. Confounding in Factorial Experiments

When block size is smaller than number of treatments, an effect (usually high-order interaction) is confounded (mixed) with block effect.

Types

$2^3$ Example with ABC Confounded

Block I: $(1),ab,ac,bc$   Block II: $a,b,c,abc.$ Then ABC contrast = block contrast → ABC confounded.
EXAMPLE 1 $2^3$ in blocks of 4, confound ABC: I:(1),ab,ac,bc; II:a,b,c,abc. Main effects estimated, ABC lost.
EXAMPLE 2 $2^4$ in blocks of 8: confound highest interaction ABCD. Block I = even-defining contrasts; Block II = odd.

🌍 Where it's used in real life

  1. Fitting a big experiment into small batches.
  2. Industrial trials limited to a few runs a day.
  3. Agricultural blocks smaller than the treatment count.
  4. Giving up high-order effects to save runs.
  5. Clinical trials spread across limited facilities.

15. Incomplete Block Designs & BIBD

When block size $k<v$ (treatments), incomplete block designs are used.

Balanced Incomplete Block Design (BIBD)

Parameters $(v,b,r,k,\lambda)$: Necessary conditions: $$bk=vr,\quad \lambda(v-1)=r(k-1),\quad b\ge v\ (\text{Fisher's inequality}).$$ Efficiency factor: $E=\lambda v/(rk).$

Intuition. When a block is too small to hold every treatment, no single block can compare all treatments directly. A BIBD restores balance by arranging that every pair of treatments appears together in the same number of blocks ($\lambda$); this symmetry makes all pairwise treatment comparisons estimable with equal precision, which is what “balanced” refers to.

Connectedness and Orthogonality

Intra-block Analysis

Treatment effect estimator: $\hat\tau_i=\frac{kQ_i}{\lambda v}$ where $Q_i=T_i-\frac{1}{k}\sum_{j\in B_i}B_j$ (adjusted treatment total).

Inter-block Analysis & Recovery

Block totals also carry treatment information. Combined estimator weights intra-block by $1/\sigma_e^2$ and inter-block by $1/(\sigma_e^2+k\sigma_b^2).$
EXAMPLE 1 $(v,b,r,k,\lambda)=(7,7,3,3,1)$ — Fano plane. Each pair of treatments occurs in exactly one block.
EXAMPLE 2 $(v,b,r,k,\lambda)=(4,6,3,2,1)$. Verify: $bk=12=vr=12$; $\lambda(v-1)=3=r(k-1)=3.$ ✓ Efficiency $E=\lambda v/(rk)=4/6=2/3.$

🌍 Where it's used in real life

  1. Taste tests where each judge tries only some products.
  2. Balanced scheduling of tournament pairings.
  3. Multi-site trials where each site tests a subset.
  4. Comparing many crop varieties in small blocks.
  5. Paired-comparison survey designs.