Topics Covered
Contents
- 1. Simple Random Sampling (SRS)
- 2. Stratified Random Sampling
- 3. Systematic Sampling
- 4. Ratio & Regression Estimation
- 5. Cluster Sampling
- 6. Double Sampling (Two-Phase)
- 7. PPS Sampling (Varying Probability)
- 8. One-way ANOVA
- 9. Two-way ANOVA
- 10. Principles of Design of Experiments
- 11. CRD, RBD, LSD
- 12. Missing Plot Techniques
- 13. Factorial Experiments — $2^2$, $2^3$
- 14. Confounding
- 15. BIBD, Connectedness, Orthogonality
Topic Overview — What & Why
Unit III combines two pillars of applied statistics — how data are gathered (sampling) and how experiments are arranged (design). A flawed design or sample makes even the most sophisticated analysis worthless.
- SRS, stratified, systematic sampling: different ways to pick observations from a population, each with its own variance formula and trade-offs (cost, simplicity, precision).
- Ratio & regression estimation: use auxiliary information $X$ correlated with the variable of interest to reduce variance — standard tool in agricultural and economic surveys.
- Cluster sampling: when individual selection is costly (e.g., visiting villages), pick groups; pay the price of higher variance for cheaper fieldwork.
- Double sampling: two-phase scheme — cheap proxy first, expensive measurement second.
- Sampling with varying probabilities (PPS): when units differ in size, give larger ones higher selection probability; Hansen-Hurwitz / Horvitz-Thompson estimators give unbiased totals.
- One-way and two-way ANOVA: partition variation into "explained" and "residual"; basis of comparing means across groups.
- Principles of design (Replication, Randomisation, Local Control): Fisher's three laws — ensure unbiased estimates of error and protect against systematic bias.
- CRD, RBD, LSD: the canonical designs — from no blocking, to one blocking factor, to two simultaneously.
- Missing plot techniques: patch-up methods (Yates) for analysis when one or more observations are lost.
- Factorial & confounding: study several factors simultaneously; sacrifice high-order interactions to fit large designs into small blocks.
- BIBD, connectedness, orthogonality: when block size is smaller than the number of treatments, BIBDs ensure every pair of treatments appears together equally often, preserving estimability.
1. Simple Random Sampling (SRS)
Why this section? SRS is the baseline against which every other sampling design is measured. Understanding the mean, variance, and unbiased variance estimator under SRS is the first thing every survey statistician must master.
Sample Mean as Estimator of Population Mean
For SRSWOR: $$\bar y=\tfrac{1}{n}\sum_{i=1}^n y_i,\quad E(\bar y)=\bar Y,\quad V(\bar y)=\frac{N-n}{Nn}S^2=\frac{1-f}{n}S^2,$$ where $f=n/N$ is sampling fraction and $S^2=\tfrac{1}{N-1}\sum_{i=1}^N(Y_i-\bar Y)^2.$For SRSWR: $V(\bar y)=\sigma^2/n$ where $\sigma^2=\tfrac{1}{N}\sum(Y_i-\bar Y)^2.$ SRSWOR is more efficient.
Estimator of Variance
$\hat V(\bar y)=\tfrac{1-f}{n}s^2$ where $s^2=\tfrac{1}{n-1}\sum(y_i-\bar y)^2.$🌍 Where it's used in real life
- Lottery-style choice of survey respondents.
- Auditing a random sample of invoices.
- Exit polls picking random voters.
- Quality check on random items from a batch.
- Randomly choosing patients for a trial.
2. Stratified Random Sampling
Divide population of $N$ into $L$ non-overlapping strata of sizes $N_1,\ldots,N_L$. Take SRS of $n_h$ from stratum $h$.
Stratified Mean and Variance
$$\bar y_{st}=\sum_{h=1}^L W_h\bar y_h,\quad W_h=N_h/N.$$ $$V(\bar y_{st})=\sum_h W_h^2\frac{1-f_h}{n_h}S_h^2,\quad f_h=n_h/N_h.$$Allocation Methods
- Equal allocation: $n_h=n/L.$
- Proportional: $n_h=nW_h.$
- Neyman (optimum for fixed $n$): $n_h\propto W_hS_h$, $n_h=\frac{nN_hS_h}{\sum N_kS_k}.$
- Optimum with cost $C=\sum c_h n_h$: $n_h\propto\frac{W_h S_h}{\sqrt{c_h}}.$
Efficiency
$V(\bar y_{Ney})\le V(\bar y_{prop})\le V(\bar y_{SRS}).$🌍 Where it's used in real life
- Surveying by age or income groups so all are represented.
- National health surveys split by state.
- Market research across customer segments.
- School exam sampling across big and small schools.
- Polls balanced by demographic groups.
3. Systematic Sampling
Choose random start $r\in\{1,\ldots,k\}$ where $k=N/n$, then take every $k$th unit. Yields one of $k$ possible samples.
Variance
$$V(\bar y_{sys})=\frac{S^2}{n}\big[1+(n-1)\rho_w\big]$$ where $\rho_w$ is intraclass correlation among units in the same systematic sample.- Efficient when $\rho_w<0$ (population heterogeneous within sample).
- Easy to implement; one sample → no unbiased variance estimator unless special assumptions.
🌍 Where it's used in real life
- Picking every 10th item on an assembly line.
- Auditing every 50th transaction.
- Selecting every k-th house in a locality survey.
- Sampling records from a register.
- Environmental readings at fixed distances.
4. Ratio and Regression Estimators
Ratio Estimator
Use auxiliary $X$ correlated with $Y$, with known $\bar X$: $$\hat{\bar Y}_R=\frac{\bar y}{\bar x}\bar X.$$ Approximate variance: $$V(\hat{\bar Y}_R)\approx\frac{1-f}{n}\big(S_y^2+R^2 S_x^2-2R S_{xy}\big),\quad R=\bar Y/\bar X.$$ Efficient if $\rho_{xy}>\frac{1}{2}\frac{C_x}{C_y}$ (CV ratio condition).Regression Estimator
$$\hat{\bar Y}_{lr}=\bar y+b(\bar X-\bar x),\quad b=S_{xy}/S_x^2.$$ $V(\hat{\bar Y}_{lr})\approx\frac{1-f}{n}S_y^2(1-\rho^2)$ — always at least as efficient as the ratio estimator.🌍 Where it's used in real life
- Estimating crop yield using field area as a helper.
- Estimating this year's sales from last year's totals.
- Forest timber volume from tree counts.
- Household spending from known income.
- Fish-population estimates from tagged samples.
5. Cluster Sampling
Population grouped into $N$ clusters; SRS of $n$ clusters; all units within selected clusters surveyed.
Equal Clusters of size $M$
$$\bar y=\tfrac{1}{nM}\sum_{i\in s}\sum_j y_{ij},\quad V(\bar y)=\frac{1-f}{n}\frac{S_b^2}{M},$$ where $S_b^2=\frac{1}{N-1}\sum_i M(\bar y_i-\bar Y)^2$ measures between-cluster variance.Intracluster correlation $\rho$
$V(\bar y_{cl})=\frac{1-f}{nM}S^2[1+(M-1)\rho].$ Cluster sampling efficient only if $\rho<0$ (clusters internally heterogeneous).Unequal Clusters
Use ratio-type estimator $\hat{\bar Y}=\sum y_i/\sum M_i$ or PPS sampling.🌍 Where it's used in real life
- Surveying whole villages instead of scattered homes.
- Sampling entire schools or classrooms.
- Health surveys by chosen city blocks.
- Auditing every item in randomly chosen cartons.
- Market surveys through selected retail outlets.
6. Double (Two-Phase) Sampling
First sample of $n'$ units gathers cheap auxiliary $X$; second sample of $n\subset n'$ measures $Y$. Useful for ratio/regression when $\bar X$ is unknown.
- Double sampling for stratification.
- Double sampling with regression: $\hat{\bar Y}_{lr}=\bar y+b(\bar x'-\bar x)$.
🌍 Where it's used in real life
- A cheap screening survey followed by detailed study.
- Quick income question, then a full expenditure survey.
- Soil test in the field, then lab analysis.
- Rapid disease screening, then confirmatory tests.
- Web survey, then in-depth phone interviews.
7. Sampling with Varying Probability (PPS)
PPSWR — Hansen–Hurwitz Estimator
Probability of selecting unit $i$ is $p_i$. Estimator of total: $$\hat Y_{HH}=\frac{1}{n}\sum_{i=1}^n\frac{y_i}{p_i},\quad V=\frac{1}{n}\sum_{i=1}^N p_i\Big(\frac{Y_i}{p_i}-Y\Big)^2.$$PPSWOR — Horvitz–Thompson
$\pi_i$ = inclusion probability; joint $\pi_{ij}.$ $$\hat Y_{HT}=\sum_{i\in s}\frac{y_i}{\pi_i},\quad V=\sum_i\sum_j(\pi_{ij}-\pi_i\pi_j)\frac{Y_i Y_j}{\pi_i\pi_j}.$$ $\hat Y_{HT}$ is unbiased.Yates–Grundy Variance Form
$$\hat V_{YG}=\sum_{i<j}\frac{\pi_i\pi_j-\pi_{ij}}{\pi_{ij}}\Big(\frac{y_i}{\pi_i}-\frac{y_j}{\pi_j}\Big)^2.$$ Non-negative iff $\pi_{ij}\le\pi_i\pi_j.$Ordered vs Unordered Estimators
For PPSWOR with order, Murthy's unordered estimator improves efficiency by averaging over all orderings; non-negative variance estimator possible (Yates–Grundy form for fixed-size designs).🌍 Where it's used in real life
- Sampling large factories with higher probability.
- Choosing cities in proportion to population.
- Auditing big accounts more often than small ones.
- Retail audits that weight large stores more.
- Forest plots chosen in proportion to size.
8. One-way ANOVA (Fixed Effects)
Model: $y_{ij}=\mu+\alpha_i+\varepsilon_{ij},\,\varepsilon_{ij}\overset{\text{iid}}{\sim}N(0,\sigma^2),\,\sum\alpha_i=0.$ $i=1,\ldots,k;\,j=1,\ldots,n_i.$
| Source | df | SS | MS | F |
|---|---|---|---|---|
| Between | $k-1$ | $\sum n_i(\bar y_{i\cdot}-\bar y_{\cdot\cdot})^2$ | SSB/(k-1) | MSB/MSE |
| Within (Error) | $N-k$ | $\sum\sum(y_{ij}-\bar y_{i\cdot})^2$ | SSE/(N-k) | |
| Total | $N-1$ | $\sum\sum(y_{ij}-\bar y_{\cdot\cdot})^2$ |
Test $H_0:\alpha_1=\cdots=\alpha_k=0$ using $F\sim F_{k-1,N-k}$ under $H_0.$
🌍 Where it's used in real life
- Comparing average yields of several fertilisers.
- Finding which teaching method scores best.
- Comparing mean sales across store layouts.
- Drug vs placebo vs alternative recovery times.
- Comparing average call times across centres.
9. Two-way ANOVA
Without Interaction
$y_{ij}=\mu+\alpha_i+\beta_j+\varepsilon_{ij}$.| Source | df |
|---|---|
| Rows (A) | $a-1$ |
| Columns (B) | $b-1$ |
| Error | $(a-1)(b-1)$ |
| Total | $ab-1$ |
With Interaction (replicated)
$y_{ijk}=\mu+\alpha_i+\beta_j+(\alpha\beta)_{ij}+\varepsilon_{ijk}$, df: A:$a-1$, B:$b-1$, AB:$(a-1)(b-1)$, Error:$ab(n-1)$, Total:$abn-1.$🌍 Where it's used in real life
- Crop yield by fertiliser and irrigation together.
- Exam scores by method and school.
- Product ratings by design and price level.
- Reaction time by drug and age group.
- Fuel efficiency by engine and fuel type.
10. Principles of Design of Experiments
Three fundamental principles (Fisher):
- Replication: repeating treatments gives variance estimates and increases precision.
- Randomization: protects against systematic bias; validates inferential procedures.
- Local Control (Blocking): grouping experimental units into blocks of homogeneous units to reduce error variance.
🌍 Where it's used in real life
- Designing fair clinical trials (randomisation).
- Field trials blocked by soil fertility.
- A/B/n testing on websites (replication).
- Industrial process-improvement experiments.
- Taste tests that control for tasting order.
11. CRD, RBD, Latin Square Design
Completely Randomized Design (CRD)
- $N$ homogeneous units, $t$ treatments allocated at random.
- Analysis = one-way ANOVA.
- Suitable for laboratory experiments where units are uniform.
- Error df = $N-t.$
Randomized Block Design (RBD)
- $b$ blocks of $t$ units; randomization within each block.
- Analysis = two-way ANOVA without interaction.
- Error df = $(b-1)(t-1).$
- Efficiency relative to CRD: $RE=\frac{(b-1)\hat\sigma_b^2+b(t-1)\hat\sigma_e^2}{(bt-1)\hat\sigma_e^2}.$
Latin Square Design (LSD)
- $t\times t$ arrangement; each treatment appears once per row and once per column.
- Controls two sources of variation simultaneously.
- df: rows=$t-1$, cols=$t-1$, treatments=$t-1$, error=$(t-1)(t-2)$, total=$t^2-1.$
- Restriction: $t$ must equal number of treatments.
A B C D B A D C C D A B D C B A
🌍 Where it's used in real life
- Lab experiments on uniform samples (CRD).
- Field trials blocked into fertile strips (RBD).
- Trials controlling both rows and columns of a field (LSD).
- Machine trials blocking by operator and shift.
- Comparing tyres across the four car positions.
12. Missing Plot Techniques
Yates's method: estimate missing value by minimizing residual SS, then proceed with standard ANOVA but reduce error df by number of missing observations.
RBD: Single Missing Value
$$\hat y=\frac{bB+tT-G}{(b-1)(t-1)},$$ where $B,T,G$ are totals of the block, treatment, and grand total computed excluding the missing value.LSD: Single Missing Value
$$\hat y=\frac{t(R+C+T)-2G}{(t-1)(t-2)}.$$🌍 Where it's used in real life
- Estimating a lost yield reading in a field trial.
- Recovering a spoiled lab measurement.
- Handling an animal that drops out of a study.
- Filling a damaged plot in agriculture.
- Analysing data despite an equipment failure.
13. Factorial Experiments — $2^2$ and $2^3$
$2^2$ Design
Two factors A, B each at 2 levels. Treatments: $(1),a,b,ab.$ Main effects and interaction: $$A=\tfrac{1}{2r}[(a-1)(b+1)]=\tfrac{1}{2r}\big(a+ab-(1)-b\big).$$ $$B=\tfrac{1}{2r}\big(b+ab-(1)-a\big),\quad AB=\tfrac{1}{2r}\big((1)+ab-a-b\big).$$$2^3$ Design
8 treatment combinations: $(1),a,b,c,ab,ac,bc,abc.$ Effects (main A, B, C; two-factor AB, AC, BC; three-factor ABC) computed using Yates's algorithm.Yates's Algorithm
Sequence treatments in standard order; perform $k$ rounds of pairwise sum and difference to get factorial effect totals.🌍 Where it's used in real life
- Testing temperature and pressure effects together.
- Recipe tuning (sugar level × baking time).
- Website tests of headline × image × button.
- Chemical yield by catalyst × temperature.
- Crop trials of variety × spacing × fertiliser.
14. Confounding in Factorial Experiments
When block size is smaller than number of treatments, an effect (usually high-order interaction) is confounded (mixed) with block effect.
Types
- Complete confounding: same effect confounded in every replicate — totally lost.
- Partial confounding: different effects confounded in different replicates → recoverable information.
- Balanced confounding: all interactions of given order confounded equally often.
$2^3$ Example with ABC Confounded
Block I: $(1),ab,ac,bc$ Block II: $a,b,c,abc.$ Then ABC contrast = block contrast → ABC confounded.🌍 Where it's used in real life
- Fitting a big experiment into small batches.
- Industrial trials limited to a few runs a day.
- Agricultural blocks smaller than the treatment count.
- Giving up high-order effects to save runs.
- Clinical trials spread across limited facilities.
15. Incomplete Block Designs & BIBD
When block size $k<v$ (treatments), incomplete block designs are used.
Balanced Incomplete Block Design (BIBD)
Parameters $(v,b,r,k,\lambda)$:- $v$ treatments, $b$ blocks, each of size $k$.
- Each treatment in $r$ blocks.
- Every pair of treatments occurs together in exactly $\lambda$ blocks.
Intuition. When a block is too small to hold every treatment, no single block can compare all treatments directly. A BIBD restores balance by arranging that every pair of treatments appears together in the same number of blocks ($\lambda$); this symmetry makes all pairwise treatment comparisons estimable with equal precision, which is what “balanced” refers to.
Connectedness and Orthogonality
- Connected design: all elementary contrasts among treatment effects estimable.
- Orthogonal: all treatment contrasts orthogonal to block contrasts (RBD, LSD); intra-block analysis suffices.
- BIBD is connected but non-orthogonal — recovery of inter-block info is possible.
Intra-block Analysis
Treatment effect estimator: $\hat\tau_i=\frac{kQ_i}{\lambda v}$ where $Q_i=T_i-\frac{1}{k}\sum_{j\in B_i}B_j$ (adjusted treatment total).Inter-block Analysis & Recovery
Block totals also carry treatment information. Combined estimator weights intra-block by $1/\sigma_e^2$ and inter-block by $1/(\sigma_e^2+k\sigma_b^2).$🌍 Where it's used in real life
- Taste tests where each judge tries only some products.
- Balanced scheduling of tournament pairings.
- Multi-site trials where each site tests a subset.
- Comparing many crop varieties in small blocks.
- Paired-comparison survey designs.