Topics Covered
Contents
- 1. Unbiasedness
- 2. Consistency
- 3. Method of Moments
- 4. Maximum Likelihood Estimation
- 5. Efficiency & UMVUE
- 6. Cramér–Rao Lower Bound
- 7. Sufficiency & Factorization Theorem
- 8. Minimal Sufficiency & Ancillarity
- 9. Completeness
- 10. Rao–Blackwell Theorem
- 11. Lehmann–Scheffé Theorem
- 12. Basu's Theorem
- 13. Method of Pivoting (Interval Estimation)
- 14. Confidence Intervals
- 15. Large-Sample CIs
- 16. Order Statistics & EDF
- 17. Rank Correlation (Spearman, Kendall)
Topic Overview — What & Why
Unit IV is about how to infer unknown population parameters from a sample. We need formal criteria to compare estimators (good vs bad) and a toolkit to construct optimal ones.
- Unbiasedness: on average, the estimator hits the truth. A clean, intuitive criterion — but not the only one (mean-square-error trades off bias and variance).
- Consistency: with enough data, the estimator converges to the true parameter. The minimum bar any estimator should clear.
- Method of Moments & MLE: two recipes for constructing estimators. MoM is simple; MLE is asymptotically efficient.
- Efficiency & UMVUE: among unbiased estimators, the one with smallest variance is "best". UMVUE is the gold standard.
- Cramér-Rao Lower Bound: a fundamental limit — no unbiased estimator can have smaller variance than $1/(nI(\theta))$. Tells us when we have done as well as possible.
- Sufficiency & factorisation theorem: a sufficient statistic captures all sample information about the parameter; you can compress data without losing inference power.
- Minimal sufficiency & ancillarity: the most-compressed sufficient summary; ancillary statistics carry no parametric information.
- Completeness: a technical condition that, combined with sufficiency, guarantees uniqueness of UMVUEs.
- Rao-Blackwell theorem: conditioning any unbiased estimator on a sufficient statistic strictly improves it (in MSE). The "improvement" engine.
- Lehmann-Scheffé theorem: if you have a complete sufficient statistic and any unbiased function of it, you have the UMVUE.
- Basu's theorem: a complete sufficient statistic is independent of any ancillary statistic — surprisingly powerful for proving independence.
- Pivoting & confidence intervals: turn estimators into ranges with stated coverage; one- and two-sample, normal and large-sample cases.
- Order statistics & EDF: the building blocks of nonparametric inference; Glivenko-Cantelli says EDF converges uniformly to the true CDF.
- Spearman & Kendall rank correlations: distribution-free measures of monotone association; useful when scales are ordinal or distributions are non-normal.
1. Unbiasedness
Why this section? Unbiasedness is often the first criterion students meet for evaluating an estimator. It says nothing about variability, so it must be combined with variance considerations — this is exactly what the rest of the unit develops.
Mean Squared Error: $\text{MSE}(T)=E(T-g(\theta))^2=\text{Var}(T)+\text{Bias}^2(T).$
🌍 Where it's used in real life
- Reporting an unbiased average income from a sample.
- Calibrating instruments to remove systematic error.
- Unbiased variance estimates in quality control.
- Fair estimates of vote share in polling.
- Estimating a defect rate without built-in bias.
2. Consistency
Sufficient condition: $E(T_n)\to\theta$ and $\text{Var}(T_n)\to 0.$
🌍 Where it's used in real life
- Trusting estimates to improve with bigger samples.
- Sensor readings settling down with more measurements.
- Poll accuracy rising with sample size.
- Model parameters stabilising as data grows.
- Long-run frequency estimating a true probability.
3. Method of Moments (MoM)
Equate sample moments to population moments and solve for parameters: $$m'_r=\tfrac{1}{n}\sum X_i^r=\mu'_r(\theta_1,\ldots,\theta_k),\quad r=1,\ldots,k.$$
🌍 Where it's used in real life
- Fast first estimates when fitting a distribution.
- Fitting income data to a gamma model.
- Estimating rates from average counts.
- Starting values for more complex fitting.
- Calibrating input distributions for simulation.
4. Maximum Likelihood Estimation
Properties
- Invariance: if $\hat\theta$ is MLE of $\theta$, then $g(\hat\theta)$ is MLE of $g(\theta).$
- Consistent under regularity conditions.
- Asymptotically normal: $\sqrt n(\hat\theta-\theta)\xrightarrow{d}N(0,1/I(\theta))$ where $I$ is Fisher information.
- Asymptotically efficient (achieves CRLB).
🌍 Where it's used in real life
- Fitting logistic models in medicine and credit scoring.
- Estimating failure rates in reliability.
- Training many machine-learning models.
- Estimating click and conversion probabilities.
- Estimating parameters in genetics.
5. Efficiency and UMVUE
Efficiency of unbiased $T$: $e(T)=\frac{1/I(\theta)}{V(T)}.$ $T$ is efficient if $e=1.$
🌍 Where it's used in real life
- Choosing the most precise estimator for given data.
- Minimising cost while hitting a target accuracy.
- Best unbiased estimate of a defect rate.
- Efficient survey estimators to save budget.
- Comparing estimators in simulation studies.
6. Cramér–Rao Lower Bound (CRLB)
Intuition. Fisher information $I(\theta)$ measures the average curvature (sharpness) of the log-likelihood in $\theta$: the more sharply the likelihood peaks around the true value, the more the data “pin down” the parameter, and the smaller the variance any unbiased estimator can achieve. The CRLB is therefore the precision ceiling imposed by the model itself, and an estimator meeting it is doing as well as the information in the data allows.
Equality holds iff there exists an exponential family structure with $T$ as canonical statistic.
🌍 Where it's used in real life
- The best possible accuracy of GPS positioning.
- Limits of radar and sonar estimation.
- Designing sensors to a precision target.
- Planning sample size for a required precision.
- Benchmarking estimator quality in signal processing.
7. Sufficiency & Factorization Theorem
Intuition. A sufficient statistic compresses all the parameter-relevant information in the sample into a smaller summary: once $T$ is known, the leftover randomness in the raw data carries no further information about $\theta$. The factorization theorem lets you certify sufficiency just by inspecting how $\theta$ enters the likelihood — if $\theta$ “touches” the data only through $T(x)$, then $T$ is sufficient.
Exponential Family
$f(x;\theta)=\exp\{\eta(\theta)T(x)-A(\theta)\}h(x)$ — $T$ is sufficient and complete.🌍 Where it's used in real life
- Summarising data by a few numbers with no loss.
- Storing totals instead of full datasets.
- Data compression for inference.
- Reporting the sample sum or mean in QC.
- Streaming statistics that keep running summaries.
8. Minimal Sufficiency & Ancillarity
$A$ is ancillary if its distribution does not depend on $\theta$.
Lehmann–Scheffé Method
$T(X)$ is minimal sufficient iff: $f(x;\theta)/f(y;\theta)$ is free of $\theta$ $\iff$ $T(x)=T(y).$🌍 Where it's used in real life
- Finding the smallest summary inference needs.
- Efficient data reduction in big-data pipelines.
- Spotting which extra data adds nothing.
- Designing compact monitoring statistics.
- Simplifying models to their essentials.
9. Completeness
Intuition. Completeness says the family is “rich enough” that the only unbiased estimator of $0$ built from $T$ is the trivial one ($g(T)\equiv 0$). This rules out two different unbiased functions of $T$ estimating the same quantity, which is precisely why a statistic that is both complete and sufficient delivers a unique UMVUE (Lehmann–Scheffé).
Exponential family is complete (under usual rank conditions on the natural parameter space).
🌍 Where it's used in real life
- Guaranteeing a unique best unbiased estimator.
- Avoiding ambiguous estimates in surveys.
- Theory behind reliable quality estimates.
- Ensuring a model is identifiable.
- Foundation for building UMVUEs.
10. Rao–Blackwell Theorem
- $T^*$ is a statistic (no $\theta$ dependence by sufficiency).
- $E(T^*)=g(\theta).$
- $V(T^*)\le V(W),$ with equality iff $W=T^*$ a.s.
🌍 Where it's used in real life
- Turning a rough estimator into a better one.
- Variance reduction in Monte-Carlo simulation.
- Sharpening survey estimates.
- Refining predicted probabilities in ML.
- Better reliability estimates from summaries.
11. Lehmann–Scheffé Theorem
Two equivalent paths to UMVUE:
- Find any unbiased estimator and Rao-Blackwellize using the complete sufficient $T$.
- Find $h(T)$ such that $E[h(T)]=g(\theta).$
🌍 Where it's used in real life
- Constructing the best unbiased estimator in practice.
- Best estimate of a probability like P(no defect).
- Optimal survey estimators.
- Reliability-metric estimation.
- A yardstick to judge other estimators.
12. Basu's Theorem
🌍 Where it's used in real life
- Proving sample mean and variance are independent (normal).
- Simplifying derivations of sampling distributions.
- Justifying separate study of location and spread.
- Underlying theory of the t-test.
- Independence arguments in probability proofs.
13. Method of Pivoting
🌍 Where it's used in real life
- Building a confidence interval for a mean.
- Margin of error in a poll.
- Tolerance intervals in manufacturing.
- Confidence bounds on a failure rate.
- Interval estimates for lab measurements.
14. Confidence Intervals — One and Two Sample
| Parameter | Conditions | $1-\alpha$ CI |
|---|---|---|
| $\mu$ | $\sigma$ known | $\bar X\pm z_{\alpha/2}\sigma/\sqrt n$ |
| $\mu$ | $\sigma$ unknown | $\bar X\pm t_{n-1,\alpha/2}S/\sqrt n$ |
| $\sigma^2$ | — | $\left(\frac{(n-1)S^2}{\chi^2_{n-1,\alpha/2}},\frac{(n-1)S^2}{\chi^2_{n-1,1-\alpha/2}}\right)$ |
| $\mu_1-\mu_2$ | $\sigma_1=\sigma_2$ unknown | $\bar X-\bar Y\pm t_{n_1+n_2-2,\alpha/2}S_p\sqrt{\tfrac{1}{n_1}+\tfrac{1}{n_2}}$ |
| $\sigma_1^2/\sigma_2^2$ | — | $\left(\frac{S_1^2/S_2^2}{F_{n_1-1,n_2-1,\alpha/2}},\frac{S_1^2/S_2^2}{F_{n_1-1,n_2-1,1-\alpha/2}}\right)$ |
Pooled Variance
$S_p^2=\frac{(n_1-1)S_1^2+(n_2-1)S_2^2}{n_1+n_2-2}.$🌍 Where it's used in real life
- The "±3% margin of error" in election polls.
- A drug's likely effect range in a trial.
- Confidence range for average delivery time.
- Control limits for a process mean.
- Confidence interval for a conversion rate.
15. Large-Sample Confidence Intervals
Based on CLT and consistency: if $\hat\theta$ is asymptotically normal, $$\hat\theta\pm z_{\alpha/2}\,\text{SE}(\hat\theta).$$
Common Cases
- Proportion: $\hat p\pm z_{\alpha/2}\sqrt{\hat p(1-\hat p)/n}.$
- Poisson rate: $\hat\lambda\pm z_{\alpha/2}\sqrt{\hat\lambda/n}.$
- Wilson score interval (proportion): $\frac{\hat p+z^2/2n\pm z\sqrt{\hat p(1-\hat p)/n+z^2/4n^2}}{1+z^2/n}.$
🌍 Where it's used in real life
- Big-survey approve/disapprove proportions.
- Conversion-rate intervals in web analytics.
- Incidence-rate intervals in epidemiology.
- Event-rate confidence (accidents per month).
- Large-sample intervals in A/B testing.
16. Order Statistics & Empirical Distribution Function
Distribution of Order Statistics
$X_{(1)}\le X_{(2)}\le\cdots\le X_{(n)}.$ $$f_{X_{(k)}}(x)=\frac{n!}{(k-1)!(n-k)!}f(x)F(x)^{k-1}(1-F(x))^{n-k}.$$ Joint of $(X_{(j)},X_{(k)})$, $j<k$: $$f_{j,k}(x,y)=\frac{n!}{(j-1)!(k-j-1)!(n-k)!}f(x)f(y)F(x)^{j-1}[F(y)-F(x)]^{k-j-1}[1-F(y)]^{n-k}.$$Empirical Distribution Function (EDF)
$$F_n(x)=\tfrac{1}{n}\sum_{i=1}^n I(X_i\le x).$$ $E(F_n(x))=F(x),\,V(F_n(x))=F(x)(1-F(x))/n.$ By Glivenko–Cantelli: $\sup|F_n-F|\xrightarrow{a.s.}0.$Kolmogorov–Smirnov
$D_n=\sup_x|F_n(x)-F(x)|;\,\sqrt n D_n\xrightarrow{d}$ Kolmogorov distribution.🌍 Where it's used in real life
- Percentiles in growth charts and exam ranks.
- Extreme-value analysis of floods and earthquakes.
- Warranty design from minimum lifetimes.
- Median and quartiles in salary reports.
- Goodness-of-fit tests (Kolmogorov–Smirnov).
17. Rank Correlation: Spearman & Kendall
Spearman's $\rho_s$
Replace observations $(X_i,Y_i)$ by ranks $(R_i,S_i)$: $$\rho_s=1-\frac{6\sum d_i^2}{n(n^2-1)},\quad d_i=R_i-S_i.$$ $\rho_s\in[-1,1].$ Equals Pearson correlation between ranks.Kendall's $\tau$
Count concordant ($C$) and discordant ($D$) pairs: $$\tau=\frac{C-D}{\binom{n}{2}}=\frac{2(C-D)}{n(n-1)}.$$ For pair $(i,j)$: concordant if $\text{sgn}(X_i-X_j)=\text{sgn}(Y_i-Y_j)$.Properties
- Invariant under monotone transformations.
- Distribution free under $H_0$: $X,Y$ independent.
- Asymptotically: $\sqrt{n-1}\rho_s\xrightarrow{d}N(0,1)$, $\frac{3\tau\sqrt{n(n-1)}}{\sqrt{2(2n+5)}}\xrightarrow{d}N(0,1)$ under $H_0.$
🌍 Where it's used in real life
- Agreement between two judges' rankings.
- Correlating exam rank with interview rank.
- Customer preference vs price rank.
- Association in non-normal data (income vs health).
- Comparing sports ranking systems.