Skip to the content

Topics Covered

Exponential Family Lehmann–Scheffé Method Pooled Variance Distribution of Order Statistics Empirical Distribution Function (EDF) Kolmogorov–Smirnov Spearman's ρₛ Kendall's τ

Topic Overview — What & Why

Unit IV is about how to infer unknown population parameters from a sample. We need formal criteria to compare estimators (good vs bad) and a toolkit to construct optimal ones.

  • Unbiasedness: on average, the estimator hits the truth. A clean, intuitive criterion — but not the only one (mean-square-error trades off bias and variance).
  • Consistency: with enough data, the estimator converges to the true parameter. The minimum bar any estimator should clear.
  • Method of Moments & MLE: two recipes for constructing estimators. MoM is simple; MLE is asymptotically efficient.
  • Efficiency & UMVUE: among unbiased estimators, the one with smallest variance is "best". UMVUE is the gold standard.
  • Cramér-Rao Lower Bound: a fundamental limit — no unbiased estimator can have smaller variance than $1/(nI(\theta))$. Tells us when we have done as well as possible.
  • Sufficiency & factorisation theorem: a sufficient statistic captures all sample information about the parameter; you can compress data without losing inference power.
  • Minimal sufficiency & ancillarity: the most-compressed sufficient summary; ancillary statistics carry no parametric information.
  • Completeness: a technical condition that, combined with sufficiency, guarantees uniqueness of UMVUEs.
  • Rao-Blackwell theorem: conditioning any unbiased estimator on a sufficient statistic strictly improves it (in MSE). The "improvement" engine.
  • Lehmann-Scheffé theorem: if you have a complete sufficient statistic and any unbiased function of it, you have the UMVUE.
  • Basu's theorem: a complete sufficient statistic is independent of any ancillary statistic — surprisingly powerful for proving independence.
  • Pivoting & confidence intervals: turn estimators into ranges with stated coverage; one- and two-sample, normal and large-sample cases.
  • Order statistics & EDF: the building blocks of nonparametric inference; Glivenko-Cantelli says EDF converges uniformly to the true CDF.
  • Spearman & Kendall rank correlations: distribution-free measures of monotone association; useful when scales are ordinal or distributions are non-normal.

1. Unbiasedness

Why this section? Unbiasedness is often the first criterion students meet for evaluating an estimator. It says nothing about variability, so it must be combined with variance considerations — this is exactly what the rest of the unit develops.

An estimator $T(X_1,\ldots,X_n)$ of $g(\theta)$ is unbiased if $E_\theta[T]=g(\theta)$ for all $\theta\in\Theta.$ Bias = $E(T)-g(\theta).$

Mean Squared Error: $\text{MSE}(T)=E(T-g(\theta))^2=\text{Var}(T)+\text{Bias}^2(T).$

Bias–variance: why MSE = Variance + Bias² Low variance High variance Low bias High bias true value estimates
Bias vs variance. The green centre is the true $g(\theta)$; blue marks are an estimator's values across samples. Bias is how far their average sits from the centre; variance is their scatter. MSE combines both, so a slightly biased but tight estimator can beat an unbiased but noisy one.
EXAMPLE 1 $X_1,\ldots,X_n\sim N(\mu,\sigma^2).$ $\bar X$ is unbiased for $\mu$. $S^2=\frac{1}{n-1}\sum(X_i-\bar X)^2$ is unbiased for $\sigma^2$; $\hat\sigma^2_{ML}=\frac{1}{n}\sum(X_i-\bar X)^2$ is biased.
EXAMPLE 2 For Bernoulli$(p)$, $\bar X$ is unbiased for $p$; but $\bar X^2$ is NOT unbiased for $p^2$ (since $E\bar X^2=p^2+p(1-p)/n$).

🌍 Where it's used in real life

  1. Reporting an unbiased average income from a sample.
  2. Calibrating instruments to remove systematic error.
  3. Unbiased variance estimates in quality control.
  4. Fair estimates of vote share in polling.
  5. Estimating a defect rate without built-in bias.

2. Consistency

$T_n$ is consistent for $\theta$ if $T_n\xrightarrow{P}\theta$ as $n\to\infty.$
Sufficient condition: $E(T_n)\to\theta$ and $\text{Var}(T_n)\to 0.$
EXAMPLE 1 $\bar X_n$ is consistent for $\mu$ by WLLN: $E(\bar X)=\mu,\text{Var}(\bar X)=\sigma^2/n\to 0.$
EXAMPLE 2 For $X_i\sim U(0,\theta)$, $T=\max(X_i)$ is consistent: $P(T<\theta-\varepsilon)=(1-\varepsilon/\theta)^n\to 0.$

🌍 Where it's used in real life

  1. Trusting estimates to improve with bigger samples.
  2. Sensor readings settling down with more measurements.
  3. Poll accuracy rising with sample size.
  4. Model parameters stabilising as data grows.
  5. Long-run frequency estimating a true probability.

3. Method of Moments (MoM)

Equate sample moments to population moments and solve for parameters: $$m'_r=\tfrac{1}{n}\sum X_i^r=\mu'_r(\theta_1,\ldots,\theta_k),\quad r=1,\ldots,k.$$

EXAMPLE 1 $X\sim N(\mu,\sigma^2).$ MoM: $\hat\mu=\bar X$, $\hat\sigma^2=\tfrac{1}{n}\sum X_i^2-\bar X^2$.
EXAMPLE 2 $X\sim\text{Gamma}(\alpha,\beta)$ with mean $\alpha/\beta$, variance $\alpha/\beta^2.$ MoM: $\hat\alpha=\bar X^2/s^2,\hat\beta=\bar X/s^2.$

🌍 Where it's used in real life

  1. Fast first estimates when fitting a distribution.
  2. Fitting income data to a gamma model.
  3. Estimating rates from average counts.
  4. Starting values for more complex fitting.
  5. Calibrating input distributions for simulation.

4. Maximum Likelihood Estimation

Likelihood: $L(\theta)=\prod f(x_i;\theta).$ The MLE $\hat\theta$ maximizes $L$ (or $\log L$). Solve $\partial\log L/\partial\theta=0.$

Properties

EXAMPLE 1 Bernoulli$(p)$: $\log L=\sum X_i\log p+(n-\sum X_i)\log(1-p)$. $\hat p_{MLE}=\bar X.$
EXAMPLE 2 $N(\mu,\sigma^2)$: $\hat\mu=\bar X$, $\hat\sigma^2=\frac{1}{n}\sum(X_i-\bar X)^2$ (biased; corrected by $n/(n-1)$).

🌍 Where it's used in real life

  1. Fitting logistic models in medicine and credit scoring.
  2. Estimating failure rates in reliability.
  3. Training many machine-learning models.
  4. Estimating click and conversion probabilities.
  5. Estimating parameters in genetics.

5. Efficiency and UMVUE

Efficiency of unbiased $T$: $e(T)=\frac{1/I(\theta)}{V(T)}.$ $T$ is efficient if $e=1.$

UMVUE (Uniformly Minimum Variance Unbiased Estimator): the unbiased estimator with the smallest variance for every $\theta\in\Theta.$
EXAMPLE 1 $\bar X$ is UMVUE for $\mu$ in $N(\mu,\sigma^2)$ — both unbiased and attains CRLB.
EXAMPLE 2 For Poisson$(\lambda)$: $\bar X$ is UMVUE for $\lambda$. To estimate $e^{-\lambda}=P(X=0)$, the UMVUE is $(1-1/n)^{T}$ where $T=\sum X_i.$

🌍 Where it's used in real life

  1. Choosing the most precise estimator for given data.
  2. Minimising cost while hitting a target accuracy.
  3. Best unbiased estimate of a defect rate.
  4. Efficient survey estimators to save budget.
  5. Comparing estimators in simulation studies.

6. Cramér–Rao Lower Bound (CRLB)

Under regularity conditions, for any unbiased estimator $T$ of $g(\theta)$: $$V(T)\ge\frac{[g'(\theta)]^2}{nI(\theta)},$$ where Fisher information $I(\theta)=E\left[\left(\frac{\partial\log f}{\partial\theta}\right)^2\right]=-E\left[\frac{\partial^2\log f}{\partial\theta^2}\right].$

Intuition. Fisher information $I(\theta)$ measures the average curvature (sharpness) of the log-likelihood in $\theta$: the more sharply the likelihood peaks around the true value, the more the data “pin down” the parameter, and the smaller the variance any unbiased estimator can achieve. The CRLB is therefore the precision ceiling imposed by the model itself, and an estimator meeting it is doing as well as the information in the data allows.

Equality holds iff there exists an exponential family structure with $T$ as canonical statistic.

EXAMPLE 1 Bernoulli$(p)$: $I(p)=\frac{1}{p(1-p)}.$ CRLB $=p(1-p)/n.$ $V(\bar X)=p(1-p)/n$ — attains bound.
EXAMPLE 2 Normal $\sigma^2$ known, $\mu$ unknown: $I(\mu)=1/\sigma^2,$ CRLB $=\sigma^2/n=V(\bar X).$

🌍 Where it's used in real life

  1. The best possible accuracy of GPS positioning.
  2. Limits of radar and sonar estimation.
  3. Designing sensors to a precision target.
  4. Planning sample size for a required precision.
  5. Benchmarking estimator quality in signal processing.

7. Sufficiency & Factorization Theorem

$T(X)$ is sufficient for $\theta$ if conditional distribution of $X$ given $T$ is independent of $\theta.$
Neyman–Fisher Factorization: $T$ is sufficient iff $$f(x;\theta)=g(T(x),\theta)\,h(x)$$ for some functions $g,h.$

Intuition. A sufficient statistic compresses all the parameter-relevant information in the sample into a smaller summary: once $T$ is known, the leftover randomness in the raw data carries no further information about $\theta$. The factorization theorem lets you certify sufficiency just by inspecting how $\theta$ enters the likelihood — if $\theta$ “touches” the data only through $T(x)$, then $T$ is sufficient.

Exponential Family

$f(x;\theta)=\exp\{\eta(\theta)T(x)-A(\theta)\}h(x)$ — $T$ is sufficient and complete.
EXAMPLE 1 Bernoulli: $f=p^x(1-p)^{1-x}.$ Joint $=p^{\sum x_i}(1-p)^{n-\sum x_i}.$ So $T=\sum X_i$ is sufficient.
EXAMPLE 2 $N(\mu,\sigma^2)$: joint factors with $T=(\sum X_i,\sum X_i^2)$ sufficient — equivalently $(\bar X,S^2).$

🌍 Where it's used in real life

  1. Summarising data by a few numbers with no loss.
  2. Storing totals instead of full datasets.
  3. Data compression for inference.
  4. Reporting the sample sum or mean in QC.
  5. Streaming statistics that keep running summaries.

8. Minimal Sufficiency & Ancillarity

$T$ is minimal sufficient if it is a function of every other sufficient statistic.
$A$ is ancillary if its distribution does not depend on $\theta$.

Lehmann–Scheffé Method

$T(X)$ is minimal sufficient iff: $f(x;\theta)/f(y;\theta)$ is free of $\theta$ $\iff$ $T(x)=T(y).$
EXAMPLE 1 For $N(\mu,1)$: $\bar X$ is minimal sufficient. $S^2$ is ancillary.
EXAMPLE 2 For $U(\theta,\theta+1)$: $(X_{(1)},X_{(n)})$ is minimal sufficient; range $X_{(n)}-X_{(1)}$ is ancillary.

🌍 Where it's used in real life

  1. Finding the smallest summary inference needs.
  2. Efficient data reduction in big-data pipelines.
  3. Spotting which extra data adds nothing.
  4. Designing compact monitoring statistics.
  5. Simplifying models to their essentials.

9. Completeness

A family $\{f(x;\theta):\theta\in\Theta\}$ is complete if $E_\theta[g(T)]=0\,\forall\theta\Rightarrow g(T)=0$ a.s. for all $\theta.$

Intuition. Completeness says the family is “rich enough” that the only unbiased estimator of $0$ built from $T$ is the trivial one ($g(T)\equiv 0$). This rules out two different unbiased functions of $T$ estimating the same quantity, which is precisely why a statistic that is both complete and sufficient delivers a unique UMVUE (Lehmann–Scheffé).

Exponential family is complete (under usual rank conditions on the natural parameter space).

EXAMPLE 1 $\sum X_i$ is complete sufficient for Bernoulli$(p)$, Poisson$(\lambda)$, Exp$(\lambda).$
EXAMPLE 2 For $U(0,\theta)$: $X_{(n)}$ is complete sufficient — $E[g(X_{(n)})]=0$ implies $g\equiv 0.$

🌍 Where it's used in real life

  1. Guaranteeing a unique best unbiased estimator.
  2. Avoiding ambiguous estimates in surveys.
  3. Theory behind reliable quality estimates.
  4. Ensuring a model is identifiable.
  5. Foundation for building UMVUEs.

10. Rao–Blackwell Theorem

Let $T$ be sufficient for $\theta$, $W$ unbiased for $g(\theta).$ Define $T^*=E(W\mid T).$ Then:
  1. $T^*$ is a statistic (no $\theta$ dependence by sufficiency).
  2. $E(T^*)=g(\theta).$
  3. $V(T^*)\le V(W),$ with equality iff $W=T^*$ a.s.
EXAMPLE 1 Bernoulli$(p)$, want UMVUE of $p^2.$ $X_1 X_2$ is unbiased; $T=\sum X_i$ sufficient. $E(X_1X_2\mid T=t)=\frac{t(t-1)}{n(n-1)}.$ Improved estimator.
EXAMPLE 2 $X_i\sim$ Poisson$(\lambda),$ unbiased estimator of $e^{-\lambda}$: $W=I(X_1=0).$ $T=\sum X_i$ sufficient. Since $X_1\mid T=t\sim$ Bin$(t,1/n)$, $E(W\mid T=t)=P(X_1=0\mid T=t)=\left(\frac{n-1}{n}\right)^t=(1-1/n)^t.$

🌍 Where it's used in real life

  1. Turning a rough estimator into a better one.
  2. Variance reduction in Monte-Carlo simulation.
  3. Sharpening survey estimates.
  4. Refining predicted probabilities in ML.
  5. Better reliability estimates from summaries.

11. Lehmann–Scheffé Theorem

If $T$ is complete sufficient for $\theta$ and $W=h(T)$ is unbiased for $g(\theta)$, then $W$ is the unique UMVUE.

Two equivalent paths to UMVUE:

EXAMPLE 1 $U(0,\theta)$, $X_{(n)}$ complete sufficient, $E(X_{(n)})=\frac{n}{n+1}\theta.$ UMVUE of $\theta$ is $\frac{n+1}{n}X_{(n)}.$
EXAMPLE 2 $X_i\sim$ Poisson$(\lambda),$ $T=\sum X_i$ complete sufficient. UMVUE of $\lambda$ is $\bar X=T/n.$

🌍 Where it's used in real life

  1. Constructing the best unbiased estimator in practice.
  2. Best estimate of a probability like P(no defect).
  3. Optimal survey estimators.
  4. Reliability-metric estimation.
  5. A yardstick to judge other estimators.

12. Basu's Theorem

If $T$ is complete sufficient and $A$ is ancillary, then $T$ and $A$ are independent.
EXAMPLE 1 For $N(\mu,1)$: $\bar X$ is complete sufficient; $S^2$ is ancillary $\Rightarrow\bar X$ and $S^2$ independent.
EXAMPLE 2 For Exp$(\theta)$: $\sum X_i$ complete sufficient; $X_{(1)}/\sum X_i$ has distribution free of $\theta$ → independent of $\sum X_i.$

🌍 Where it's used in real life

  1. Proving sample mean and variance are independent (normal).
  2. Simplifying derivations of sampling distributions.
  3. Justifying separate study of location and spread.
  4. Underlying theory of the t-test.
  5. Independence arguments in probability proofs.

13. Method of Pivoting

A pivot $Q(X,\theta)$ has a distribution that does not depend on $\theta.$ Choose $a,b$ s.t. $P(a\le Q\le b)=1-\alpha$, then invert to obtain CI.
EXAMPLE 1 $X_i\sim N(\mu,\sigma^2)$ ($\sigma$ known): $Q=(\bar X-\mu)/(\sigma/\sqrt n)\sim N(0,1).$ $1-\alpha$ CI: $\bar X\pm z_{\alpha/2}\sigma/\sqrt n.$
EXAMPLE 2 $\sigma$ unknown: $Q=(\bar X-\mu)/(S/\sqrt n)\sim t_{n-1}.$ CI: $\bar X\pm t_{n-1,\alpha/2}S/\sqrt n.$

🌍 Where it's used in real life

  1. Building a confidence interval for a mean.
  2. Margin of error in a poll.
  3. Tolerance intervals in manufacturing.
  4. Confidence bounds on a failure rate.
  5. Interval estimates for lab measurements.

14. Confidence Intervals — One and Two Sample

A confidence interval for a parameter $\theta$ is a random interval $(L,U)$ computed from the sample such that $P(L\le\theta\le U)=1-\alpha$. The value $1-\alpha$ is the confidence coefficient (e.g. 95%): over many repeated samples, that proportion of the intervals will contain the true $\theta$.
ParameterConditions$1-\alpha$ CI
$\mu$$\sigma$ known$\bar X\pm z_{\alpha/2}\sigma/\sqrt n$
$\mu$$\sigma$ unknown$\bar X\pm t_{n-1,\alpha/2}S/\sqrt n$
$\sigma^2$—$\left(\frac{(n-1)S^2}{\chi^2_{n-1,\alpha/2}},\frac{(n-1)S^2}{\chi^2_{n-1,1-\alpha/2}}\right)$
$\mu_1-\mu_2$$\sigma_1=\sigma_2$ unknown$\bar X-\bar Y\pm t_{n_1+n_2-2,\alpha/2}S_p\sqrt{\tfrac{1}{n_1}+\tfrac{1}{n_2}}$
$\sigma_1^2/\sigma_2^2$—$\left(\frac{S_1^2/S_2^2}{F_{n_1-1,n_2-1,\alpha/2}},\frac{S_1^2/S_2^2}{F_{n_1-1,n_2-1,1-\alpha/2}}\right)$

Pooled Variance

$S_p^2=\frac{(n_1-1)S_1^2+(n_2-1)S_2^2}{n_1+n_2-2}.$
EXAMPLE 1 $n=25,\bar X=10,S=2.$ 95% CI for $\mu$: $10\pm 2.064(2/5)=10\pm 0.826=(9.17,10.83).$
EXAMPLE 2 $n_1=10,n_2=12,\bar X=15,\bar Y=12,S_p=3.$ 95% CI for $\mu_1-\mu_2$: $3\pm 2.086(3)\sqrt{1/10+1/12}=3\pm 2.69.$

🌍 Where it's used in real life

  1. The "±3% margin of error" in election polls.
  2. A drug's likely effect range in a trial.
  3. Confidence range for average delivery time.
  4. Control limits for a process mean.
  5. Confidence interval for a conversion rate.

15. Large-Sample Confidence Intervals

Based on CLT and consistency: if $\hat\theta$ is asymptotically normal, $$\hat\theta\pm z_{\alpha/2}\,\text{SE}(\hat\theta).$$

Common Cases

EXAMPLE 1 Survey: 60 successes in 100 trials. $\hat p=0.6.$ 95% CI: $0.6\pm 1.96\sqrt{0.24/100}=0.6\pm 0.096=(0.504,0.696).$
EXAMPLE 2 For Poisson, observe $X=49$ events in 1 unit. 95% CI: $49\pm 1.96\sqrt{49}=49\pm 13.72=(35.28,62.72).$

🌍 Where it's used in real life

  1. Big-survey approve/disapprove proportions.
  2. Conversion-rate intervals in web analytics.
  3. Incidence-rate intervals in epidemiology.
  4. Event-rate confidence (accidents per month).
  5. Large-sample intervals in A/B testing.

16. Order Statistics & Empirical Distribution Function

The order statistics of a sample $X_1,\ldots,X_n$ are the observations re-arranged in increasing order, $X_{(1)}\le X_{(2)}\le\cdots\le X_{(n)}$, where $X_{(1)}=\min$ and $X_{(n)}=\max$. They are the building blocks of the sample median, range, quantiles and many non-parametric methods.

Distribution of Order Statistics

$X_{(1)}\le X_{(2)}\le\cdots\le X_{(n)}.$ $$f_{X_{(k)}}(x)=\frac{n!}{(k-1)!(n-k)!}f(x)F(x)^{k-1}(1-F(x))^{n-k}.$$ Joint of $(X_{(j)},X_{(k)})$, $j<k$: $$f_{j,k}(x,y)=\frac{n!}{(j-1)!(k-j-1)!(n-k)!}f(x)f(y)F(x)^{j-1}[F(y)-F(x)]^{k-j-1}[1-F(y)]^{n-k}.$$

Empirical Distribution Function (EDF)

$$F_n(x)=\tfrac{1}{n}\sum_{i=1}^n I(X_i\le x).$$ $E(F_n(x))=F(x),\,V(F_n(x))=F(x)(1-F(x))/n.$ By Glivenko–Cantelli: $\sup|F_n-F|\xrightarrow{a.s.}0.$

Kolmogorov–Smirnov

$D_n=\sup_x|F_n(x)-F(x)|;\,\sqrt n D_n\xrightarrow{d}$ Kolmogorov distribution.
EXAMPLE 1 $X_i\sim U(0,1),\,n=5.$ $f_{X_{(3)}}(x)=\frac{5!}{2!2!}x^2(1-x)^2=30x^2(1-x)^2,\,0\le x\le 1.$
EXAMPLE 2 For Exp$(1),\,n=2.$ $X_{(1)}=\min(X_1,X_2)\sim$ Exp$(2).$ $X_{(2)}-X_{(1)}\sim$ Exp$(1)$ independent of $X_{(1)}.$

🌍 Where it's used in real life

  1. Percentiles in growth charts and exam ranks.
  2. Extreme-value analysis of floods and earthquakes.
  3. Warranty design from minimum lifetimes.
  4. Median and quartiles in salary reports.
  5. Goodness-of-fit tests (Kolmogorov–Smirnov).

17. Rank Correlation: Spearman & Kendall

A rank correlation measures the strength and direction of a monotonic relationship between two variables using only the ranks of the observations, not their exact values. Being rank-based, it is distribution-free and resistant to outliers.

Spearman's $\rho_s$

Replace observations $(X_i,Y_i)$ by ranks $(R_i,S_i)$: $$\rho_s=1-\frac{6\sum d_i^2}{n(n^2-1)},\quad d_i=R_i-S_i.$$ $\rho_s\in[-1,1].$ Equals Pearson correlation between ranks.

Kendall's $\tau$

Count concordant ($C$) and discordant ($D$) pairs: $$\tau=\frac{C-D}{\binom{n}{2}}=\frac{2(C-D)}{n(n-1)}.$$ For pair $(i,j)$: concordant if $\text{sgn}(X_i-X_j)=\text{sgn}(Y_i-Y_j)$.

Properties

EXAMPLE 1 Ranks $(1,2,3,4,5)$ vs $(2,1,4,3,5)$: $d=( -1,1,-1,1,0),\,\sum d^2=4.$ $\rho_s=1-\frac{24}{120}=0.8.$
EXAMPLE 2 Same data: pairs (1,2),(1,3),(1,4),(1,5),(2,3),(2,4),(2,5),(3,4),(3,5),(4,5). Concordant=8, Discordant=2. $\tau=(8-2)/10=0.6.$

🌍 Where it's used in real life

  1. Agreement between two judges' rankings.
  2. Correlating exam rank with interview rank.
  3. Customer preference vs price rank.
  4. Association in non-normal data (income vs health).
  5. Comparing sports ranking systems.