Skip to the content

Topics Covered

Estimator vs Estimate Unbiasedness Consistency Efficiency Sufficiency Method of Moments MLE Rao-Cramer Factorization Theorem Regularity Conditions Fisher Information Confidence Intervals
On this page
  1. 1. Estimator and Estimate
  2. 2. Criteria of a Good Estimator
  3. 3. Methods of Estimation
  4. 4. Rao–Cramer Inequality
  5. 5. Confidence Intervals (CIs)
  6. Key Take-aways

1. Estimator and Estimate

DEFINITIONS

Two kinds of estimation:

1.1 Parameter Space (\(\Theta\))

The parameter space \(\Theta\) is the set of all admissible values a parameter can take. Estimation is the problem of using the sample to pick the value of \(\theta\in\Theta\) that generated the data.

Example. If \(X \sim N(\mu,\sigma^2)\) with density \(f(x;\mu,\sigma^2)\), then

\[ \Theta = \{(\mu,\sigma^2)\;:\; -\infty < \mu < \infty,\ \sigma > 0\}. \]

An estimator is a map from the sample space into \(\Theta\); "best estimator" means the point of \(\Theta\) that best satisfies the criteria U–C–E–S below.

2. Criteria of a Good Estimator

2.1 Unbiasedness

\(\hat\theta\) is unbiased for \(\theta\) if \(E(\hat\theta) = \theta\) for all \(\theta\).

EXAMPLE 1

Sample mean \(\bar X = \frac{1}{n}\sum X_i\) is unbiased for \(\mu\): \(E(\bar X) = \mu\).

EXAMPLE 2

Sample variance \(s^2 = \frac{1}{n-1}\sum(X_i - \bar X)^2\) is unbiased for \(\sigma^2\). The naïve version with \(1/n\) is biased: \(E(\sigma_n^2) = \frac{n-1}{n}\sigma^2\).

DERIVATION — why we divide by \(n-1\)

Let \(s^2 = \frac1n\sum(X_i-\bar X)^2 = \frac1n\sum X_i^2 - \bar X^2\). Using \(E(X_i^2)=\sigma^2+\mu^2\) and \(E(\bar X^2)=\dfrac{\sigma^2}{n}+\mu^2\),

\[ E(s^2) = \frac1n\sum E(X_i^2) - E(\bar X^2) = (\sigma^2+\mu^2) - \Big(\tfrac{\sigma^2}{n}+\mu^2\Big) = \Big(1-\tfrac1n\Big)\sigma^2 = \frac{n-1}{n}\,\sigma^2 . \]

So \(s^2\) is biased (it underestimates \(\sigma^2\)). Rescaling removes the bias:

\[ E\!\left[\frac{n}{n-1}\,s^2\right] = \sigma^2 \quad\Longrightarrow\quad S^2 = \frac{1}{n-1}\sum_{i=1}^n (X_i-\bar X)^2 \ \text{ is unbiased for } \sigma^2 . \]
PROBLEM 1 — an unbiased estimator squared is generally biased

If \(T\) is unbiased for \(\theta\), show that \(T^2\) is not unbiased for \(\theta^2\).

Solution. \(E(T)=\theta\). From \(\operatorname{Var}(T)=E(T^2)-[E(T)]^2\),

\[ E(T^2) = \theta^2 + \operatorname{Var}(T). \]

Since \(\operatorname{Var}(T)>0\) (unless \(T\) is degenerate), \(E(T^2)\neq\theta^2\); \(T^2\) overestimates \(\theta^2\) by exactly \(\operatorname{Var}(T)\).

PROBLEM 2 — an unbiased estimator of \(\theta^2\) (Bernoulli)

Let \(X_1,\dots,X_n\) take values \(1\) and \(0\) with probabilities \(\theta\) and \(1-\theta\). Show that \(\dfrac{\sum X_i\left(\sum X_i-1\right)}{n(n-1)}\) is unbiased for \(\theta^2\).

Solution. \(T=\sum X_i\sim\text{Bin}(n,\theta)\), so \(E(T)=n\theta\), \(\operatorname{Var}(T)=n\theta(1-\theta)\). Then

\[ E\!\left[\frac{T(T-1)}{n(n-1)}\right] = \frac{E(T^2)-E(T)}{n(n-1)} = \frac{\big[\operatorname{Var}(T)+(E T)^2\big]-E(T)}{n(n-1)}. \] \[ = \frac{n\theta(1-\theta)+n^2\theta^2-n\theta}{n(n-1)} = \frac{n^2\theta^2-n\theta^2}{n(n-1)} = \frac{n\theta^2(n-1)}{n(n-1)} = \theta^2 . \]

Hence the statistic is unbiased for \(\theta^2\). (Verified numerically: for \(n=10,\ \theta=0.62\) the expectation equals \(0.3844=\theta^2\).)

PROBLEM 3 — exponential mean (unbiased, variance \(\theta^2/n\))

For a sample of size \(n\) from \(f(x)=\dfrac{e^{-x/\theta}}{\theta},\ 0<x<\infty\), show that \(\bar X\) is unbiased for \(\theta\) with variance \(\theta^2/n\).

Solution. \(E(X)=\displaystyle\int_0^\infty x\,\frac{e^{-x/\theta}}{\theta}\,dx=\theta\) and \(E(X^2)=\displaystyle\int_0^\infty x^2\,\frac{e^{-x/\theta}}{\theta}\,dx=2\theta^2\), so \(\operatorname{Var}(X)=2\theta^2-\theta^2=\theta^2\). Therefore

\[ E(\bar X)=\frac1n\sum E(X_i)=\theta, \qquad \operatorname{Var}(\bar X)=\frac1{n^2}\sum\operatorname{Var}(X_i)=\frac{\theta^2}{n}. \]

So \(\bar X\) is an unbiased, consistent estimator of \(\theta\).

2.2 Consistency

\(\hat\theta_n\) is consistent if \(\hat\theta_n \xrightarrow{P} \theta\) as \(n \to \infty\). By Chebyshev: an unbiased estimator is consistent if \(\text{Var}(\hat\theta_n) \to 0\).

EXAMPLE 1

\(\bar X\) is consistent for \(\mu\): unbiased and \(\text{Var}(\bar X) = \sigma^2/n \to 0\).

EXAMPLE 2

The sample proportion \(\hat p = X/n\) (where \(X\) is the count of successes in Bernoulli trials) is consistent for \(p\) with variance \(p(1-p)/n \to 0\).

THEOREM 1 — INVARIANCE PROPERTY OF CONSISTENT ESTIMATORS

Statement. If \(t_n\) is a consistent estimator of \(\theta\) and \(\psi\) is a function continuous at \(\theta\), then \(\psi(t_n)\) is a consistent estimator of \(\psi(\theta)\).

Proof. Consistency means \(t_n \xrightarrow{P}\theta\): for every \(\varepsilon>0,\ \eta>0\) there is \(m\) such that \(P(|t_n-\theta|\le\varepsilon) > 1-\eta\) for all \(n\ge m\). Because \(\psi\) is continuous at \(\theta\), for every \(\varepsilon_1>0\) there is \(\varepsilon>0\) with

\[ |t_n-\theta|\le\varepsilon \;\Longrightarrow\; |\psi(t_n)-\psi(\theta)|\le\varepsilon_1 . \]

Hence the event \(\{|t_n-\theta|\le\varepsilon\}\) is contained in \(\{|\psi(t_n)-\psi(\theta)|\le\varepsilon_1\}\), so

\[ P\big(|\psi(t_n)-\psi(\theta)|\le\varepsilon_1\big) \;\ge\; P\big(|t_n-\theta|\le\varepsilon\big) \;>\; 1-\eta \quad (n\ge m). \]

Therefore \(\psi(t_n)\xrightarrow{P}\psi(\theta)\), i.e. \(\psi(t_n)\) is consistent for \(\psi(\theta)\). \(\blacksquare\)

THEOREM 2 — SUFFICIENT CONDITION FOR CONSISTENCY

Statement. If \(E(t_n)\to\theta\) and \(\operatorname{Var}(t_n)\to 0\) as \(n\to\infty\), then \(t_n\) is a consistent estimator of \(\theta\).

Proof. By Chebyshev's inequality applied to \(t_n\), for any \(\varepsilon>0\)

\[ P\big(|t_n-\theta|\ge\varepsilon\big) \;\le\; \frac{E\big[(t_n-\theta)^2\big]}{\varepsilon^2} \;=\; \frac{\operatorname{Var}(t_n) + \big(E(t_n)-\theta\big)^2}{\varepsilon^2}, \]

since the mean-squared error splits as variance plus squared bias. As \(n\to\infty\) the bias \(E(t_n)-\theta\to0\) and \(\operatorname{Var}(t_n)\to0\), so the right-hand side \(\to0\). Hence \(P(|t_n-\theta|\ge\varepsilon)\to0\), i.e. \(t_n\xrightarrow{P}\theta\). \(\blacksquare\)

DERIVATION 1 — \(\bar X\) is unbiased and consistent for \(\mu\)

For \(X_1,\dots,X_n\) with \(E(X_i)=\mu,\ \operatorname{Var}(X_i)=\sigma^2\),

\[ E(\bar X) = \frac1n\sum_{i=1}^n E(X_i) = \frac1n\,(n\mu) = \mu \quad\Rightarrow\quad \text{unbiased.} \] \[ \operatorname{Var}(\bar X) = \frac1{n^2}\sum_{i=1}^n \operatorname{Var}(X_i) = \frac{\sigma^2}{n} \xrightarrow[n\to\infty]{} 0. \]

Unbiased with variance \(\to0\); by Theorem 2, \(\bar X\) is consistent for \(\mu\).

DERIVATION 2 — \(s^2\) is consistent for \(\sigma^2\) (normal population)

For a normal sample, \(\dfrac{n s^2}{\sigma^2}\sim\chi^2_{(n-1)}\), whose mean is \(n-1\) and variance \(2(n-1)\). Hence

\[ E(s^2) = \frac{n-1}{n}\,\sigma^2 \to \sigma^2, \qquad \operatorname{Var}(s^2) = \frac{2(n-1)}{n^2}\,\sigma^4 \to 0 . \]

Mean \(\to\sigma^2\) and variance \(\to0\); by Theorem 2, \(s^2\) is consistent for \(\sigma^2\) (the unbiased \(S^2\) is consistent too).

2.3 Efficiency

Among unbiased estimators, the one with minimum variance is most efficient. Relative efficiency:

\[ \text{RE}(\hat\theta_1, \hat\theta_2) = \dfrac{\text{Var}(\hat\theta_2)}{\text{Var}(\hat\theta_1)}. \]

If RE > 1, \(\hat\theta_1\) is more efficient than \(\hat\theta_2\).

EXAMPLE 1

For a normal population, both \(\bar X\) and the sample median are unbiased for \(\mu\). \(\text{Var}(\bar X) = \sigma^2/n\), \(\text{Var}(\text{median}) \approx \pi\sigma^2/(2n) \approx 1.57\,\sigma^2/n\). Hence \(\bar X\) is about 57 % more efficient.

EXAMPLE 2

If \(T_1\) has variance 5 and \(T_2\) has variance 8 (both unbiased), \(\text{RE}(T_1,T_2) = 8/5 = 1.6\), so \(T_1\) is 60 % more efficient.

COEFFICIENT OF EFFICIENCY

Among consistent estimators, the one with the smallest variance is the most efficient. If \(t_1\) is the most efficient estimator (variance \(V_1\)) and \(t_2\) is another (variance \(V_2\)), the efficiency of \(t_2\) is

\[ e = \frac{V_1}{V_2}, \qquad 0 < e \le 1 . \]

Example. For a normal population both \(\bar X\) and the sample median \(M_d\) are unbiased and consistent for \(\mu\), with \(\operatorname{Var}(\bar X)=\sigma^2/n\) and \(\operatorname{Var}(M_d)=\dfrac{\pi\sigma^2}{2n}\approx 1.571\,\dfrac{\sigma^2}{n}\). Since \(\operatorname{Var}(\bar X)<\operatorname{Var}(M_d)\), the mean is most efficient, and the efficiency of the median is

\[ e = \frac{\operatorname{Var}(\bar X)}{\operatorname{Var}(M_d)} = \frac{\sigma^2/n}{\pi\sigma^2/2n} = \frac{2}{\pi} \approx 0.637 . \]
PROBLEM — comparing estimators of \(\mu\) (find \(\lambda\); find the best)

A random sample of size 5 is drawn from a population with mean \(\mu\), variance \(\sigma^2\). Consider

\[ t_1=\frac{X_1+X_2+X_3+X_4+X_5}{5},\quad t_2=\frac{X_1+X_2}{2}+X_3,\quad t_3=\frac{2X_1+X_2+\lambda X_3}{3}. \]

(a) Value of \(\lambda\) making \(t_3\) unbiased. \(E(t_3)=\dfrac{2\mu+\mu+\lambda\mu}{3}=\dfrac{(3+\lambda)\mu}{3}\). Setting this \(=\mu\) gives \(3+\lambda=3\), so \(\boxed{\lambda=0}\).

(b) Are \(t_1,t_2\) unbiased? \(E(t_1)=\mu\) — unbiased. \(E(t_2)=\dfrac{\mu+\mu}{2}+\mu=2\mu\neq\mu\) — \(t_2\) is biased.

(c) Best estimator. Compare variances of the unbiased ones \(t_1\) and \(t_3\) (with \(\lambda=0\)):

\[ \operatorname{Var}(t_1)=\frac{5\sigma^2}{25}=\frac{\sigma^2}{5}=0.20\,\sigma^2,\qquad \operatorname{Var}(t_3)=\frac{(2^2+1^2)\sigma^2}{9}=\frac{5\sigma^2}{9}\approx0.556\,\sigma^2. \]

Since \(\operatorname{Var}(t_1)<\operatorname{Var}(t_3)\), the sample mean \(t_1\) is the best (most efficient) estimator of \(\mu\).

2.4 Sufficiency

A statistic \(T\) is sufficient for \(\theta\) if the conditional distribution of the sample given \(T\) does not depend on \(\theta\). It captures all the information in the data about \(\theta\).

FISHER–NEYMAN FACTORISATION THEOREM

\(T(\mathbf X)\) is sufficient for \(\theta\) iff the joint pdf factors as

\[ f(\mathbf x; \theta) \;=\; g(T(\mathbf x); \theta) \cdot h(\mathbf x), \]

where \(h\) does not depend on \(\theta\).

EXAMPLE 1

For Bernoulli\((p)\) sample, \(\sum X_i\) is sufficient for \(p\). Joint pmf factors as \(p^{\sum x_i}(1-p)^{n-\sum x_i}\).

EXAMPLE 2

For \(N(\mu, \sigma^2)\) with known \(\sigma^2\), \(\sum X_i\) (or \(\bar X\)) is sufficient for \(\mu\).

EXAMPLE 3 (Poisson)

For \(X_1,\dots,X_n \sim\) Poisson\((\lambda)\), \(L(\lambda) = e^{-n\lambda}\lambda^{\sum x_i}/\prod x_i!\). Here \(g = e^{-n\lambda}\lambda^{\sum x_i}\) depends on the data only through \(T = \sum X_i\), and \(h = 1/\prod x_i!\) is free of \(\lambda\); hence \(\sum X_i\) (equivalently \(\bar X\)) is sufficient for \(\lambda\).

PROBLEM 1 — sufficient statistic for the exponential \(\lambda\)

Let \(x_1,x_2,\dots,x_n\) be a random sample from \(f(x,\lambda)=\lambda e^{-\lambda x},\ x>0,\ \lambda>0\). Find a sufficient statistic for \(\lambda\).

Solution. The likelihood function of \(x_1,x_2,\dots,x_n\) is

\[ L = f(x_1,\lambda)\,f(x_2,\lambda)\cdots f(x_n,\lambda) = \lambda e^{-\lambda x_1}\cdot\lambda e^{-\lambda x_2}\cdots\lambda e^{-\lambda x_n} = \lambda^{n} e^{-\lambda\sum_{i=1}^{n} x_i}. \] \[ = \underbrace{e^{-\lambda\sum_{i=1}^{n} x_i}\cdot\lambda^{n}}_{g_\lambda[t(x)]}\cdot \underbrace{1}_{h(x)} . \]

Here \(g_\lambda[t(x)] = e^{-\lambda\sum x_i}\lambda^{n}\) is a function of \(\lambda\) and of the data only through \(t(x)=\sum_{i=1}^{n} x_i\), and \(h(x)=1\) is independent of \(\lambda\).

Therefore, by the factorization theorem, \(\displaystyle\sum_{i=1}^{n} x_i\) is a sufficient statistic for the population parameter \(\lambda\) of the exponential distribution.

PROBLEM 2 — sufficient statistics for \(\mu\) and \(\sigma^2\) in a normal population

Let \(x_1,x_2,\dots,x_n\) be a random sample of size \(n\) drawn from a normal population with parameters \(\mu\) and \(\sigma^2\), so that

\[ f(x_i,\mu,\sigma^2) = \frac{1}{\sigma\sqrt{2\pi}}\, e^{-\frac12\left(\frac{x_i-\mu}{\sigma}\right)^2}, \qquad -\infty<x_i<\infty,\ -\infty<\mu<\infty,\ \sigma>0 . \]

Solution. The likelihood function is

\[ L = f(x_1;\mu,\sigma^2)\,f(x_2;\mu,\sigma^2)\cdots f(x_n;\mu,\sigma^2) = \frac{1}{(\sigma\sqrt{2\pi})^{n}}\, e^{-\frac{1}{2\sigma^{2}}\sum_{i=1}^{n}(x_i-\mu)^2}. \]

Expanding the square, \(\sum (x_i-\mu)^2 = \sum x_i^2 - 2\mu\sum x_i + n\mu^2\), so

\[ L = \frac{1}{(\sigma\sqrt{2\pi})^{n}}\, e^{-\frac{1}{2\sigma^{2}}\left[\sum_{i=1}^{n} x_i^{2}-2\mu\sum_{i=1}^{n} x_i+n\mu^{2}\right]} \;=\; g_{(\mu,\sigma^2)}[t_1(x),t_2(x)]\cdot h(x), \]

where \(g_{(\mu,\sigma^2)}\) is the whole expression above — a function of \(\mu,\sigma^2\) and of the data only through

\[ t_1(x)=\sum_{i=1}^{n} x_i, \qquad t_2(x)=\sum_{i=1}^{n} x_i^{2}, \]

and \(h(x)=1\) is independent of \(\mu\) and \(\sigma^2\).

Therefore, by the factorization theorem, \(\sum x_i\) is a sufficient statistic for \(\mu\) and \(\sum x_i^{2}\) is a sufficient statistic for \(\sigma^2\).

PROBLEM 3 — a case where the sufficient statistic is a product

Find a sufficient statistic for \(\theta\) for the density \(f(x,\theta)=\theta x^{\theta-1},\ 0<x<1,\ \theta>0\).

Solution. The likelihood function is

\[ L = f(x_1,\theta)\,f(x_2,\theta)\cdots f(x_n,\theta) = \theta x_1^{\theta-1}\cdot\theta x_2^{\theta-1}\cdots\theta x_n^{\theta-1} = \theta^{n}\,(x_1x_2\cdots x_n)^{\theta-1}. \] \[ = \frac{\theta^{n}(x_1x_2\cdots x_n)^{\theta}}{x_1x_2\cdots x_n} = \theta^{n}\left(\prod_{i=1}^{n} x_i\right)^{\!\theta}\cdot \frac{1}{\prod_{i=1}^{n} x_i} \;=\; g_\theta[t(x)]\cdot h(x), \]

where \(g_\theta[t(x)] = \theta^{n}\left(\prod x_i\right)^{\theta}\) is a function of \(\theta\) and of the data only through \(t(x)=\prod_{i=1}^{n} x_i\), and \(h(x)=1/\prod x_i\) depends on the data alone and is independent of \(\theta\).

Therefore, by the factorization theorem, \(\displaystyle\prod_{i=1}^{n} x_i\) is a sufficient statistic for \(\theta\). Note that a sufficient statistic need not be a sum — here it is a product.

2.5 Minimum-Variance Unbiased Estimator (MVUE)

Among all unbiased estimators of \(\theta\), the one with the smallest variance is the MVUE. Two theorems locate it:

Example. For Poisson\((\lambda)\), \(\bar X\) is an unbiased function of the complete sufficient statistic \(\sum X_i\); by Lehmann–Scheffé it is the MVUE of \(\lambda\).

3. Methods of Estimation

THE SIX METHODS

There are several methods of obtaining a good — or the best — estimator for a population parameter:

  1. Method of maximum likelihood estimation
  2. Method of moments
  3. Method of least squares
  4. Method of minimum variance
  5. Method of minimum chi-square
  6. Method of inverse probability

The first two are the ones this unit develops, and the two the syllabus examines. The third, the principle of least squares, is already covered under curve fitting.

3.1 Method of Moments (MoM)

Introduced by Prof. Karl Pearson. Let \(x_1,x_2,\dots,x_n\) be a random sample drawn from a population with density \(f(x;\theta_1,\theta_2,\dots,\theta_k)\), where \(\theta_1,\dots,\theta_k\) are \(k\) parameters. The first \(k\) moments of the population about the origin are

\[ \mu'_r = E(X^r) = \int_{-\infty}^{\infty} x^r f(x;\theta_1,\theta_2,\dots,\theta_k)\,dx, \qquad r=1,2,\dots,k, \]

and each \(\mu'_r\) is in turn a function of \(\theta_1,\theta_2,\dots,\theta_k\). The corresponding moments of the sample are

\[ a_r = \frac{1}{n}\sum_{i=1}^{n} x_i^{\,r}, \qquad r=1,2,\dots,k . \]

Equating \(\mu'_r\) with \(a_r\) gives \(k\) equations. Solving them for the \(k\) parameters yields the estimates \(\hat\theta_1,\hat\theta_2,\dots,\hat\theta_k\), each a function of \(x_1,x_2,\dots,x_n\). That is the whole method: match as many sample moments as you have parameters.

\[ m'_r \;=\; \dfrac{1}{n}\sum X_i^r \;=\; \mu'_r(\theta). \]
EXAMPLE 1 (Exponential)

For \(X \sim \text{Exp}(\lambda)\), \(\mu'_1 = 1/\lambda\). MoM: \(\bar X = 1/\hat\lambda\) ⇒ \(\hat\lambda = 1/\bar X\).

EXAMPLE 2 (Uniform)

For \(X \sim U(0, \theta)\), \(\mu'_1 = \theta/2\). MoM: \(\hat\theta = 2 \bar X\), which is unbiased. The MLE \(\max(X_i)\) is biased (its mean is \(n\theta/(n+1)\)) but has much smaller variance, so it wins on mean squared error.

PROBLEM — estimate \(\mu\) and \(\sigma^2\) of a normal population by the method of moments

Let \(x_1,x_2,\dots,x_n\) be a random sample drawn from a normal population with mean \(\mu\) and variance \(\sigma^2\). There are two parameters, so two moments are needed.

Population moments. From the normal distribution,

\[ \mu'_1 = E(X) = \mu, \qquad \mu'_2 = E(X^2) = \mu^2 + \sigma^2 . \]

Sample moments.

\[ a_1 = \frac1n\sum_{i=1}^{n} x_i = \bar x, \qquad a_2 = \frac1n\sum_{i=1}^{n} x_i^{2} . \]

Equate them.

\[ \mu = \bar x, \qquad \sigma^2 + \mu^2 = \frac1n\sum_{i=1}^{n} x_i^{2} . \]

The first equation gives the estimate of \(\mu\) directly:

\[ \hat\mu = \bar x . \]

Substituting it into the second and rearranging,

\[ \hat\sigma^{2} = \frac1n\sum_{i=1}^{n} x_i^{2} - \hat\mu^{2} = \frac1n\sum_{i=1}^{n} x_i^{2} - \bar x^{2} = \frac1n\sum_{i=1}^{n}(x_i-\bar x)^{2} = s^{2}. \]

So the method of moments gives \(\hat\mu=\bar x\) and \(\hat\sigma^2=s^2\) — the same pair the method of maximum likelihood produces below. The two methods agree here; they do not always.

3.2 Method of Maximum Likelihood (MLE)

The most general method of obtaining a good estimator. It was first formulated by Gauss and introduced by R. A. Fisher.

The likelihood function is \(L(\theta) = \prod f(x_i; \theta)\). The principle of maximum likelihood is to find, for an unknown parameter \(\theta=(\theta_1,\theta_2,\dots,\theta_K)\), the value \(\hat\theta=(\hat\theta_1,\hat\theta_2,\dots,\hat\theta_K)\) that maximises \(L(x,\theta)\) over variations in the parameter:

\[ L(\hat\theta) \ \ge\ L(\theta) \qquad \forall\ \theta\in\Theta . \]

If a function \(\hat\theta = \hat\theta(x_1,x_2,\dots,x_n)\) of the sample values maximises \(L\), that \(\hat\theta\) is taken as the estimator of \(\theta\), and is called the maximum likelihood estimator of \(\theta\).

WHY THE LOG

By the principle of maxima and minima, maximising \(L\) with respect to \(\theta\) requires

\[ \frac{\partial L}{\partial\theta}=0 \quad\text{and}\quad \frac{\partial^{2} L}{\partial\theta^{2}}<0 . \]

Since \(L>0\) and \(\log L\) is a non-decreasing function of \(L\), the same turning point is found by differentiating the logarithm instead:

\[ \frac{1}{L}\frac{\partial L}{\partial\theta}=0 \;\Longrightarrow\; \frac{\partial \log L}{\partial\theta}=0, \qquad\text{and}\qquad \frac{\partial^{2}\log L}{\partial\theta^{2}}<0 . \]

This is why every problem below differentiates \(\log L\) and not \(L\): the product becomes a sum, and the answer is the same.

Regularity Conditions for the MLE

The properties of the MLE hold under the following assumptions, known as the regularity conditions.

  1. The first and second order derivatives \(\dfrac{\partial \log L}{\partial\theta}\) and \(\dfrac{\partial^{2}\log L}{\partial\theta^{2}}\) exist and are continuous functions of \(\theta\) for \((x,\theta)\in R\).
  2. \(\dfrac{\partial^{3}\log L}{\partial\theta^{3}}\) should exist and \(\left|\dfrac{\partial^{3}\log L}{\partial\theta^{3}}\right|<M(x)\), where \(E[M(x)]<K\) and \(K>0\).
  3. For every \(\theta\in R\), \[ E\!\left[\frac{-\partial^{2}\log L}{\partial\theta^{2}}\right] = \int_{-\infty}^{\infty}\!\int_{-\infty}^{\infty}\cdots\int_{-\infty}^{\infty} \frac{-\partial^{2}\log L}{\partial\theta^{2}}\,L(X;\theta)\,dx_1\,dx_2\cdots dx_n = I(\theta), \] where \(I(\theta)\) is a finite non-zero quantity known as the information supplied by the sample \(x_1,x_2,\dots,x_n\).
  4. The range of integration is independent of \(\theta\).

Condition 4 is the one most often violated in practice. For \(X\sim U(0,\theta)\) the range of the data depends on \(\theta\), which is exactly why the MLE there is \(\max(X_i)\) and not something found by setting a derivative to zero.

Properties of MLE

  1. Under the regularity conditions, MLEs are consistent, but they need not be unbiased.
    Example. The MLE of the population variance \(\sigma^2\) in a normal population is the sample variance \(s^2=\frac1n\sum(x_i-\bar x)^2\). That \(s^2\) is a consistent estimator but not an unbiased one.
  2. Any consistent solution of the likelihood equation gives a solution with probability tending to unity as \(n\to\infty\).
  3. A consistent solution of the likelihood equation is asymptotically normally distributed about the true value \(\theta_0\). Hence \(\hat\theta\) is asymptotically \(N\!\left(\theta_0,\ \dfrac{1}{I(\theta_0)}\right)\).

    Note. The variance of the MLE is

    \[ V(\hat\theta) = \frac{1}{I(\theta)} = \frac{1}{E\!\left(\dfrac{-\partial^{2}\log L}{\partial\theta^{2}}\right)} . \]
  4. If the MLE exists, it is the most efficient estimator in the class of such estimators.
  5. If a sufficient estimator exists, it is a function of the MLE.
  6. If \(t\) is the MLE of \(\theta\) and \(\psi(\theta)\) is a one-to-one function of \(\theta\), then \(\psi(t)\) is the MLE of \(\psi(\theta)\). This is called the invariance property of the MLE.
  7. If for a given population a minimum variance bound (MVB) estimator \(t\) exists for \(\theta\), then the likelihood equation will have a solution equal to \(t\).

MLE for Standard Distributions

EXAMPLE 1 (Bernoulli)

\(L(p) = p^{\sum x_i}(1-p)^{n-\sum x_i}\). \(\frac{d}{dp}\log L = \sum x_i / p - (n - \sum x_i)/(1-p) = 0\) ⇒ \(\hat p = \bar X\).

EXAMPLE 2 (Normal mean, σ known)

\(\log L = -\frac{n}{2}\log(2\pi\sigma^2) - \frac{1}{2\sigma^2}\sum(x_i - \mu)^2\). \(\partial/\partial\mu\) = 0 gives \(\hat\mu = \bar X\).

Worked MLE Problems

PROBLEM 1 — MLE in a normal population

Obtain the MLE for (a) \(\mu\) when \(\sigma^2\) is known, (b) \(\sigma^2\) when \(\mu\) is known, and (c) \(\mu\) and \(\sigma^2\) simultaneously, in a normal population with mean \(\mu\) and variance \(\sigma^2\).

Solution. Let \(x_1,x_2,\dots,x_n\) be a random sample drawn from a normal population with density

\[ f(x_i,\mu,\sigma^2) = \frac{1}{\sigma\sqrt{2\pi}}\cdot e^{-\frac12\left(\frac{x_i-\mu}{\sigma}\right)^{2}}, \qquad i=1,2,\dots,n. \]

The likelihood function is

\[ L = f(x_1;\mu,\sigma^2)\,f(x_2;\mu,\sigma^2)\cdots f(x_n;\mu,\sigma^2) = \frac{1}{(\sigma\sqrt{2\pi})^{n}}\cdot e^{-\frac{1}{2\sigma^{2}}\sum_{i=1}^{n}(x_i-\mu)^{2}} . \]

Taking logarithms and simplifying,

\[ \log L = -n\log(\sigma\sqrt{2\pi}) - \frac{1}{2\sigma^{2}}\sum_{i=1}^{n}(x_i-\mu)^{2} \] \[ = -n\log\big[(\sigma\sqrt{2\pi})^{2}\big]^{1/2} - \frac{1}{2\sigma^{2}}\sum_{i=1}^{n}(x_i-\mu)^{2} \] \[ = -\frac{n}{2}\log(\sigma^{2}2\pi) - \frac{1}{2\sigma^{2}}\sum_{i=1}^{n}(x_i-\mu)^{2} \] \[ = -\frac{n}{2}\log\sigma^{2} - \frac{n}{2}\log 2\pi - \frac{1}{2\sigma^{2}}\sum_{i=1}^{n}(x_i-\mu)^{2} . \]

(a) \(\mu\) when \(\sigma^2\) is known.

\[ \frac{\partial \log L}{\partial\mu}=0 \;\Longrightarrow\; \frac{-1}{2\sigma^{2}}\cdot 2\sum_{i=1}^{n}(x_i-\mu)(-1)=0 \] \[ \Longrightarrow\; \sum_{i=1}^{n}(x_i-\mu)=0 \;\Longrightarrow\; \sum_{i=1}^{n} x_i - n\mu = 0 \;\Longrightarrow\; n\mu = \sum_{i=1}^{n} x_i \] \[ \Longrightarrow\; \mu = \frac1n\sum_{i=1}^{n} x_i = \bar x \qquad\therefore\quad \hat\mu = \bar x . \]

That is, the MLE for the population mean \(\mu\) is the sample mean \(\bar x\) in a normal population.

(b) \(\sigma^2\) when \(\mu\) is known.

\[ \frac{\partial \log L}{\partial\sigma^{2}}=0 \;\Longrightarrow\; \frac{-n}{2}\cdot\frac{1}{\sigma^{2}} - \left(\frac{-1}{2(\sigma^{2})^{2}}\right)\sum_{i=1}^{n}(x_i-\mu)^{2}=0 \] \[ \Longrightarrow\; -n + \frac{\sum_{i=1}^{n}(x_i-\mu)^{2}}{\sigma^{2}} = 0 \;\Longrightarrow\; \frac{\sum_{i=1}^{n}(x_i-\mu)^{2}}{\sigma^{2}} = n \] \[ \therefore\quad \hat\sigma^{2} = \frac1n\sum_{i=1}^{n}(x_i-\mu)^{2}, \qquad \mu \text{ known}. \]

(c) Simultaneous estimation of \(\mu\) and \(\sigma^2\). Solve both likelihood equations together:

\[ \frac{\partial \log L}{\partial\mu}=0 \;\Longrightarrow\; \hat\mu = \bar x, \] \[ \frac{\partial \log L}{\partial\sigma^{2}}=0 \;\Longrightarrow\; \hat\sigma^{2} = \frac1n\sum_{i=1}^{n}(x_i-\mu)^{2} = \frac1n\sum_{i=1}^{n}(x_i-\hat\mu)^{2} = \frac1n\sum_{i=1}^{n}(x_i-\bar x)^{2} = s^{2}. \]

So the simultaneous MLE of \(\mu\) and \(\sigma^2\) is \((\bar x,\ s^2)\), the sample mean and the sample variance — note the divisor \(n\), not \(n-1\), which is exactly why the MLE of \(\sigma^2\) is biased.

PROBLEM 2 — MLE for \(\lambda\) in a Poisson population

Solution. Let \(x_1,x_2,\dots,x_n\) be a random sample of size \(n\) from a Poisson population, whose probability mass function is

\[ f(x_i,\lambda) = \frac{e^{-\lambda}\lambda^{x_i}}{x_i!}, \qquad x_i = 0,1,2,\dots,\ \lambda>0 . \]

The likelihood function is

\[ L = f(x_1,\lambda)\,f(x_2,\lambda)\cdots f(x_n,\lambda) = \frac{e^{-n\lambda}\cdot\lambda^{\sum_{i=1}^{n} x_i}}{x_1!\,x_2!\cdots x_n!} \] \[ \log L = -n\lambda + \sum_{i=1}^{n} x_i\log\lambda - \log(x_1!\,x_2!\cdots x_n!) \] \[ \frac{\partial \log L}{\partial\lambda}=0 \;\Longrightarrow\; -n + \frac{1}{\lambda}\sum_{i=1}^{n} x_i = 0 \;\Longrightarrow\; \frac{1}{\lambda}\sum_{i=1}^{n} x_i = n \] \[ \Longrightarrow\; \lambda = \frac{\sum_{i=1}^{n} x_i}{n} = \bar x \qquad\therefore\quad \hat\lambda = \bar x . \]

That is, the MLE for \(\lambda\) is the sample mean \(\bar x\) in the Poisson population.

The printed page gives the support as \(x_i>0\). A Poisson variable takes the value 0, so the support is \(x_i=0,1,2,\dots\), as written above. Nothing else in the derivation changes.

PROBLEM 3 — MLE for \(\lambda\) in an exponential population

Solution. Let \(x_1,x_2,\dots,x_n\) be a random sample of size \(n\) from an exponential population with density

\[ f(x_i,\lambda) = \lambda e^{-\lambda x_i}, \qquad x_i>0,\ \lambda>0,\ i=1,2,\dots,n. \]

The likelihood function is

\[ L = f(x_1,\lambda)\,f(x_2,\lambda)\cdots f(x_n,\lambda) = \lambda^{n} e^{-\lambda\sum_{i=1}^{n} x_i} \] \[ \log L = n\log\lambda - \lambda\sum_{i=1}^{n} x_i \] \[ \frac{\partial \log L}{\partial\lambda}=0 \;\Longrightarrow\; \frac{n}{\lambda} - \sum_{i=1}^{n} x_i = 0 \;\Longrightarrow\; \frac{n}{\lambda} = \sum_{i=1}^{n} x_i \] \[ \Longrightarrow\; \lambda = \frac{n}{\sum_{i=1}^{n} x_i} = \frac{1}{\bar x} \qquad\therefore\quad \hat\lambda = \frac{1}{\bar x} . \]

That is, the MLE for \(\lambda\) is the reciprocal of the sample mean. Compare this with the method-of-moments estimate in §3.1 above: for the exponential the two methods agree.

4. Rao–Cramer Inequality

STATEMENT

For an unbiased estimator \(\hat\theta\) of \(\theta\) (under regularity conditions),

\[ \text{Var}(\hat\theta) \;\ge\; \dfrac{1}{n\, I(\theta)}, \]

where the Fisher information is

\[ I(\theta) \;=\; E\!\left[\left(\dfrac{\partial \log f(X;\theta)}{\partial \theta}\right)^2\right] \;=\; -E\!\left[\dfrac{\partial^2 \log f(X;\theta)}{\partial \theta^2}\right]. \]
THE GENERAL FORM

The bound above is a special case. In general, if \(t\) is an unbiased estimator of \(\gamma(\theta)\), where \(\gamma(\theta)\) is a function of \(\theta\), then

\[ \operatorname{Var}(t) \;\ge\; \frac{\left[\dfrac{\partial}{\partial\theta}\gamma(\theta)\right]^{2}} {E\!\left(\dfrac{\partial}{\partial\theta}\log L\right)^{2}} \;=\; \frac{[\gamma'(\theta)]^{2}}{I(\theta)}, \]

where \(I(\theta)\) is the information on \(\theta\) supplied by the sample. Putting \(\gamma(\theta)=\theta\) gives \(\gamma'(\theta)=1\) and recovers the familiar \(\operatorname{Var}(\hat\theta)\ge 1/I(\theta)\).

Regularity Conditions for the Bound

  1. The parameter space \(\Theta\) is a non-degenerate open interval on the real line.
  2. For almost all \(X=(x_1,x_2,\dots,x_n)\) and for all \(\theta\in\Theta\), \(\dfrac{\partial}{\partial\theta}L(x,\theta)\) exists; the exceptional set, if any, is independent of \(\theta\).
  3. The range of integration is independent of the parameter \(\theta\), so that \(f(x,\theta)\) is differentiable under the integral sign.
  4. Differentiation and integration are interchangeable.
  5. \(I(\theta) = E\!\left[\dfrac{\partial}{\partial\theta}\log L\right]^{2}\) exists and is positive for all \(\theta\in\Theta\).

Proof of the Rao–Cramer Inequality

Let \(X\) be a random variable with probability density function \(f(x,\theta)\) and let \(L\) be the likelihood function of the random sample \(x_1,x_2,\dots,x_n\):

\[ L = L(x,\theta) = \prod_{i=1}^{n} f(x_i,\theta). \]

Since \(L\) is the joint probability density function of \(x_1,x_2,\dots,x_n\),

\[ \int\!\int\cdots\int L\,dx_1\,dx_2\cdots dx_n = 1 \qquad\text{i.e.}\qquad \int L\,dx = 1, \]

where \(\int dx\) is shorthand for the \(n\) integrals \(\int\!\int\cdots\int dx_1\,dx_2\cdots dx_n\).

Step 1 — the score has mean zero. Differentiating with respect to \(\theta\),

\[ \frac{\partial}{\partial\theta}\int L\,dx = 0 \;\Longrightarrow\; \int \frac{\partial L}{\partial\theta}\,dx = 0 \qquad \text{[using regularity condition 4]} \]

Multiplying and dividing by \(L\),

\[ \int \frac{1}{L}\frac{\partial L}{\partial\theta}\cdot L\,dx = 0 \;\Longrightarrow\; \int \frac{\partial \log L}{\partial\theta}\,L\,dx = 0 \]

— the middle step is just the chain rule of calculus — and therefore

\[ E\!\left[\frac{\partial \log L}{\partial\theta}\right] = 0 . \tag{1} \]

Step 2 — the score against the estimator. Let \(t=t(x_1,x_2,\dots,x_n)\) be an unbiased estimator of \(\gamma(\theta)\), so that \(E(t)=\gamma(\theta)\), that is \(\int t\,L\,dx = \gamma(\theta)\). Differentiating with respect to \(\theta\),

\[ \frac{\partial}{\partial\theta}\int t\,L\,dx = \gamma'(\theta) \;\Longrightarrow\; \int t\cdot\frac{\partial L}{\partial\theta}\,dx = \gamma'(\theta) \qquad \text{(using regularity condition 4)} \] \[ \Longrightarrow\; \int t\,\frac{1}{L}\frac{\partial L}{\partial\theta}\cdot L\,dx = \gamma'(\theta) \;\Longrightarrow\; \int t\,\frac{\partial \log L}{\partial\theta}\,L\,dx = \gamma'(\theta) \] \[ \Longrightarrow\; E\!\left[t\,\frac{\partial \log L}{\partial\theta}\right] = \gamma'(\theta) . \tag{2} \]

Step 3 — the covariance. The covariance of \(t\) and \(\dfrac{\partial \log L}{\partial\theta}\) is

\[ \operatorname{Cov}\!\left(t,\frac{\partial \log L}{\partial\theta}\right) = E\!\left(t\,\frac{\partial \log L}{\partial\theta}\right) - E(t)\cdot E\!\left(\frac{\partial \log L}{\partial\theta}\right) = \gamma'(\theta) \qquad \text{[from (1) and (2)]} \]

Step 4 — Cauchy–Schwarz. For any two random variables the squared correlation is at most one, \(\rho^{2}\le 1\), that is

\[ \left[\frac{\operatorname{Cov}(X,Y)}{\sqrt{\operatorname{Var}(X)\operatorname{Var}(Y)}}\right]^{2}\le 1 \;\Longrightarrow\; [\operatorname{Cov}(X,Y)]^{2}\le \operatorname{Var}(X)\cdot\operatorname{Var}(Y). \]

Take \(X=t\) and \(Y=\dfrac{\partial \log L}{\partial\theta}\):

\[ \left[\operatorname{Cov}\!\left(t,\frac{\partial \log L}{\partial\theta}\right)\right]^{2} \le \operatorname{Var}(t)\cdot\operatorname{Var}\!\left(\frac{\partial \log L}{\partial\theta}\right) \] \[ [\gamma'(\theta)]^{2} \le \operatorname{Var}(t) \left[ E\!\left(\frac{\partial \log L}{\partial\theta}\right)^{2} - \left\{E\!\left(\frac{\partial \log L}{\partial\theta}\right)\right\}^{2}\right] \le \operatorname{Var}(t)\left[E\!\left(\frac{\partial \log L}{\partial\theta}\right)^{2}\right] \qquad \text{[from (1)]} \]

Step 5 — rearrange.

\[ \operatorname{Var}(t)\,E\!\left(\frac{\partial \log L}{\partial\theta}\right)^{2} \ge [\gamma'(\theta)]^{2} \;\Longrightarrow\; \operatorname{Var}(t) \ge \frac{[\gamma'(\theta)]^{2}}{E\!\left[\dfrac{\partial \log L}{\partial\theta}\right]^{2}} . \quad\blacksquare \]

The Cramér–Rao inequality therefore provides a lower bound \(\dfrac{[\gamma'(\theta)]^{2}}{I(\theta)}\) to the variance of an unbiased estimator of \(\gamma(\theta)\).

Two notes on the printed proof. It opens by writing the likelihood as a sum, \(L=\sum f(x_i,\theta)\); it must be the product used above, since the very next line calls \(L\) a joint density and the proof does not work otherwise. And at Step 4 it writes \(\gamma^{2}\le1\) for the squared correlation, where \(\gamma\) already stands for \(\gamma(\theta)\) elsewhere in the same proof; the correlation is written \(\rho\) above to keep the two apart.

Properties

  1. Defines the minimum possible variance for any unbiased estimator (the "lower bound").
  2. An estimator that attains the bound is called most efficient (MVUE).
  3. MLEs asymptotically attain this bound.
EXAMPLE 1 (Bernoulli)

\(I(p) = 1/[p(1-p)]\). Lower bound = \(p(1-p)/n\). Variance of \(\bar X\) (for Bernoulli) = \(p(1-p)/n\) — bound attained.

EXAMPLE 2 (Normal mean, σ known)

\(I(\mu) = 1/\sigma^2\). Lower bound = \(\sigma^2/n\). Variance of \(\bar X\) = \(\sigma^2/n\) — attained.

5. Confidence Intervals (CIs)

DEFINITION

A \((1-\alpha)100\%\) confidence interval for \(\theta\) is a random interval \([L, U]\) such that

\[ P(L \le \theta \le U) \;=\; 1 - \alpha. \]

Common confidence levels: 90 % (\(z = 1.645\)), 95 % (\(z = 1.96\)), 99 % (\(z = 2.576\)).

Standard CI Formulas

−1.96 0 +1.96 95% α/2 = 2.5% α/2 = 2.5% 95% Confidence Interval (Z scale)
Fig 1.1 — Central 95 % area gives the standard CI
EXAMPLE 1 (Mean, σ known)

Sample of 64 light bulbs has \(\bar X = 1\,250\) hours; \(\sigma = 80\). 95 % CI for \(\mu\):

\(1\,250 \pm 1.96 (80/\sqrt{64}) = 1\,250 \pm 19.6 = (1\,230.4,\; 1\,269.6)\) hours.

EXAMPLE 2 (Proportion)

Out of 400 voters, 240 favour candidate A. \(\hat p = 0.6\). 95 % CI:

\(0.6 \pm 1.96\sqrt{0.6 \cdot 0.4 / 400} = 0.6 \pm 0.048 = (0.552,\; 0.648)\).

Key Take-aways

Where this goes next. The confidence interval you have just built already contains a test, without calling it one: if someone claims a value for \(\theta\) and that value falls outside your 95% interval, you have rejected the claim at the 5% level. Unit 2 makes that decision explicit and gives it a vocabulary — null hypothesis, critical region, and the two different ways of being wrong. The estimator built here becomes the test statistic there.

→ Unit 2 — Testing of Hypothesis