Two kinds of estimation:
The parameter space \(\Theta\) is the set of all admissible values a parameter can take. Estimation is the problem of using the sample to pick the value of \(\theta\in\Theta\) that generated the data.
Example. If \(X \sim N(\mu,\sigma^2)\) with density \(f(x;\mu,\sigma^2)\), then
\[ \Theta = \{(\mu,\sigma^2)\;:\; -\infty < \mu < \infty,\ \sigma > 0\}. \]An estimator is a map from the sample space into \(\Theta\); "best estimator" means the point of \(\Theta\) that best satisfies the criteria U–C–E–S below.
\(\hat\theta\) is unbiased for \(\theta\) if \(E(\hat\theta) = \theta\) for all \(\theta\).
Sample mean \(\bar X = \frac{1}{n}\sum X_i\) is unbiased for \(\mu\): \(E(\bar X) = \mu\).
Sample variance \(s^2 = \frac{1}{n-1}\sum(X_i - \bar X)^2\) is unbiased for \(\sigma^2\). The naïve version with \(1/n\) is biased: \(E(\sigma_n^2) = \frac{n-1}{n}\sigma^2\).
Let \(s^2 = \frac1n\sum(X_i-\bar X)^2 = \frac1n\sum X_i^2 - \bar X^2\). Using \(E(X_i^2)=\sigma^2+\mu^2\) and \(E(\bar X^2)=\dfrac{\sigma^2}{n}+\mu^2\),
\[ E(s^2) = \frac1n\sum E(X_i^2) - E(\bar X^2) = (\sigma^2+\mu^2) - \Big(\tfrac{\sigma^2}{n}+\mu^2\Big) = \Big(1-\tfrac1n\Big)\sigma^2 = \frac{n-1}{n}\,\sigma^2 . \]So \(s^2\) is biased (it underestimates \(\sigma^2\)). Rescaling removes the bias:
\[ E\!\left[\frac{n}{n-1}\,s^2\right] = \sigma^2 \quad\Longrightarrow\quad S^2 = \frac{1}{n-1}\sum_{i=1}^n (X_i-\bar X)^2 \ \text{ is unbiased for } \sigma^2 . \]If \(T\) is unbiased for \(\theta\), show that \(T^2\) is not unbiased for \(\theta^2\).
Solution. \(E(T)=\theta\). From \(\operatorname{Var}(T)=E(T^2)-[E(T)]^2\),
\[ E(T^2) = \theta^2 + \operatorname{Var}(T). \]Since \(\operatorname{Var}(T)>0\) (unless \(T\) is degenerate), \(E(T^2)\neq\theta^2\); \(T^2\) overestimates \(\theta^2\) by exactly \(\operatorname{Var}(T)\).
Let \(X_1,\dots,X_n\) take values \(1\) and \(0\) with probabilities \(\theta\) and \(1-\theta\). Show that \(\dfrac{\sum X_i\left(\sum X_i-1\right)}{n(n-1)}\) is unbiased for \(\theta^2\).
Solution. \(T=\sum X_i\sim\text{Bin}(n,\theta)\), so \(E(T)=n\theta\), \(\operatorname{Var}(T)=n\theta(1-\theta)\). Then
\[ E\!\left[\frac{T(T-1)}{n(n-1)}\right] = \frac{E(T^2)-E(T)}{n(n-1)} = \frac{\big[\operatorname{Var}(T)+(E T)^2\big]-E(T)}{n(n-1)}. \] \[ = \frac{n\theta(1-\theta)+n^2\theta^2-n\theta}{n(n-1)} = \frac{n^2\theta^2-n\theta^2}{n(n-1)} = \frac{n\theta^2(n-1)}{n(n-1)} = \theta^2 . \]Hence the statistic is unbiased for \(\theta^2\). (Verified numerically: for \(n=10,\ \theta=0.62\) the expectation equals \(0.3844=\theta^2\).)
For a sample of size \(n\) from \(f(x)=\dfrac{e^{-x/\theta}}{\theta},\ 0<x<\infty\), show that \(\bar X\) is unbiased for \(\theta\) with variance \(\theta^2/n\).
Solution. \(E(X)=\displaystyle\int_0^\infty x\,\frac{e^{-x/\theta}}{\theta}\,dx=\theta\) and \(E(X^2)=\displaystyle\int_0^\infty x^2\,\frac{e^{-x/\theta}}{\theta}\,dx=2\theta^2\), so \(\operatorname{Var}(X)=2\theta^2-\theta^2=\theta^2\). Therefore
\[ E(\bar X)=\frac1n\sum E(X_i)=\theta, \qquad \operatorname{Var}(\bar X)=\frac1{n^2}\sum\operatorname{Var}(X_i)=\frac{\theta^2}{n}. \]So \(\bar X\) is an unbiased, consistent estimator of \(\theta\).
\(\hat\theta_n\) is consistent if \(\hat\theta_n \xrightarrow{P} \theta\) as \(n \to \infty\). By Chebyshev: an unbiased estimator is consistent if \(\text{Var}(\hat\theta_n) \to 0\).
\(\bar X\) is consistent for \(\mu\): unbiased and \(\text{Var}(\bar X) = \sigma^2/n \to 0\).
The sample proportion \(\hat p = X/n\) (where \(X\) is the count of successes in Bernoulli trials) is consistent for \(p\) with variance \(p(1-p)/n \to 0\).
Statement. If \(t_n\) is a consistent estimator of \(\theta\) and \(\psi\) is a function continuous at \(\theta\), then \(\psi(t_n)\) is a consistent estimator of \(\psi(\theta)\).
Proof. Consistency means \(t_n \xrightarrow{P}\theta\): for every \(\varepsilon>0,\ \eta>0\) there is \(m\) such that \(P(|t_n-\theta|\le\varepsilon) > 1-\eta\) for all \(n\ge m\). Because \(\psi\) is continuous at \(\theta\), for every \(\varepsilon_1>0\) there is \(\varepsilon>0\) with
\[ |t_n-\theta|\le\varepsilon \;\Longrightarrow\; |\psi(t_n)-\psi(\theta)|\le\varepsilon_1 . \]Hence the event \(\{|t_n-\theta|\le\varepsilon\}\) is contained in \(\{|\psi(t_n)-\psi(\theta)|\le\varepsilon_1\}\), so
\[ P\big(|\psi(t_n)-\psi(\theta)|\le\varepsilon_1\big) \;\ge\; P\big(|t_n-\theta|\le\varepsilon\big) \;>\; 1-\eta \quad (n\ge m). \]Therefore \(\psi(t_n)\xrightarrow{P}\psi(\theta)\), i.e. \(\psi(t_n)\) is consistent for \(\psi(\theta)\). \(\blacksquare\)
Statement. If \(E(t_n)\to\theta\) and \(\operatorname{Var}(t_n)\to 0\) as \(n\to\infty\), then \(t_n\) is a consistent estimator of \(\theta\).
Proof. By Chebyshev's inequality applied to \(t_n\), for any \(\varepsilon>0\)
\[ P\big(|t_n-\theta|\ge\varepsilon\big) \;\le\; \frac{E\big[(t_n-\theta)^2\big]}{\varepsilon^2} \;=\; \frac{\operatorname{Var}(t_n) + \big(E(t_n)-\theta\big)^2}{\varepsilon^2}, \]since the mean-squared error splits as variance plus squared bias. As \(n\to\infty\) the bias \(E(t_n)-\theta\to0\) and \(\operatorname{Var}(t_n)\to0\), so the right-hand side \(\to0\). Hence \(P(|t_n-\theta|\ge\varepsilon)\to0\), i.e. \(t_n\xrightarrow{P}\theta\). \(\blacksquare\)
For \(X_1,\dots,X_n\) with \(E(X_i)=\mu,\ \operatorname{Var}(X_i)=\sigma^2\),
\[ E(\bar X) = \frac1n\sum_{i=1}^n E(X_i) = \frac1n\,(n\mu) = \mu \quad\Rightarrow\quad \text{unbiased.} \] \[ \operatorname{Var}(\bar X) = \frac1{n^2}\sum_{i=1}^n \operatorname{Var}(X_i) = \frac{\sigma^2}{n} \xrightarrow[n\to\infty]{} 0. \]Unbiased with variance \(\to0\); by Theorem 2, \(\bar X\) is consistent for \(\mu\).
For a normal sample, \(\dfrac{n s^2}{\sigma^2}\sim\chi^2_{(n-1)}\), whose mean is \(n-1\) and variance \(2(n-1)\). Hence
\[ E(s^2) = \frac{n-1}{n}\,\sigma^2 \to \sigma^2, \qquad \operatorname{Var}(s^2) = \frac{2(n-1)}{n^2}\,\sigma^4 \to 0 . \]Mean \(\to\sigma^2\) and variance \(\to0\); by Theorem 2, \(s^2\) is consistent for \(\sigma^2\) (the unbiased \(S^2\) is consistent too).
Among unbiased estimators, the one with minimum variance is most efficient. Relative efficiency:
\[ \text{RE}(\hat\theta_1, \hat\theta_2) = \dfrac{\text{Var}(\hat\theta_2)}{\text{Var}(\hat\theta_1)}. \]If RE > 1, \(\hat\theta_1\) is more efficient than \(\hat\theta_2\).
For a normal population, both \(\bar X\) and the sample median are unbiased for \(\mu\). \(\text{Var}(\bar X) = \sigma^2/n\), \(\text{Var}(\text{median}) \approx \pi\sigma^2/(2n) \approx 1.57\,\sigma^2/n\). Hence \(\bar X\) is about 57 % more efficient.
If \(T_1\) has variance 5 and \(T_2\) has variance 8 (both unbiased), \(\text{RE}(T_1,T_2) = 8/5 = 1.6\), so \(T_1\) is 60 % more efficient.
Among consistent estimators, the one with the smallest variance is the most efficient. If \(t_1\) is the most efficient estimator (variance \(V_1\)) and \(t_2\) is another (variance \(V_2\)), the efficiency of \(t_2\) is
\[ e = \frac{V_1}{V_2}, \qquad 0 < e \le 1 . \]Example. For a normal population both \(\bar X\) and the sample median \(M_d\) are unbiased and consistent for \(\mu\), with \(\operatorname{Var}(\bar X)=\sigma^2/n\) and \(\operatorname{Var}(M_d)=\dfrac{\pi\sigma^2}{2n}\approx 1.571\,\dfrac{\sigma^2}{n}\). Since \(\operatorname{Var}(\bar X)<\operatorname{Var}(M_d)\), the mean is most efficient, and the efficiency of the median is
\[ e = \frac{\operatorname{Var}(\bar X)}{\operatorname{Var}(M_d)} = \frac{\sigma^2/n}{\pi\sigma^2/2n} = \frac{2}{\pi} \approx 0.637 . \]A random sample of size 5 is drawn from a population with mean \(\mu\), variance \(\sigma^2\). Consider
\[ t_1=\frac{X_1+X_2+X_3+X_4+X_5}{5},\quad t_2=\frac{X_1+X_2}{2}+X_3,\quad t_3=\frac{2X_1+X_2+\lambda X_3}{3}. \](a) Value of \(\lambda\) making \(t_3\) unbiased. \(E(t_3)=\dfrac{2\mu+\mu+\lambda\mu}{3}=\dfrac{(3+\lambda)\mu}{3}\). Setting this \(=\mu\) gives \(3+\lambda=3\), so \(\boxed{\lambda=0}\).
(b) Are \(t_1,t_2\) unbiased? \(E(t_1)=\mu\) — unbiased. \(E(t_2)=\dfrac{\mu+\mu}{2}+\mu=2\mu\neq\mu\) — \(t_2\) is biased.
(c) Best estimator. Compare variances of the unbiased ones \(t_1\) and \(t_3\) (with \(\lambda=0\)):
\[ \operatorname{Var}(t_1)=\frac{5\sigma^2}{25}=\frac{\sigma^2}{5}=0.20\,\sigma^2,\qquad \operatorname{Var}(t_3)=\frac{(2^2+1^2)\sigma^2}{9}=\frac{5\sigma^2}{9}\approx0.556\,\sigma^2. \]Since \(\operatorname{Var}(t_1)<\operatorname{Var}(t_3)\), the sample mean \(t_1\) is the best (most efficient) estimator of \(\mu\).
A statistic \(T\) is sufficient for \(\theta\) if the conditional distribution of the sample given \(T\) does not depend on \(\theta\). It captures all the information in the data about \(\theta\).
\(T(\mathbf X)\) is sufficient for \(\theta\) iff the joint pdf factors as
\[ f(\mathbf x; \theta) \;=\; g(T(\mathbf x); \theta) \cdot h(\mathbf x), \]where \(h\) does not depend on \(\theta\).
For Bernoulli\((p)\) sample, \(\sum X_i\) is sufficient for \(p\). Joint pmf factors as \(p^{\sum x_i}(1-p)^{n-\sum x_i}\).
For \(N(\mu, \sigma^2)\) with known \(\sigma^2\), \(\sum X_i\) (or \(\bar X\)) is sufficient for \(\mu\).
For \(X_1,\dots,X_n \sim\) Poisson\((\lambda)\), \(L(\lambda) = e^{-n\lambda}\lambda^{\sum x_i}/\prod x_i!\). Here \(g = e^{-n\lambda}\lambda^{\sum x_i}\) depends on the data only through \(T = \sum X_i\), and \(h = 1/\prod x_i!\) is free of \(\lambda\); hence \(\sum X_i\) (equivalently \(\bar X\)) is sufficient for \(\lambda\).
Let \(x_1,x_2,\dots,x_n\) be a random sample from \(f(x,\lambda)=\lambda e^{-\lambda x},\ x>0,\ \lambda>0\). Find a sufficient statistic for \(\lambda\).
Solution. The likelihood function of \(x_1,x_2,\dots,x_n\) is
\[ L = f(x_1,\lambda)\,f(x_2,\lambda)\cdots f(x_n,\lambda) = \lambda e^{-\lambda x_1}\cdot\lambda e^{-\lambda x_2}\cdots\lambda e^{-\lambda x_n} = \lambda^{n} e^{-\lambda\sum_{i=1}^{n} x_i}. \] \[ = \underbrace{e^{-\lambda\sum_{i=1}^{n} x_i}\cdot\lambda^{n}}_{g_\lambda[t(x)]}\cdot \underbrace{1}_{h(x)} . \]Here \(g_\lambda[t(x)] = e^{-\lambda\sum x_i}\lambda^{n}\) is a function of \(\lambda\) and of the data only through \(t(x)=\sum_{i=1}^{n} x_i\), and \(h(x)=1\) is independent of \(\lambda\).
Therefore, by the factorization theorem, \(\displaystyle\sum_{i=1}^{n} x_i\) is a sufficient statistic for the population parameter \(\lambda\) of the exponential distribution.
Let \(x_1,x_2,\dots,x_n\) be a random sample of size \(n\) drawn from a normal population with parameters \(\mu\) and \(\sigma^2\), so that
\[ f(x_i,\mu,\sigma^2) = \frac{1}{\sigma\sqrt{2\pi}}\, e^{-\frac12\left(\frac{x_i-\mu}{\sigma}\right)^2}, \qquad -\infty<x_i<\infty,\ -\infty<\mu<\infty,\ \sigma>0 . \]Solution. The likelihood function is
\[ L = f(x_1;\mu,\sigma^2)\,f(x_2;\mu,\sigma^2)\cdots f(x_n;\mu,\sigma^2) = \frac{1}{(\sigma\sqrt{2\pi})^{n}}\, e^{-\frac{1}{2\sigma^{2}}\sum_{i=1}^{n}(x_i-\mu)^2}. \]Expanding the square, \(\sum (x_i-\mu)^2 = \sum x_i^2 - 2\mu\sum x_i + n\mu^2\), so
\[ L = \frac{1}{(\sigma\sqrt{2\pi})^{n}}\, e^{-\frac{1}{2\sigma^{2}}\left[\sum_{i=1}^{n} x_i^{2}-2\mu\sum_{i=1}^{n} x_i+n\mu^{2}\right]} \;=\; g_{(\mu,\sigma^2)}[t_1(x),t_2(x)]\cdot h(x), \]where \(g_{(\mu,\sigma^2)}\) is the whole expression above — a function of \(\mu,\sigma^2\) and of the data only through
\[ t_1(x)=\sum_{i=1}^{n} x_i, \qquad t_2(x)=\sum_{i=1}^{n} x_i^{2}, \]and \(h(x)=1\) is independent of \(\mu\) and \(\sigma^2\).
Therefore, by the factorization theorem, \(\sum x_i\) is a sufficient statistic for \(\mu\) and \(\sum x_i^{2}\) is a sufficient statistic for \(\sigma^2\).
Find a sufficient statistic for \(\theta\) for the density \(f(x,\theta)=\theta x^{\theta-1},\ 0<x<1,\ \theta>0\).
Solution. The likelihood function is
\[ L = f(x_1,\theta)\,f(x_2,\theta)\cdots f(x_n,\theta) = \theta x_1^{\theta-1}\cdot\theta x_2^{\theta-1}\cdots\theta x_n^{\theta-1} = \theta^{n}\,(x_1x_2\cdots x_n)^{\theta-1}. \] \[ = \frac{\theta^{n}(x_1x_2\cdots x_n)^{\theta}}{x_1x_2\cdots x_n} = \theta^{n}\left(\prod_{i=1}^{n} x_i\right)^{\!\theta}\cdot \frac{1}{\prod_{i=1}^{n} x_i} \;=\; g_\theta[t(x)]\cdot h(x), \]where \(g_\theta[t(x)] = \theta^{n}\left(\prod x_i\right)^{\theta}\) is a function of \(\theta\) and of the data only through \(t(x)=\prod_{i=1}^{n} x_i\), and \(h(x)=1/\prod x_i\) depends on the data alone and is independent of \(\theta\).
Therefore, by the factorization theorem, \(\displaystyle\prod_{i=1}^{n} x_i\) is a sufficient statistic for \(\theta\). Note that a sufficient statistic need not be a sum — here it is a product.
Among all unbiased estimators of \(\theta\), the one with the smallest variance is the MVUE. Two theorems locate it:
Example. For Poisson\((\lambda)\), \(\bar X\) is an unbiased function of the complete sufficient statistic \(\sum X_i\); by Lehmann–Scheffé it is the MVUE of \(\lambda\).
There are several methods of obtaining a good — or the best — estimator for a population parameter:
The first two are the ones this unit develops, and the two the syllabus examines. The third, the principle of least squares, is already covered under curve fitting.
Introduced by Prof. Karl Pearson. Let \(x_1,x_2,\dots,x_n\) be a random sample drawn from a population with density \(f(x;\theta_1,\theta_2,\dots,\theta_k)\), where \(\theta_1,\dots,\theta_k\) are \(k\) parameters. The first \(k\) moments of the population about the origin are
\[ \mu'_r = E(X^r) = \int_{-\infty}^{\infty} x^r f(x;\theta_1,\theta_2,\dots,\theta_k)\,dx, \qquad r=1,2,\dots,k, \]and each \(\mu'_r\) is in turn a function of \(\theta_1,\theta_2,\dots,\theta_k\). The corresponding moments of the sample are
\[ a_r = \frac{1}{n}\sum_{i=1}^{n} x_i^{\,r}, \qquad r=1,2,\dots,k . \]Equating \(\mu'_r\) with \(a_r\) gives \(k\) equations. Solving them for the \(k\) parameters yields the estimates \(\hat\theta_1,\hat\theta_2,\dots,\hat\theta_k\), each a function of \(x_1,x_2,\dots,x_n\). That is the whole method: match as many sample moments as you have parameters.
For \(X \sim \text{Exp}(\lambda)\), \(\mu'_1 = 1/\lambda\). MoM: \(\bar X = 1/\hat\lambda\) ⇒ \(\hat\lambda = 1/\bar X\).
For \(X \sim U(0, \theta)\), \(\mu'_1 = \theta/2\). MoM: \(\hat\theta = 2 \bar X\), which is unbiased. The MLE \(\max(X_i)\) is biased (its mean is \(n\theta/(n+1)\)) but has much smaller variance, so it wins on mean squared error.
Let \(x_1,x_2,\dots,x_n\) be a random sample drawn from a normal population with mean \(\mu\) and variance \(\sigma^2\). There are two parameters, so two moments are needed.
Population moments. From the normal distribution,
\[ \mu'_1 = E(X) = \mu, \qquad \mu'_2 = E(X^2) = \mu^2 + \sigma^2 . \]Sample moments.
\[ a_1 = \frac1n\sum_{i=1}^{n} x_i = \bar x, \qquad a_2 = \frac1n\sum_{i=1}^{n} x_i^{2} . \]Equate them.
\[ \mu = \bar x, \qquad \sigma^2 + \mu^2 = \frac1n\sum_{i=1}^{n} x_i^{2} . \]The first equation gives the estimate of \(\mu\) directly:
\[ \hat\mu = \bar x . \]Substituting it into the second and rearranging,
\[ \hat\sigma^{2} = \frac1n\sum_{i=1}^{n} x_i^{2} - \hat\mu^{2} = \frac1n\sum_{i=1}^{n} x_i^{2} - \bar x^{2} = \frac1n\sum_{i=1}^{n}(x_i-\bar x)^{2} = s^{2}. \]So the method of moments gives \(\hat\mu=\bar x\) and \(\hat\sigma^2=s^2\) — the same pair the method of maximum likelihood produces below. The two methods agree here; they do not always.
The most general method of obtaining a good estimator. It was first formulated by Gauss and introduced by R. A. Fisher.
The likelihood function is \(L(\theta) = \prod f(x_i; \theta)\). The principle of maximum likelihood is to find, for an unknown parameter \(\theta=(\theta_1,\theta_2,\dots,\theta_K)\), the value \(\hat\theta=(\hat\theta_1,\hat\theta_2,\dots,\hat\theta_K)\) that maximises \(L(x,\theta)\) over variations in the parameter:
\[ L(\hat\theta) \ \ge\ L(\theta) \qquad \forall\ \theta\in\Theta . \]If a function \(\hat\theta = \hat\theta(x_1,x_2,\dots,x_n)\) of the sample values maximises \(L\), that \(\hat\theta\) is taken as the estimator of \(\theta\), and is called the maximum likelihood estimator of \(\theta\).
By the principle of maxima and minima, maximising \(L\) with respect to \(\theta\) requires
\[ \frac{\partial L}{\partial\theta}=0 \quad\text{and}\quad \frac{\partial^{2} L}{\partial\theta^{2}}<0 . \]Since \(L>0\) and \(\log L\) is a non-decreasing function of \(L\), the same turning point is found by differentiating the logarithm instead:
\[ \frac{1}{L}\frac{\partial L}{\partial\theta}=0 \;\Longrightarrow\; \frac{\partial \log L}{\partial\theta}=0, \qquad\text{and}\qquad \frac{\partial^{2}\log L}{\partial\theta^{2}}<0 . \]This is why every problem below differentiates \(\log L\) and not \(L\): the product becomes a sum, and the answer is the same.
The properties of the MLE hold under the following assumptions, known as the regularity conditions.
Condition 4 is the one most often violated in practice. For \(X\sim U(0,\theta)\) the range of the data depends on \(\theta\), which is exactly why the MLE there is \(\max(X_i)\) and not something found by setting a derivative to zero.
Note. The variance of the MLE is
\[ V(\hat\theta) = \frac{1}{I(\theta)} = \frac{1}{E\!\left(\dfrac{-\partial^{2}\log L}{\partial\theta^{2}}\right)} . \]\(L(p) = p^{\sum x_i}(1-p)^{n-\sum x_i}\). \(\frac{d}{dp}\log L = \sum x_i / p - (n - \sum x_i)/(1-p) = 0\) ⇒ \(\hat p = \bar X\).
\(\log L = -\frac{n}{2}\log(2\pi\sigma^2) - \frac{1}{2\sigma^2}\sum(x_i - \mu)^2\). \(\partial/\partial\mu\) = 0 gives \(\hat\mu = \bar X\).
Obtain the MLE for (a) \(\mu\) when \(\sigma^2\) is known, (b) \(\sigma^2\) when \(\mu\) is known, and (c) \(\mu\) and \(\sigma^2\) simultaneously, in a normal population with mean \(\mu\) and variance \(\sigma^2\).
Solution. Let \(x_1,x_2,\dots,x_n\) be a random sample drawn from a normal population with density
\[ f(x_i,\mu,\sigma^2) = \frac{1}{\sigma\sqrt{2\pi}}\cdot e^{-\frac12\left(\frac{x_i-\mu}{\sigma}\right)^{2}}, \qquad i=1,2,\dots,n. \]The likelihood function is
\[ L = f(x_1;\mu,\sigma^2)\,f(x_2;\mu,\sigma^2)\cdots f(x_n;\mu,\sigma^2) = \frac{1}{(\sigma\sqrt{2\pi})^{n}}\cdot e^{-\frac{1}{2\sigma^{2}}\sum_{i=1}^{n}(x_i-\mu)^{2}} . \]Taking logarithms and simplifying,
\[ \log L = -n\log(\sigma\sqrt{2\pi}) - \frac{1}{2\sigma^{2}}\sum_{i=1}^{n}(x_i-\mu)^{2} \] \[ = -n\log\big[(\sigma\sqrt{2\pi})^{2}\big]^{1/2} - \frac{1}{2\sigma^{2}}\sum_{i=1}^{n}(x_i-\mu)^{2} \] \[ = -\frac{n}{2}\log(\sigma^{2}2\pi) - \frac{1}{2\sigma^{2}}\sum_{i=1}^{n}(x_i-\mu)^{2} \] \[ = -\frac{n}{2}\log\sigma^{2} - \frac{n}{2}\log 2\pi - \frac{1}{2\sigma^{2}}\sum_{i=1}^{n}(x_i-\mu)^{2} . \](a) \(\mu\) when \(\sigma^2\) is known.
\[ \frac{\partial \log L}{\partial\mu}=0 \;\Longrightarrow\; \frac{-1}{2\sigma^{2}}\cdot 2\sum_{i=1}^{n}(x_i-\mu)(-1)=0 \] \[ \Longrightarrow\; \sum_{i=1}^{n}(x_i-\mu)=0 \;\Longrightarrow\; \sum_{i=1}^{n} x_i - n\mu = 0 \;\Longrightarrow\; n\mu = \sum_{i=1}^{n} x_i \] \[ \Longrightarrow\; \mu = \frac1n\sum_{i=1}^{n} x_i = \bar x \qquad\therefore\quad \hat\mu = \bar x . \]That is, the MLE for the population mean \(\mu\) is the sample mean \(\bar x\) in a normal population.
(b) \(\sigma^2\) when \(\mu\) is known.
\[ \frac{\partial \log L}{\partial\sigma^{2}}=0 \;\Longrightarrow\; \frac{-n}{2}\cdot\frac{1}{\sigma^{2}} - \left(\frac{-1}{2(\sigma^{2})^{2}}\right)\sum_{i=1}^{n}(x_i-\mu)^{2}=0 \] \[ \Longrightarrow\; -n + \frac{\sum_{i=1}^{n}(x_i-\mu)^{2}}{\sigma^{2}} = 0 \;\Longrightarrow\; \frac{\sum_{i=1}^{n}(x_i-\mu)^{2}}{\sigma^{2}} = n \] \[ \therefore\quad \hat\sigma^{2} = \frac1n\sum_{i=1}^{n}(x_i-\mu)^{2}, \qquad \mu \text{ known}. \](c) Simultaneous estimation of \(\mu\) and \(\sigma^2\). Solve both likelihood equations together:
\[ \frac{\partial \log L}{\partial\mu}=0 \;\Longrightarrow\; \hat\mu = \bar x, \] \[ \frac{\partial \log L}{\partial\sigma^{2}}=0 \;\Longrightarrow\; \hat\sigma^{2} = \frac1n\sum_{i=1}^{n}(x_i-\mu)^{2} = \frac1n\sum_{i=1}^{n}(x_i-\hat\mu)^{2} = \frac1n\sum_{i=1}^{n}(x_i-\bar x)^{2} = s^{2}. \]So the simultaneous MLE of \(\mu\) and \(\sigma^2\) is \((\bar x,\ s^2)\), the sample mean and the sample variance — note the divisor \(n\), not \(n-1\), which is exactly why the MLE of \(\sigma^2\) is biased.
Solution. Let \(x_1,x_2,\dots,x_n\) be a random sample of size \(n\) from a Poisson population, whose probability mass function is
\[ f(x_i,\lambda) = \frac{e^{-\lambda}\lambda^{x_i}}{x_i!}, \qquad x_i = 0,1,2,\dots,\ \lambda>0 . \]The likelihood function is
\[ L = f(x_1,\lambda)\,f(x_2,\lambda)\cdots f(x_n,\lambda) = \frac{e^{-n\lambda}\cdot\lambda^{\sum_{i=1}^{n} x_i}}{x_1!\,x_2!\cdots x_n!} \] \[ \log L = -n\lambda + \sum_{i=1}^{n} x_i\log\lambda - \log(x_1!\,x_2!\cdots x_n!) \] \[ \frac{\partial \log L}{\partial\lambda}=0 \;\Longrightarrow\; -n + \frac{1}{\lambda}\sum_{i=1}^{n} x_i = 0 \;\Longrightarrow\; \frac{1}{\lambda}\sum_{i=1}^{n} x_i = n \] \[ \Longrightarrow\; \lambda = \frac{\sum_{i=1}^{n} x_i}{n} = \bar x \qquad\therefore\quad \hat\lambda = \bar x . \]That is, the MLE for \(\lambda\) is the sample mean \(\bar x\) in the Poisson population.
The printed page gives the support as \(x_i>0\). A Poisson variable takes the value 0, so the support is \(x_i=0,1,2,\dots\), as written above. Nothing else in the derivation changes.
Solution. Let \(x_1,x_2,\dots,x_n\) be a random sample of size \(n\) from an exponential population with density
\[ f(x_i,\lambda) = \lambda e^{-\lambda x_i}, \qquad x_i>0,\ \lambda>0,\ i=1,2,\dots,n. \]The likelihood function is
\[ L = f(x_1,\lambda)\,f(x_2,\lambda)\cdots f(x_n,\lambda) = \lambda^{n} e^{-\lambda\sum_{i=1}^{n} x_i} \] \[ \log L = n\log\lambda - \lambda\sum_{i=1}^{n} x_i \] \[ \frac{\partial \log L}{\partial\lambda}=0 \;\Longrightarrow\; \frac{n}{\lambda} - \sum_{i=1}^{n} x_i = 0 \;\Longrightarrow\; \frac{n}{\lambda} = \sum_{i=1}^{n} x_i \] \[ \Longrightarrow\; \lambda = \frac{n}{\sum_{i=1}^{n} x_i} = \frac{1}{\bar x} \qquad\therefore\quad \hat\lambda = \frac{1}{\bar x} . \]That is, the MLE for \(\lambda\) is the reciprocal of the sample mean. Compare this with the method-of-moments estimate in §3.1 above: for the exponential the two methods agree.
For an unbiased estimator \(\hat\theta\) of \(\theta\) (under regularity conditions),
\[ \text{Var}(\hat\theta) \;\ge\; \dfrac{1}{n\, I(\theta)}, \]where the Fisher information is
\[ I(\theta) \;=\; E\!\left[\left(\dfrac{\partial \log f(X;\theta)}{\partial \theta}\right)^2\right] \;=\; -E\!\left[\dfrac{\partial^2 \log f(X;\theta)}{\partial \theta^2}\right]. \]The bound above is a special case. In general, if \(t\) is an unbiased estimator of \(\gamma(\theta)\), where \(\gamma(\theta)\) is a function of \(\theta\), then
\[ \operatorname{Var}(t) \;\ge\; \frac{\left[\dfrac{\partial}{\partial\theta}\gamma(\theta)\right]^{2}} {E\!\left(\dfrac{\partial}{\partial\theta}\log L\right)^{2}} \;=\; \frac{[\gamma'(\theta)]^{2}}{I(\theta)}, \]where \(I(\theta)\) is the information on \(\theta\) supplied by the sample. Putting \(\gamma(\theta)=\theta\) gives \(\gamma'(\theta)=1\) and recovers the familiar \(\operatorname{Var}(\hat\theta)\ge 1/I(\theta)\).
Let \(X\) be a random variable with probability density function \(f(x,\theta)\) and let \(L\) be the likelihood function of the random sample \(x_1,x_2,\dots,x_n\):
\[ L = L(x,\theta) = \prod_{i=1}^{n} f(x_i,\theta). \]Since \(L\) is the joint probability density function of \(x_1,x_2,\dots,x_n\),
\[ \int\!\int\cdots\int L\,dx_1\,dx_2\cdots dx_n = 1 \qquad\text{i.e.}\qquad \int L\,dx = 1, \]where \(\int dx\) is shorthand for the \(n\) integrals \(\int\!\int\cdots\int dx_1\,dx_2\cdots dx_n\).
Step 1 — the score has mean zero. Differentiating with respect to \(\theta\),
\[ \frac{\partial}{\partial\theta}\int L\,dx = 0 \;\Longrightarrow\; \int \frac{\partial L}{\partial\theta}\,dx = 0 \qquad \text{[using regularity condition 4]} \]Multiplying and dividing by \(L\),
\[ \int \frac{1}{L}\frac{\partial L}{\partial\theta}\cdot L\,dx = 0 \;\Longrightarrow\; \int \frac{\partial \log L}{\partial\theta}\,L\,dx = 0 \]— the middle step is just the chain rule of calculus — and therefore
\[ E\!\left[\frac{\partial \log L}{\partial\theta}\right] = 0 . \tag{1} \]Step 2 — the score against the estimator. Let \(t=t(x_1,x_2,\dots,x_n)\) be an unbiased estimator of \(\gamma(\theta)\), so that \(E(t)=\gamma(\theta)\), that is \(\int t\,L\,dx = \gamma(\theta)\). Differentiating with respect to \(\theta\),
\[ \frac{\partial}{\partial\theta}\int t\,L\,dx = \gamma'(\theta) \;\Longrightarrow\; \int t\cdot\frac{\partial L}{\partial\theta}\,dx = \gamma'(\theta) \qquad \text{(using regularity condition 4)} \] \[ \Longrightarrow\; \int t\,\frac{1}{L}\frac{\partial L}{\partial\theta}\cdot L\,dx = \gamma'(\theta) \;\Longrightarrow\; \int t\,\frac{\partial \log L}{\partial\theta}\,L\,dx = \gamma'(\theta) \] \[ \Longrightarrow\; E\!\left[t\,\frac{\partial \log L}{\partial\theta}\right] = \gamma'(\theta) . \tag{2} \]Step 3 — the covariance. The covariance of \(t\) and \(\dfrac{\partial \log L}{\partial\theta}\) is
\[ \operatorname{Cov}\!\left(t,\frac{\partial \log L}{\partial\theta}\right) = E\!\left(t\,\frac{\partial \log L}{\partial\theta}\right) - E(t)\cdot E\!\left(\frac{\partial \log L}{\partial\theta}\right) = \gamma'(\theta) \qquad \text{[from (1) and (2)]} \]Step 4 — Cauchy–Schwarz. For any two random variables the squared correlation is at most one, \(\rho^{2}\le 1\), that is
\[ \left[\frac{\operatorname{Cov}(X,Y)}{\sqrt{\operatorname{Var}(X)\operatorname{Var}(Y)}}\right]^{2}\le 1 \;\Longrightarrow\; [\operatorname{Cov}(X,Y)]^{2}\le \operatorname{Var}(X)\cdot\operatorname{Var}(Y). \]Take \(X=t\) and \(Y=\dfrac{\partial \log L}{\partial\theta}\):
\[ \left[\operatorname{Cov}\!\left(t,\frac{\partial \log L}{\partial\theta}\right)\right]^{2} \le \operatorname{Var}(t)\cdot\operatorname{Var}\!\left(\frac{\partial \log L}{\partial\theta}\right) \] \[ [\gamma'(\theta)]^{2} \le \operatorname{Var}(t) \left[ E\!\left(\frac{\partial \log L}{\partial\theta}\right)^{2} - \left\{E\!\left(\frac{\partial \log L}{\partial\theta}\right)\right\}^{2}\right] \le \operatorname{Var}(t)\left[E\!\left(\frac{\partial \log L}{\partial\theta}\right)^{2}\right] \qquad \text{[from (1)]} \]Step 5 — rearrange.
\[ \operatorname{Var}(t)\,E\!\left(\frac{\partial \log L}{\partial\theta}\right)^{2} \ge [\gamma'(\theta)]^{2} \;\Longrightarrow\; \operatorname{Var}(t) \ge \frac{[\gamma'(\theta)]^{2}}{E\!\left[\dfrac{\partial \log L}{\partial\theta}\right]^{2}} . \quad\blacksquare \]The Cramér–Rao inequality therefore provides a lower bound \(\dfrac{[\gamma'(\theta)]^{2}}{I(\theta)}\) to the variance of an unbiased estimator of \(\gamma(\theta)\).
Two notes on the printed proof. It opens by writing the likelihood as a sum, \(L=\sum f(x_i,\theta)\); it must be the product used above, since the very next line calls \(L\) a joint density and the proof does not work otherwise. And at Step 4 it writes \(\gamma^{2}\le1\) for the squared correlation, where \(\gamma\) already stands for \(\gamma(\theta)\) elsewhere in the same proof; the correlation is written \(\rho\) above to keep the two apart.
\(I(p) = 1/[p(1-p)]\). Lower bound = \(p(1-p)/n\). Variance of \(\bar X\) (for Bernoulli) = \(p(1-p)/n\) — bound attained.
\(I(\mu) = 1/\sigma^2\). Lower bound = \(\sigma^2/n\). Variance of \(\bar X\) = \(\sigma^2/n\) — attained.
A \((1-\alpha)100\%\) confidence interval for \(\theta\) is a random interval \([L, U]\) such that
\[ P(L \le \theta \le U) \;=\; 1 - \alpha. \]Common confidence levels: 90 % (\(z = 1.645\)), 95 % (\(z = 1.96\)), 99 % (\(z = 2.576\)).
Sample of 64 light bulbs has \(\bar X = 1\,250\) hours; \(\sigma = 80\). 95 % CI for \(\mu\):
\(1\,250 \pm 1.96 (80/\sqrt{64}) = 1\,250 \pm 19.6 = (1\,230.4,\; 1\,269.6)\) hours.
Out of 400 voters, 240 favour candidate A. \(\hat p = 0.6\). 95 % CI:
\(0.6 \pm 1.96\sqrt{0.6 \cdot 0.4 / 400} = 0.6 \pm 0.048 = (0.552,\; 0.648)\).
Where this goes next. The confidence interval you have just built already contains a test, without calling it one: if someone claims a value for \(\theta\) and that value falls outside your 95% interval, you have rejected the claim at the 5% level. Unit 2 makes that decision explicit and gives it a vocabulary — null hypothesis, critical region, and the two different ways of being wrong. The estimator built here becomes the test statistic there.