Standard Normal — definition, MGF, mean, variance, area property; Population, Sample, Parameter, Statistic, Sampling Distribution, Standard Error; \(\chi^2\), Student's and Fisher's \(t\), \(F\) — densities derived, moments, properties, relations, limiting forms & applications.
Topics Covered
Standard NormalArea PropertyPopulation & SampleParameter / StatisticSampling DistributionStudent's tF-DistributionChi-squareStandard ErrorDistribution of s²Moments & ModeLimiting Forms
If \(X \sim N(\mu, \sigma^2)\), the standardized variable \(Z = \dfrac{X - \mu}{\sigma}\) follows the
Standard Normal distribution \(N(0, 1)\) with PDF
Population: the totality of all units of interest. Size \(N\) (finite or infinite).
Sample: a subset of the population, size \(n\).
Parameter: a numerical characteristic of the population (e.g. \(\mu, \sigma^2, p\)). Usually unknown.
Statistic: a numerical characteristic of the sample (e.g. \(\bar X, s^2, \hat p\)). A function of the sample.
Sampling Distribution
The probability distribution of a statistic, computed over all possible samples of a fixed size \(n\),
is called its sampling distribution. Its standard deviation is the standard error (SE).
Examples:
If \(X_i \sim N(\mu, \sigma^2)\): \(\bar X \sim N(\mu, \sigma^2/n)\).
For any population (finite variance) and large \(n\), CLT gives \(\bar X \approx N(\mu, \sigma^2/n)\).
SE of mean \(= \sigma/\sqrt n\).
Parameters and statistics in symbols
Write the population units as \(X_1, X_2, \ldots, X_N\) and the sample units as \(x_1, x_2, \ldots, x_n\).
A statistic is random because the sample is drawn at random: a different sample gives a different value.
That is why every statistic has a sampling distribution.
Standard errors
The standard error (S.E.) of a statistic is the standard deviation of its sampling distribution. For a
random sample (with replacement, or from a large population):
The formulas for differences need the two samples to be independent: the variance of a difference of
independent statistics is the sum of their variances.
What “exact” sampling distribution means
For large samples the central limit theorem gives an approximate normal distribution for many
statistics. For a sample from a normal population some statistics have a distribution that
can be written down exactly for every sample size, however small. These are the
exact sampling distributions:
the chi-square (\(\chi^2\)) distribution,
Student's and Fisher's \(t\) distribution,
Snedecor's \(F\) distribution,
Fisher's \(z\) distribution (\(z = \tfrac12 \log_e F\)).
Because they are exact, they are the basis of the small sample tests. Sections 3 to 5 study the first
three.
3. Chi-square (χ²) Distribution
DEFINITION
If \(Z_1, Z_2, \ldots, Z_n\) are independent \(N(0,1)\) variables, then
\[
\chi^2 \;=\; Z_1^2 + Z_2^2 + \cdots + Z_n^2
\]
follows a chi-square distribution with \(n\) degrees of freedom (df).
Setting. Let \(X_1, \ldots, X_n\) be independent with \(X_i \sim N(\mu_i, \sigma_i^2)\), and put
\(z_i = (X_i - \mu_i)/\sigma_i\). Each \(z_i \sim N(0, 1)\), the \(z_i\) are independent, and
\(\chi^2 = \sum z_i^2\).
MGF of one \(z^2\). By definition, for \(t < \tfrac12\),
\[ M_{z^2}(t) = E\big(e^{tz^2}\big) = \int_{-\infty}^{\infty} e^{tz^2}\,\frac{1}{\sqrt{2\pi}}e^{-z^2/2}\,dz = \frac{2}{\sqrt{2\pi}}\int_0^{\infty} e^{-\frac{(1-2t)}{2}z^2}\,dz , \]
using that the integrand is even.
Substitute \(y = z^2\), so \(dz = dy/(2\sqrt y)\) and \(y\) runs from 0 to \(\infty\):
\[ M_{z^2}(t) = \frac{1}{\sqrt{2\pi}}\int_0^{\infty} e^{-\frac{(1-2t)}{2}y}\,y^{\frac12 - 1}\,dy = \frac{1}{\sqrt{2\pi}}\cdot\frac{\Gamma(\frac12)}{\big(\frac{1-2t}{2}\big)^{1/2}} = (1 - 2t)^{-1/2}, \]
by the gamma integral \(\int_0^\infty e^{-ay}y^{k-1}dy = \Gamma(k)/a^k\) and \(\Gamma(\frac12) = \sqrt\pi\). The
condition \(t < \frac12\) keeps \(a = (1-2t)/2\) positive.
Multiply. The MGF of a sum of independent variables is the product of their MGFs:
\[ M_{\chi^2}(t) = \prod_{i=1}^{n} (1 - 2t)^{-1/2} = (1 - 2t)^{-n/2} . \]
Recognise it. \((1 - t/a)^{-\lambda}\) is the MGF of the gamma density
\(\frac{a^\lambda}{\Gamma(\lambda)}e^{-ax}x^{\lambda-1}\). Here \(a = \frac12\) and \(\lambda = \frac n2\). By
the uniqueness theorem of MGFs, \(\chi^2\) has the density stated above,
\(\frac{(1/2)^{n/2}}{\Gamma(n/2)}e^{-\chi^2/2}(\chi^2)^{n/2-1}\). \(\blacksquare\)
Cumulants, skewness and kurtosis
The cumulant generating function is \(K(t) = \log M(t) = -\frac n2 \log(1 - 2t)\). Use the series
\(-\log(1 - x) = x + \frac{x^2}{2} + \frac{x^3}{3} + \cdots\) with \(x = 2t\):
Skewness: \(\beta_1 = \dfrac{\mu_3^2}{\mu_2^3} = \dfrac{64n^2}{8n^3} = \dfrac{8}{n}\). Since \(\mu_3 > 0\), the
distribution is positively skewed.
Kurtosis: \(\beta_2 = \dfrac{\mu_4}{\mu_2^2} = \dfrac{12n(n+4)}{4n^2} = 3 + \dfrac{12}{n} > 3\), so it is
leptokurtic.
Both \(\beta_1 \to 0\) and \(\beta_2 \to 3\) as \(n \to \infty\), the values for a normal curve, which
agrees with the limiting result of section 7.
Characteristic function. The same gamma integral with \(it\) in place of \(t\) gives
\(\phi(t) = E(e^{it\chi^2}) = (1 - 2it)^{-n/2}\). Unlike the MGF, it exists for every real \(t\).
Mode of \(\chi^2\)
Up to a constant, \(\log f(x) = \big(\frac n2 - 1\big)\log x - \frac x2\). Setting the derivative to zero,
\[ \frac{d}{dx}\log f(x) = \frac{n/2 - 1}{x} - \frac12 = 0 \quad\Longrightarrow\quad x = n - 2 . \]
The second derivative, \(-(n/2 - 1)/x^2\), is negative for \(n > 2\), so the mode is \(n - 2\) when
\(n > 2\). For \(n = 1\) and \(n = 2\) the density decreases from \(x = 0\), so the mode is at 0.
Additive (reproductive) property, proved
Let \(\chi_1^2, \ldots, \chi_k^2\) be independent with \(n_1, \ldots, n_k\) degrees of freedom. The MGF of
their sum is the product of their MGFs:
which is the MGF of \(\chi^2\) with \(n_1 + \cdots + n_k\) degrees of freedom. By uniqueness,
\(\chi_1^2 + \cdots + \chi_k^2 \sim \chi^2_{n_1 + \cdots + n_k}\).
EXAMPLE 3 — the formulas at \(n = 10\)
For \(\chi^2_{10}\): mean \(= 10\), variance \(= 20\), \(\mu_3 = 80\), \(\mu_4 = 12 \times 10 \times 14 = 1680\),
\(\beta_1 = 8/10 = 0.8\), \(\beta_2 = 3 + 12/10 = 4.2\), mode \(= 10 - 2 = 8\). Integrating the density
numerically gives the same values.
Applications
Test for population variance.
Goodness of fit test.
Test of independence in contingency tables.
Inference about variance of a normal population: \((n-1)S^2/\sigma^2 = ns^2/\sigma^2 \sim \chi^2_{n-1}\).
Testing the homogeneity of several independent estimates of a population variance.
Testing the homogeneity of several independent estimates of a population correlation coefficient
(through Fisher's \(z\)-transformation).
EXAMPLE 1
If \(\chi^2_{10}\) is observed, find the value below which 95 % of the area lies.
From tables, \(\chi^2_{0.95, 10} = 18.31\).
EXAMPLE 2
Variance of 25 sample observations from \(N(\mu, 16)\): \(S^2 = 18\). Test statistic
\((n-1)S^2/\sigma^2 = 24 \cdot 18 / 16 = 27.0 \sim \chi^2_{24}\).
Three results from the \(\chi^2\) density
RESULT 1
If \(X \sim \chi^2_n\), then \(Y = X/2\) has the gamma density \(\dfrac{e^{-y}y^{n/2-1}}{\Gamma(n/2)}\),
\(y \ge 0\): a gamma variate with parameter \(n/2\).
Proof. \(x = 2y\), so \(|J| = |dx/dy| = 2\). Then
\(f_Y(y) = f_X(2y)\cdot 2 = \dfrac{e^{-y}(2y)^{n/2-1}\cdot 2}{2^{n/2}\Gamma(n/2)} = \dfrac{e^{-y}y^{n/2-1}}{\Gamma(n/2)}\),
since \(2^{n/2-1}\cdot 2 = 2^{n/2}\). \(\blacksquare\)
RESULT 2
If \(X \sim \chi^2_{n_1}\) and \(Y \sim \chi^2_{n_2}\) are independent, then \(U = X/Y\) is a beta variate of
the second kind with parameters \(n_1/2\) and \(n_2/2\):
Split the sum of squares. Add and subtract \(\bar x\):
\[ \sum (x_i - \mu)^2 = \sum (x_i - \bar x)^2 + n(\bar x - \mu)^2 + 2(\bar x - \mu)\sum (x_i - \bar x) , \]
and the last sum is 0, because the deviations from the mean add to zero.
Divide by \(\sigma^2\) and name the three pieces:
\[ \underbrace{\sum \Big(\frac{x_i - \mu}{\sigma}\Big)^2}_{V} = \underbrace{\frac{ns^2}{\sigma^2}}_{W} + \underbrace{\Big(\frac{\bar x - \mu}{\sigma/\sqrt n}\Big)^2}_{Z^2} . \]
Identify \(V\) and \(Z^2\). Each \((x_i - \mu)/\sigma \sim N(0, 1)\), independently, so \(V \sim \chi^2_n\).
Since \(\bar x \sim N(\mu, \sigma^2/n)\), the ratio \((\bar x - \mu)/(\sigma/\sqrt n) \sim N(0, 1)\), so
\(Z^2 \sim \chi^2_1\).
Use independence. For a normal sample, \(\bar x\) and \(s^2\) are independent (proved by the Helmert
transformation in Distribution Theory, Unit 3), so \(W\) and \(Z^2\) are independent and
\(M_V(t) = M_W(t)\,M_{Z^2}(t)\).
Solve for \(M_W\):
\[ M_W(t) = \frac{(1 - 2t)^{-n/2}}{(1 - 2t)^{-1/2}} = (1 - 2t)^{-(n-1)/2}, \]
the MGF of \(\chi^2_{n-1}\). By uniqueness, \(W = ns^2/\sigma^2 \sim \chi^2_{n-1}\). One degree of freedom
is lost because the deviations \(x_i - \bar x\) satisfy one linear restriction. \(\blacksquare\)
Density of \(s^2\). Put \(y = ns^2/\sigma^2\), so \(|dy/ds^2| = n/\sigma^2\), and use the
\(\chi^2_{n-1}\) density of \(y\):
So \(s^2\) slightly underestimates \(\sigma^2\) (which is why \(S^2\), with divisor \(n - 1\), is used for
estimation), and for large \(n\), \(E(s^2) \approx \sigma^2\) and \(\text{Var}(s^2) \approx 2\sigma^4/n\).
EXAMPLE 4 — samples of 10 from \(N(\mu, 4)\)
\(E(s^2) = \frac{9}{10}\times 4 = 3.6\) and \(\text{Var}(s^2) = \frac{2 \times 9}{100}\times 16 = 2.88\). A
simulation of 200,000 samples of 10 from \(N(0, 4)\) gives a mean of 3.598 and a variance of 2.865 for
\(s^2\), in agreement.
4. Student's t-Distribution
DEFINITION
If \(Z \sim N(0,1)\) and \(\chi^2 \sim \chi^2_n\) are independent, then
\[
t \;=\; \dfrac{Z}{\sqrt{\chi^2 / n}} \;\sim\; t_n
\]
follows the Student's t-distribution with \(n\) degrees of freedom.
Skewness 0; kurtosis \(\beta_2 > 3\) for \(n > 4\).
Student's \(t\) and Fisher's \(t\)
Fisher's \(t\) is the definition above: \(t = \xi\big/\sqrt{\chi^2/n} \sim t_n\), with
\(\xi \sim N(0, 1)\) independent of \(\chi^2 \sim \chi^2_n\).
Student's \(t\) (W. S. Gosset, 1908) is the statistic for a sample of \(n\) from
\(N(\mu, \sigma^2)\):
\[ t = \frac{\bar x - \mu}{S/\sqrt n} \sim t_{n-1}, \qquad S^2 = \frac{1}{n-1}\sum (x_i - \bar x)^2 . \]
Student's \(t\) is a particular case of Fisher's. Take \(\xi = \dfrac{\bar x - \mu}{\sigma/\sqrt n} \sim N(0,1)\)
and \(\chi^2 = \dfrac{ns^2}{\sigma^2} \sim \chi^2_{n-1}\), which are independent. Then
\[ \frac{\xi}{\sqrt{\chi^2/(n-1)}} = \frac{(\bar x - \mu)\sqrt n/\sigma}{\sqrt{ns^2/(\sigma^2(n-1))}} = \frac{\bar x - \mu}{\sqrt{ns^2/(n(n-1))}} = \frac{\bar x - \mu}{S/\sqrt n}, \]
using \(ns^2 = (n-1)S^2\). The unknown \(\sigma\) cancels, which is what makes \(t\) usable in practice.
So Student's \(t\) has Fisher's \(t\) distribution with \(n - 1\) degrees of freedom.
Deriving the \(t\) density
Start from the independent densities \(f(\xi) = \frac{1}{\sqrt{2\pi}}e^{-\xi^2/2}\) and
\(f(u) = \frac{1}{2^{n/2}\Gamma(n/2)}e^{-u/2}u^{n/2-1}\), where \(u = \chi^2\).
Transformation. \(t = \xi/\sqrt{u/n}\) and \(u = u\), so \(\xi = t\sqrt{u/n}\) and
\(J = \begin{vmatrix} \sqrt{u/n} & \frac{t}{2\sqrt{nu}} \\ 0 & 1 \end{vmatrix} = \sqrt{u/n}\).
Joint density. \(f(t, u) = f\big(t\sqrt{u/n}\big)\,f(u)\,\sqrt{u/n}\); collecting the powers of
\(u\) and the exponentials,
\[ f(t, u) = \frac{1}{\sqrt{2\pi n}\;2^{n/2}\,\Gamma(\frac n2)}\;u^{\frac{n+1}{2}-1}\,e^{-\frac u2\left(1 + \frac{t^2}{n}\right)} . \]
Integrate out \(u\) with the gamma integral, \(a = \frac12(1 + t^2/n)\), \(k = \frac{n+1}{2}\):
\[ f(t) = \frac{\Gamma(\frac{n+1}{2})\,2^{(n+1)/2}}{\sqrt{2\pi n}\;2^{n/2}\,\Gamma(\frac n2)}\Big(1 + \frac{t^2}{n}\Big)^{-\frac{n+1}{2}} = \frac{\Gamma(\frac{n+1}{2})}{\sqrt{n\pi}\;\Gamma(\frac n2)}\Big(1 + \frac{t^2}{n}\Big)^{-\frac{n+1}{2}} . \]
Write it with the beta function. \(B(\frac12, \frac n2) = \Gamma(\frac12)\Gamma(\frac n2)/\Gamma(\frac{n+1}{2})\) and
\(\Gamma(\frac12) = \sqrt\pi\), so the constant is \(1/(\sqrt n\,B(\frac12, \frac n2))\), as stated above.
\(\blacksquare\)
Another route. \(t^2/n = \xi^2/\chi^2\) is a ratio of independent \(\chi^2_1\) and \(\chi^2_n\) variates,
so by Result 2 of section 3 it is a beta variate of the second kind, \(\beta_2(\frac12, \frac n2)\). Changing
variable from \(t^2/n\) to \(t\) needs care: each value of \(t^2\) comes from two values, \(t\) and \(-t\), and the
density is shared equally between them. That halving cancels the factor 2 in \(|d(t^2/n)/dt| = 2|t|/n\) and
gives the same density.
Moments of \(t\)
Odd moments. \(t^{2r+1}f(t)\) is an odd function, so \(\mu'_{2r+1} = 0\) whenever the moment
exists (for \(2r + 1 < n\)). With \(r = 0\): the mean is 0 for \(n > 1\).
Even moments. Using the even integrand and the substitution \(1 + t^2/n = 1/y\), which turns the
integral into a beta integral,
\[ \mu'_{2r} = \frac{n^r}{B(\frac12, \frac n2)}\int_0^1 y^{\frac n2 - r - 1}(1-y)^{r - \frac12}\,dy = n^r\,\frac{\Gamma(\frac n2 - r)\,\Gamma(r + \frac12)}{\Gamma(\frac12)\,\Gamma(\frac n2)}, \qquad n > 2r . \]
\(r = 1\): using \(\Gamma(\frac32) = \frac12\Gamma(\frac12)\) and \(\Gamma(\frac n2) = (\frac n2 - 1)\Gamma(\frac n2 - 1)\),
\(\mu_2 = \dfrac{n}{n-2}\) for \(n > 2\).
\(r = 2\): \(\mu_4 = \dfrac{3n^2}{(n-2)(n-4)}\) for \(n > 4\).
\(\beta_1 = 0\) (symmetric), and \(\beta_2 = \dfrac{\mu_4}{\mu_2^2} = \dfrac{3(n-2)}{n-4} > 3\) for
\(n > 4\): leptokurtic.
No MGF. Only the moments of order less than \(n\) exist, and \(E(e^{st})\) is infinite for every
\(s \ne 0\), because \(e^{st}\) grows faster than any power of \(t\) while the density falls only like
\(|t|^{-(n+1)}\). So the \(t\) distribution has no moment generating function.
Inference about mean when SD is unknown (single mean, paired samples, two means with equal variances).
Confidence intervals: \(\bar x \pm t_{\alpha/2, n-1}\, S/\sqrt n\).
Tests of significance of a regression coefficient.
Tests of significance of an observed sample correlation coefficient and of a partial correlation
coefficient. (A multiple correlation coefficient is tested with \(F\), section 5.)
EXAMPLE 1
Sample of 16 from \(N(\mu, \sigma^2)\): \(\bar x = 50,\; s = 8\). Test \(H_0: \mu = 45\) at 5 %.
Reciprocal property: if \(F \sim F_{n_1,n_2}\) then \(1/F \sim F_{n_2, n_1}\).
\(F_{1, n_2} = t_{n_2}^2\) (the square of t with \(n_2\) df).
Deriving the \(F\) density
Start from independent \(X \sim \chi^2_{n_1}\) and \(Y \sim \chi^2_{n_2}\), and
\(F = \dfrac{X/n_1}{Y/n_2}\).
Transformation. \(F\) and \(u = y\), so \(x = \frac{n_1}{n_2}Fu\), \(y = u\), and
\(J = \begin{vmatrix} \frac{n_1}{n_2}u & \frac{n_1}{n_2}F \\ 0 & 1 \end{vmatrix} = \frac{n_1}{n_2}u\).
Joint density. \(f(F, u) = f_X\big(\frac{n_1}{n_2}Fu\big)\,f_Y(u)\,\frac{n_1}{n_2}u\), which collects to
\[ f(F, u) = \frac{(n_1/n_2)^{n_1/2}\,F^{n_1/2-1}}{2^{(n_1+n_2)/2}\,\Gamma(\frac{n_1}{2})\Gamma(\frac{n_2}{2})}\;u^{\frac{n_1+n_2}{2}-1}\,e^{-\frac u2\left(1 + \frac{n_1}{n_2}F\right)} . \]
Integrate out \(u\) with the gamma integral, \(a = \frac12(1 + \frac{n_1}{n_2}F)\),
\(k = \frac{n_1+n_2}{2}\). The powers of 2 cancel and the gammas form \(B(\frac{n_1}{2}, \frac{n_2}{2})\),
leaving the density stated above. \(\blacksquare\)
Moments of \(F\)
Substitute \(y = \frac{n_1}{n_2}F\) in \(\mu'_r = \int_0^\infty F^r f(F)\,dF\). The integral becomes
\(\big(\frac{n_2}{n_1}\big)^r \int_0^\infty \frac{y^{n_1/2 + r - 1}}{(1+y)^{(n_1+n_2)/2}}\,dy\big/B(\frac{n_1}{2}, \frac{n_2}{2})\),
a beta integral of the second kind:
Up to a constant, \(\log f = \big(\frac{n_1}{2} - 1\big)\log F - \frac{n_1+n_2}{2}\log\big(1 + \frac{n_1}{n_2}F\big)\).
Setting the derivative to zero and multiplying through by \(2F(1 + n_1F/n_2)\):
Both factors \(\frac{n_1-2}{n_1}\) and \(\frac{n_2}{n_2+2}\) are below 1, so the mode is always less than 1,
while the mean \(\frac{n_2}{n_2-2}\) is above 1: the distribution is positively skewed.
Reciprocal property, proved
If \(F = \dfrac{X/n_1}{Y/n_2} \sim F_{n_1, n_2}\), then \(\dfrac1F = \dfrac{Y/n_2}{X/n_1}\) is again a ratio of
independent \(\chi^2\) variates, each divided by its degrees of freedom, now with \(Y\) on top. So
\(1/F \sim F_{n_2, n_1}\). This is why tables give only upper percentage points: the lower point is
\(F_{1-\alpha}(n_1, n_2) = 1/F_{\alpha}(n_2, n_1)\).
EXAMPLE 3 — the formulas at \((n_1, n_2) = (5, 10)\)
Mean \(= 10/8 = 1.25\); \(\mu'_2 = \dfrac{100 \times 7}{5 \times 8 \times 6} = \dfrac{35}{12} = 2.917\);
variance \(= \dfrac{2 \times 100 \times 13}{5 \times 64 \times 6} = \dfrac{65}{48} = 1.354\); mode
\(= \dfrac{10 \times 3}{5 \times 12} = 0.5\). Numerical integration of the density gives the same values.
Applications
Test of equality of two population variances.
One-way and two-way ANOVA.
Test of overall significance in regression.
Testing the linearity of regression, and the significance of an observed multiple correlation
coefficient or correlation ratio.
EXAMPLE 1
For two samples with \(s_1^2 = 16\) (\(n_1 = 11\)) and \(s_2^2 = 9\) (\(n_2 = 13\)): \(F = 16/9 = 1.78\) on df (10, 12).
Critical value at 5 % is \(F_{0.05} = 2.75\). Since \(1.78 < 2.75\), variances are not significantly different.
EXAMPLE 2
For df (5, 10), \(F_{0.05} = 3.33\). The probability that \(F > 3.33\) is 0.05.
Fig 5.2 — The three exact sampling distributions, drawn from their densities.
\(\chi^2\) is right-skewed with mode \(n - 2\) and moves right as \(n\) grows; \(t\) is symmetric with heavier tails
than \(N(0, 1)\) and approaches it as \(n\) grows; \(F\) is right-skewed with its mode below 1.
6. Relationships among Sampling Distributions
Standard Normal: \(Z = (X - \mu)/\sigma\).
\(\chi^2_n = \sum Z_i^2\) (sum of \(n\) squared standard normals).
For large df: \(t \to Z\); \((\chi^2 - n)/\sqrt{2n} \to Z\).
Proof that \(t^2 \sim F_{1, n}\)
Write \(t = \xi/\sqrt{\chi^2/n}\). Squaring,
\[ t^2 = \frac{\xi^2/1}{\chi^2/n} . \]
The numerator \(\xi^2\) is the square of a standard normal variate, so \(\xi^2 \sim \chi^2_1\); it is independent
of \(\chi^2 \sim \chi^2_n\). So \(t^2\) is a ratio of independent \(\chi^2\) variates, each divided by its degrees of
freedom: \(t^2 \sim F_{1, n}\). In tables, \(t_{\alpha/2}(n)^2 = F_{\alpha}(1, n)\); for example
\(2.228^2 = 4.96 = F_{0.05}(1, 10)\).
7. Limiting Cases of the Sampling Distributions
STUDENT'S t → NORMAL
As the degrees of freedom \(n \to \infty\), Student's \(t\) tends to the standard normal:
Reason: \(\chi^2_n / n\) is the mean of \(n\) i.i.d. terms \(Z_i^2\) (each with
mean 1), so by the law of large numbers \(\chi^2_n / n \to 1\); the denominator \(\to 1\) and hence
\(t_n \to Z\). The heavy tails of \(t\) thin out to match the normal (in practice \(t_n \approx Z\)
for \(n \ge 30\)).
Reason: \(\chi^2_n = \sum_{i=1}^{n} Z_i^2\) is a sum of \(n\) i.i.d. variables,
each with mean 1 and variance 2, so the Central Limit Theorem applies. (Fisher's sharper
approximation: \(\sqrt{2\chi^2_n} \approx N(\sqrt{2n - 1},\, 1)\).)
F → CHI-SQUARE
For fixed numerator degrees of freedom \(n_1\), as the denominator df \(n_2 \to \infty\),
Reason: \(F_{n_1, n_2} = \dfrac{\chi^2_{n_1}/n_1}{\chi^2_{n_2}/n_2}\); as
\(n_2 \to \infty\), \(\chi^2_{n_2}/n_2 \to 1\), so \(F_{n_1, n_2} \to \chi^2_{n_1}/n_1\). If in
addition \(n_1 \to \infty\), then \(F_{n_1, n_2} \to 1\).
The limits proved from the generating function and the densities
\(\chi^2 \to\) normal, through the CGF. Let \(z = (\chi^2 - n)/\sqrt{2n}\). Then
\(M_z(t) = e^{-tn/\sqrt{2n}}M_{\chi^2}(t/\sqrt{2n}) = e^{-t\sqrt{n/2}}\big(1 - t\sqrt{2/n}\big)^{-n/2}\), and
The first two terms cancel, and every term after \(t^2/2\) has a power of \(\sqrt n\) in its denominator, so
\(K_z(t) \to t^2/2\), the CGF of \(N(0, 1)\). By the continuity theorem for MGFs, \(z\) tends in distribution
to \(N(0, 1)\).
LEMMA
For fixed \(K\), \(\dfrac{\Gamma(n + K)}{\Gamma(n)\,n^K} \to 1\) as \(n \to \infty\); that is,
\(\Gamma(n + K)/\Gamma(n)\) grows like \(n^K\).
using \((1 + K/n)^n \to e^K\) and \((1 + K/n)^{c} \to 1\) for a fixed exponent \(c\). (With an exponent
that grows with \(n\) the second limit fails: \((1 + 1/n)^n \to e\), not 1.) \(\blacksquare\)
\(t \to\) normal, through the density. In
\(f(t) = \dfrac{\Gamma(\frac{n+1}{2})}{\sqrt{n\pi}\,\Gamma(\frac n2)}\big(1 + \frac{t^2}{n}\big)^{-n/2}\big(1 + \frac{t^2}{n}\big)^{-1/2}\):
by the lemma with \(n/2\) and \(K = \frac12\), \(\Gamma(\frac n2 + \frac12)/\Gamma(\frac n2) \approx (\frac n2)^{1/2}\), so the
constant tends to \((\frac n2)^{1/2}/\sqrt{n\pi} = 1/\sqrt{2\pi}\);
So \(f(t) \to \frac{1}{\sqrt{2\pi}}e^{-t^2/2}\), the standard normal density.
\(n_1F \to \chi^2_{n_1}\), through the density. In the \(F\) density, by the lemma with \(n_2/2\) and
\(K = n_1/2\), \(\Gamma(\frac{n_1+n_2}{2})/\Gamma(\frac{n_2}{2}) \approx (\frac{n_2}{2})^{n_1/2}\), so
\((n_1/n_2)^{n_1/2}(n_2/2)^{n_1/2} = (n_1/2)^{n_1/2}\). Also
\(\big(1 + \frac{n_1F}{n_2}\big)^{-n_2/2} \to e^{-n_1F/2}\) and \(\big(1 + \frac{n_1F}{n_2}\big)^{-n_1/2} \to 1\). Hence
Now put \(x = n_1F\), with \(|dF/dx| = 1/n_1\): the density of \(x\) is
\(\frac{(1/2)^{n_1/2}}{\Gamma(n_1/2)}e^{-x/2}x^{n_1/2-1}\), the \(\chi^2_{n_1}\) density.
For a normal sample \(ns^2/\sigma^2 = (n-1)S^2/\sigma^2 \sim \chi^2_{n-1}\), so \(E(s^2) = (1 - \frac1n)\sigma^2\) and \(\text{Var}(s^2) = 2(n-1)\sigma^4/n^2\).
Student's \(t = (\bar x - \mu)/(S/\sqrt n)\) is Fisher's \(t\) with \(n - 1\) df; \(\mu_2 = n/(n-2)\), \(\beta_2 = 3(n-2)/(n-4)\); no MGF.