Skip to the content

Topics Covered

Chi-Square Derivation Sample Mean & Variance Helmert Transformation Independence of Xbar and S² Student's t F Distribution Non-Central Forms
On this page
  1. 1. The Chi-Square Distribution
  2. 2. The Distribution of the Sample Mean and Variance
  3. 3. Student's t Distribution
  4. 4. The F Distribution
  5. 5. The Non-Central Distributions
  6. Key Take-aways
Where this unit starts. Continuous Distributions, Unit 5 introduces the standard normal and names the chi-square, \(t\) and \(F\) sampling distributions with their tables, and Inferential Statistics, Unit 4 uses them for small-sample tests. Neither derives the densities. That derivation is the work of this unit, together with the result those tests silently depend on: that \(\bar X\) and \(S^{2}\) are independent when the sample is normal.

1. The Chi-Square Distribution

DERIVING THE DENSITY OF THE SQUARE OF ONE STANDARD NORMAL

Statement. If \(Z \sim N(0,1)\) then \(Y = Z^{2} \sim \chi^{2}_{1}\), with density

\[ f_Y(y) = \frac{1}{\sqrt{2\pi y}} e^{-y/2}, \qquad y > 0. \]

Step 1 — go through the distribution function, because the map is not one-to-one. \(z \mapsto z^{2}\) sends both \(+z\) and \(-z\) to the same \(y\), so the change-of-variable formula cannot be applied directly. For \(y > 0\),

\[ F_Y(y) = P\left(Z^{2} \le y\right) = P\left(-\sqrt{y} \le Z \le \sqrt{y}\right) = \Phi(\sqrt{y}) - \Phi(-\sqrt{y}). \]

Step 2 — use symmetry. \(\Phi(-a) = 1 - \Phi(a)\), so

\[ F_Y(y) = \Phi(\sqrt{y}) - \big[1 - \Phi(\sqrt{y})\big] = 2\Phi(\sqrt{y}) - 1. \]

Step 3 — differentiate, by the chain rule. With \(\dfrac{d}{dy}\sqrt{y} = \dfrac{1}{2\sqrt{y}}\) and \(\Phi' = \varphi\),

\[ f_Y(y) = 2 \varphi(\sqrt{y}) \cdot \frac{1}{2\sqrt{y}} = \frac{\varphi(\sqrt{y})}{\sqrt{y}}. \]

Step 4 — substitute the standard normal density. \(\varphi(t) = \dfrac{1}{\sqrt{2\pi}} e^{-t^{2}/2}\), so \(\varphi(\sqrt{y})\) has \(t^{2} = y\) and

\[ f_Y(y) = \frac{1}{\sqrt{2\pi y}} e^{-y/2}, \qquad y > 0. \quad \blacksquare \]

Step 5 — recognise it. Compare with the Gamma(shape \(\alpha\), rate \(\beta\)) density \(\dfrac{\beta^{\alpha}}{\Gamma(\alpha)} y^{\alpha-1} e^{-\beta y}\). Matching exponents gives \(\alpha - 1 = -\tfrac12\), so \(\alpha = \tfrac12\), and \(\beta = \tfrac12\). The constant agrees because \(\Gamma(\tfrac12) = \sqrt{\pi}\). So \(\chi^{2}_{1} = \text{Gamma}(\tfrac12, \tfrac12)\).

FROM ONE DEGREE OF FREEDOM TO n, BY THE MGF

Statement. If \(Z_1, \ldots, Z_n\) are independent \(N(0,1)\) then \(\sum_{i} Z_i^{2} \sim \chi^{2}_{n}\), with density

\[ f(y) = \frac{1}{2^{n/2}\,\Gamma(n/2)} \, y^{n/2 - 1} e^{-y/2}, \qquad y > 0. \]

Step 1 — the MGF of one square. From Step 5 above, \(Z^{2} \sim \text{Gamma}(\tfrac12, \tfrac12)\), whose MGF is

\[ M_{Z^{2}}(t) = \left(\frac{1/2}{1/2 - t}\right)^{1/2} = (1 - 2t)^{-1/2}, \qquad t < \tfrac12. \]

Step 2 — multiply, using independence. The MGF of a sum of independent variables is the product of their MGFs, so

\[ M_{\sum Z_i^{2}}(t) = \left[(1-2t)^{-1/2}\right]^{n} = (1 - 2t)^{-n/2}. \]

Step 3 — identify. \((1-2t)^{-n/2}\) is the MGF of Gamma\((n/2, 1/2)\), whose density is the one displayed. By the uniqueness of the MGF the distribution is determined. \(\blacksquare\)

Additivity follows at once. If \(U \sim \chi^{2}_{m}\) and \(V \sim \chi^{2}_{n}\) independently, then \(M_{U+V}(t) = (1-2t)^{-m/2}(1-2t)^{-n/2} = (1-2t)^{-(m+n)/2}\), so \(U + V \sim \chi^{2}_{m+n}\).

PROPERTIES \[ E\left(\chi^{2}_{n}\right) = n, \qquad \operatorname{Var}\left(\chi^{2}_{n}\right) = 2n, \qquad \text{mode} = n - 2 \ (n \ge 2), \qquad \beta_1 = \sqrt{8/n}. \]

The skewness \(\sqrt{8/n} \to 0\), so the chi-square becomes more symmetric as the degrees of freedom grow — visible in the figure below, where \(\chi^{2}_{2}\) is a decreasing exponential and \(\chi^{2}_{9}\) is already nearly bell-shaped.

2. The Distribution of the Sample Mean and Variance

THE THEOREM EVERY SMALL-SAMPLE TEST DEPENDS ON

Statement. Let \(X_1, \ldots, X_n\) be i.i.d. \(N(\mu, \sigma^{2})\), and

\[ \bar X = \frac{1}{n}\sum_{i=1}^{n} X_i, \qquad S^{2} = \frac{1}{n-1}\sum_{i=1}^{n} \left(X_i - \bar X\right)^{2}. \]

Then (i) \(\bar X \sim N(\mu, \sigma^{2}/n)\); (ii) \(\dfrac{(n-1)S^{2}}{\sigma^{2}} \sim \chi^{2}_{n-1}\); and (iii) \(\bar X\) and \(S^{2}\) are independent.

Proof of (i). A linear combination of independent normals is normal, with \(E(\bar X) = \mu\) and \(\operatorname{Var}(\bar X) = \dfrac{1}{n^{2}} \cdot n\sigma^{2} = \dfrac{\sigma^{2}}{n}\).

Proof of (ii) and (iii), by the Helmert transformation.

Step 1. Standardise: put \(Z_i = (X_i - \mu)/\sigma\), so the \(Z_i\) are independent \(N(0,1)\), and the claims become statements about \(\sum_i (Z_i - \bar Z)^{2}\) and \(\bar Z\).

Step 2. Apply the orthogonal Helmert matrix \(H\), whose first row is \((1/\sqrt{n}, \ldots, 1/\sqrt{n})\) and whose \(k\)-th row (\(k = 2, \ldots, n\)) is

\[ \frac{1}{\sqrt{k(k-1)}}\big(\underbrace{1, 1, \ldots, 1}_{k-1}, \; -(k-1), \; 0, \ldots, 0\big). \]

Each row has unit length and any two rows are orthogonal — which is exactly what makes \(H\) orthogonal, \(H'H = I\). Set \(Y = HZ\).

Step 3. An orthogonal transformation of independent standard normals gives independent standard normals. The joint density depends on \(Z\) only through \(Z'Z\), which is unchanged since \(Y'Y = Z'H'HZ = Z'Z\), and the Jacobian is \(|\det H| = 1\). So \(Y_1, \ldots, Y_n\) are independent \(N(0,1)\).

Step 4. Read off the first component: \(Y_1 = \dfrac{1}{\sqrt{n}}\sum_i Z_i = \sqrt{n}\,\bar Z\).

Step 5. Use the length-preserving property on the remaining components:

\[ \sum_{k=2}^{n} Y_k^{2} = \sum_{i=1}^{n} Z_i^{2} - Y_1^{2} = \sum_{i=1}^{n} Z_i^{2} - n\bar Z^{2} = \sum_{i=1}^{n}\left(Z_i - \bar Z\right)^{2}, \]

the last step being the usual identity \(\sum (z_i - \bar z)^2 = \sum z_i^2 - n\bar z^2\).

Step 6. Conclude. The left-hand side of Step 5 is a sum of \(n-1\) independent squared standard normals, so it is \(\chi^{2}_{n-1}\); and since

\[ \frac{(n-1)S^{2}}{\sigma^{2}} = \sum_{i=1}^{n}\left(Z_i - \bar Z\right)^{2} = \sum_{k=2}^{n} Y_k^{2}, \]

claim (ii) holds. Claim (iii) holds because \(\bar X\) is a function of \(Y_1\) alone while \(S^{2}\) is a function of \(Y_2, \ldots, Y_n\) alone, and those are independent by Step 3. \(\blacksquare\)

Why (iii) matters. The \(t\) statistic puts \(\bar X - \mu\) over \(S/\sqrt{n}\). Its distribution is derived in section 3 from a normal divided by an independent chi-square. Without independence there is no \(t\) distribution, and every small-sample test on a normal mean loses its justification.

EXAMPLE 3.1 — HOW VARIABLE IS THE SAMPLE VARIANCE?

Given. A random sample of \(n = 10\) from \(N(\mu, \sigma^{2})\).

Asked. Find \(E(S^{2})\), \(\operatorname{Var}(S^{2})\), and \(P(S^{2} > 1.8\sigma^{2})\).

Step 1 — name the pivotal quantity. By part (ii), \(W = \dfrac{(n-1)S^{2}}{\sigma^{2}} = \dfrac{9 S^{2}}{\sigma^{2}} \sim \chi^{2}_{9}\).

Step 2 — the mean of \(S^{2}\). \(E(W) = 9\), so

\[ E\left(\frac{9S^{2}}{\sigma^{2}}\right) = 9 \;\Rightarrow\; E(S^{2}) = \sigma^{2}. \]

The divisor \(n-1\) is exactly what makes \(S^{2}\) unbiased.

Step 3 — the variance of \(S^{2}\). \(\operatorname{Var}(W) = 2 \times 9 = 18\), and \(S^{2} = \dfrac{\sigma^{2}}{9} W\), so

\[ \operatorname{Var}(S^{2}) = \left(\frac{\sigma^{2}}{9}\right)^{2} \times 18 = \frac{18\sigma^{4}}{81} = \frac{2\sigma^{4}}{9} = 0.222222\,\sigma^{4}. \]

The standard deviation of \(S^{2}\) is \(\sqrt{0.222222}\,\sigma^{2} = 0.4714\sigma^{2}\) — that is \(47\%\) of what it is estimating.

Step 4 — the tail probability. Translate the event into \(W\):

\[ P\left(S^{2} > 1.8\sigma^{2}\right) = P\left(\frac{9S^{2}}{\sigma^{2}} > 9 \times 1.8\right) = P\left(\chi^{2}_{9} > 16.2\right). \]

Step 5 — evaluate.

\[ P\left(\chi^{2}_{9} > 16.2\right) = 0.062821. \]

For orientation, the tabulated \(5\%\) point of \(\chi^{2}_{9}\) is \(16.919\), which confirms the value must be a little above \(0.05\), and it is.

Interpretation. With only ten observations there is a \(6.3\%\) chance that the sample variance overstates the true variance by \(80\%\) or more. A single small-sample variance is a weak estimate, which is why a confidence interval for \(\sigma^{2}\) built from the \(\chi^{2}_{9}\) percentage points \(3.325\) and \(16.919\) is so wide.

Chi-square densities, and the tail of Example 3.1 16.2 tail area 0.0628 df = 2 df = 5 df = 9 0 5 10 20 25 0.05 0.10 the skew falls as df rises: skewness is √(8/df), so 2.00 at df = 2 and 0.94 at df = 9 mode sits at df − 2, so at 0, 3 and 7 for the three curves
Fig 3.1 — Densities and shaded tail computed from the chi-square density function, not sketched.

3. Student's t Distribution

DERIVING THE DENSITY

Definition. If \(Z \sim N(0,1)\) and \(V \sim \chi^{2}_{n}\) are independent, then

\[ T = \frac{Z}{\sqrt{V/n}} \sim t_n. \]

Step 1 — write the joint density. By independence it is the product

\[ f_{Z,V}(z,v) = \frac{1}{\sqrt{2\pi}} e^{-z^{2}/2} \cdot \frac{1}{2^{n/2}\Gamma(n/2)} v^{n/2-1} e^{-v/2}. \]

Step 2 — transform. Put \(t = z/\sqrt{v/n}\) and keep \(u = v\). The inverse is \(z = t\sqrt{u/n}\), \(v = u\), and the Jacobian is

\[ J = \begin{vmatrix} \sqrt{u/n} & \ast \\ 0 & 1 \end{vmatrix} = \sqrt{u/n}, \]

the lower-left entry being \(\partial v/\partial t = 0\), so the determinant is the product of the diagonal.

Step 3 — substitute.

\[ f_{T,U}(t,u) = \frac{1}{\sqrt{2\pi}\,2^{n/2}\Gamma(n/2)} \, u^{n/2-1} \exp\left\{-\frac{u}{2}\left(1 + \frac{t^{2}}{n}\right)\right\}\sqrt{\frac{u}{n}}. \]

Step 4 — integrate \(u\) out. Collecting powers, \(u^{n/2-1}\sqrt{u} = u^{(n+1)/2 - 1}\), so the \(u\)-integral is a gamma integral with shape \((n+1)/2\) and rate \(\tfrac12\left(1 + t^{2}/n\right)\):

\[ \int_{0}^{\infty} u^{(n+1)/2-1} e^{-u\beta} du = \frac{\Gamma\!\left(\frac{n+1}{2}\right)}{\beta^{(n+1)/2}}, \qquad \beta = \tfrac12\left(1 + \tfrac{t^{2}}{n}\right). \]

Step 5 — collect the constants. The powers of 2 cancel against \(2^{n/2}\) and \(\beta^{(n+1)/2}\), leaving

\[ f_T(t) = \frac{\Gamma\!\left(\frac{n+1}{2}\right)}{\sqrt{n\pi}\,\Gamma\!\left(\frac{n}{2}\right)} \left(1 + \frac{t^{2}}{n}\right)^{-(n+1)/2}, \qquad t \in \mathbb{R}. \quad \blacksquare \]
PROPERTIES OF t \[ \begin{aligned} &\text{symmetric about } 0; \quad E(T) = 0 \text{ for } n > 1, \text{ undefined for } n = 1 \\ &\operatorname{Var}(T) = \frac{n}{n-2} \text{ for } n > 2, \text{ infinite for } n \le 2 \\ &E|T|^{r} \text{ exists only for } r < n \\ &n = 1 \Rightarrow t_1 \text{ is the standard Cauchy}; \qquad n \to \infty \Rightarrow t_n \to N(0,1) \\ &T^{2} \sim F_{1,n} \end{aligned} \]

At \(n = 1\) the \(t\) density is \(\dfrac{1}{\pi(1+t^{2})}\), exactly the Cauchy of Unit 1, section 6 — which is why \(t_1\) has no mean. The limit to \(N(0,1)\) follows from \(\left(1 + t^{2}/n\right)^{-(n+1)/2} \to e^{-t^{2}/2}\).

EXAMPLE 3.2 — A SMALL-SAMPLE TEST ON A MEAN

Given. \(n = 16\) observations from a normal population, \(\bar x = 102.5\), \(s = 5\). Test \(H_0: \mu = 100\) against \(H_1: \mu \ne 100\).

Asked. The statistic, its null distribution, and the significance probability.

Step 1 — why a \(t\) and not a \(z\). \(\sigma\) is unknown and replaced by \(s\). By the theorem of section 2, \(\bar X\) and \(S^{2}\) are independent, so

\[ T = \frac{\bar X - \mu}{S/\sqrt{n}} = \frac{(\bar X - \mu)/(\sigma/\sqrt{n})}{\sqrt{\frac{(n-1)S^{2}/\sigma^{2}}{n-1}}} = \frac{Z}{\sqrt{V/(n-1)}} \sim t_{n-1}, \]

which is exactly the structure required by the definition above, with \(n - 1 = 15\) degrees of freedom.

Step 2 — the standard error.

\[ \frac{s}{\sqrt{n}} = \frac{5}{\sqrt{16}} = \frac{5}{4} = 1.25. \]

Step 3 — the statistic.

\[ t = \frac{102.5 - 100}{1.25} = \frac{2.5}{1.25} = 2.000. \]

Step 4 — the significance probability.

\[ P\left(|t_{15}| > 2.000\right) = 0.063945. \]

Step 5 — decide. \(0.063945 > 0.05\), so \(H_0\) is not rejected at the \(5\%\) level. The two-sided \(5\%\) point of \(t_{15}\) is \(2.131\), and \(2.000 < 2.131\), which agrees.

Interpretation. Had \(\sigma = 5\) been known, the statistic would be the same \(2.000\) but referred to \(N(0,1)\), giving \(P(|Z| > 2) = 0.045500\) and rejection at \(5\%\). The opposite conclusion, from the same data, purely because estimating \(\sigma\) from 16 observations costs precision. That cost is what the heavier tails of the \(t\) distribution encode.

4. The F Distribution

DEFINITION AND DENSITY

If \(U \sim \chi^{2}_{m}\) and \(V \sim \chi^{2}_{n}\) are independent, then

\[ F = \frac{U/m}{V/n} \sim F_{m,n}, \qquad f(x) = \frac{\left(\frac{m}{n}\right)^{m/2}}{B\!\left(\frac{m}{2}, \frac{n}{2}\right)} \cdot \frac{x^{m/2-1}}{\left(1 + \frac{m}{n}x\right)^{(m+n)/2}}, \quad x > 0, \]

derived by the same two-step route as the \(t\): form the joint density, transform to \((F, V)\) with Jacobian, and integrate \(V\) out against a gamma integral.

PROPERTIES AND THE THREE RELATIONSHIPS \[ \begin{aligned} E(F) &= \frac{n}{n-2} \ (n > 2), \qquad \operatorname{Var}(F) = \frac{2n^{2}(m+n-2)}{m(n-2)^{2}(n-4)} \ (n > 4) \\ \frac{1}{F_{m,n}} &\sim F_{n,m} \qquad\text{(so } F_{m,n;\,1-\alpha} = 1/F_{n,m;\,\alpha}\text{)} \\ t_n^{2} &\sim F_{1,n}, \qquad \lim_{n \to \infty} m F_{m,n} = \chi^{2}_{m} \end{aligned} \]

The reciprocal relation is what lets a one-tailed table serve both tails. Note that \(E(F) = n/(n-2) > 1\) always, so an \(F\) statistic is centred slightly above 1 even when the null hypothesis of equal variances is true.

EXAMPLE 3.3 — COMPARING TWO VARIANCES

Given. Independent normal samples: \(n_1 = 6\) with \(s_1^{2} = 24\), and \(n_2 = 11\) with \(s_2^{2} = 7.2\). Test \(H_0: \sigma_1^{2} = \sigma_2^{2}\) against \(H_1: \sigma_1^{2} > \sigma_2^{2}\).

Step 1 — the degrees of freedom. \(m = n_1 - 1 = 5\) and \(n = n_2 - 1 = 10\).

Step 2 — why the ratio is an \(F\). By section 2, \((n_i - 1)S_i^{2}/\sigma_i^{2} \sim \chi^{2}_{n_i - 1}\), independently. Under \(H_0\) the common \(\sigma^{2}\) cancels:

\[ \frac{S_1^{2}}{S_2^{2}} = \frac{\left[\frac{5 S_1^{2}}{\sigma^{2}}\right] / 5}{\left[\frac{10 S_2^{2}}{\sigma^{2}}\right] / 10} = \frac{U/5}{V/10} \sim F_{5,10}. \]

Step 3 — the statistic.

\[ F = \frac{24}{7.2} = 3.333333. \]

Step 4 — the significance probability.

\[ P\left(F_{5,10} > 3.333333\right) = 0.049697. \]

Step 5 — decide. \(0.049697 < 0.05\), so \(H_0\) is rejected at the \(5\%\) level, but only just: the tabulated \(5\%\) point of \(F_{5,10}\) is \(3.33\), and the statistic is \(3.3333\).

Interpretation. A ratio of sample variances of \(3.33\) is only marginally significant with five and ten degrees of freedom — and note \(E(F_{5,10}) = 10/8 = 1.25\), so an \(F\) above 1 is expected even under \(H_0\). Variance comparisons need large samples before they say much, which is one reason the \(F\) test for equality of variances is a poor screening test before a \(t\) test.

5. The Non-Central Distributions

STATEMENTS ONLY, AS THE SYLLABUS PRESCRIBES

Each central distribution has a non-central counterpart obtained by giving the underlying normals non-zero means. The extra parameter is the non-centrality parameter \(\lambda\), and \(\lambda = 0\) recovers the central case in every instance.

Non-central chi-square. If \(Z_i \sim N(\mu_i, 1)\) independently, then \(\sum_{i=1}^{n} Z_i^{2} \sim \chi'^{2}_{n}(\lambda)\) with \(\lambda = \sum_{i} \mu_i^{2}\), and

\[ E = n + \lambda, \qquad \operatorname{Var} = 2n + 4\lambda. \]

Non-central t. \(T' = \dfrac{Z + \delta}{\sqrt{V/n}}\) with \(Z \sim N(0,1)\), \(V \sim \chi^{2}_{n}\) independent, has \(n\) degrees of freedom and non-centrality \(\delta\). It is not symmetric unless \(\delta = 0\).

Non-central F. \(F' = \dfrac{U/m}{V/n}\) with \(U \sim \chi'^{2}_{m}(\lambda)\) and \(V \sim \chi^{2}_{n}\) central and independent.

What they are for. A central distribution is the null distribution of a test statistic; the corresponding non-central distribution is its distribution under the alternative. So the power of a \(t\) test is a probability computed from the non-central \(t\), and the power of an \(F\) test in analysis of variance from the non-central \(F\). Any sample-size calculation is a non-centrality calculation in disguise.

Key Take-aways