Statement. If \(Z \sim N(0,1)\) then \(Y = Z^{2} \sim \chi^{2}_{1}\), with density
\[ f_Y(y) = \frac{1}{\sqrt{2\pi y}} e^{-y/2}, \qquad y > 0. \]Step 1 — go through the distribution function, because the map is not one-to-one. \(z \mapsto z^{2}\) sends both \(+z\) and \(-z\) to the same \(y\), so the change-of-variable formula cannot be applied directly. For \(y > 0\),
\[ F_Y(y) = P\left(Z^{2} \le y\right) = P\left(-\sqrt{y} \le Z \le \sqrt{y}\right) = \Phi(\sqrt{y}) - \Phi(-\sqrt{y}). \]Step 2 — use symmetry. \(\Phi(-a) = 1 - \Phi(a)\), so
\[ F_Y(y) = \Phi(\sqrt{y}) - \big[1 - \Phi(\sqrt{y})\big] = 2\Phi(\sqrt{y}) - 1. \]Step 3 — differentiate, by the chain rule. With \(\dfrac{d}{dy}\sqrt{y} = \dfrac{1}{2\sqrt{y}}\) and \(\Phi' = \varphi\),
\[ f_Y(y) = 2 \varphi(\sqrt{y}) \cdot \frac{1}{2\sqrt{y}} = \frac{\varphi(\sqrt{y})}{\sqrt{y}}. \]Step 4 — substitute the standard normal density. \(\varphi(t) = \dfrac{1}{\sqrt{2\pi}} e^{-t^{2}/2}\), so \(\varphi(\sqrt{y})\) has \(t^{2} = y\) and
\[ f_Y(y) = \frac{1}{\sqrt{2\pi y}} e^{-y/2}, \qquad y > 0. \quad \blacksquare \]Step 5 — recognise it. Compare with the Gamma(shape \(\alpha\), rate \(\beta\)) density \(\dfrac{\beta^{\alpha}}{\Gamma(\alpha)} y^{\alpha-1} e^{-\beta y}\). Matching exponents gives \(\alpha - 1 = -\tfrac12\), so \(\alpha = \tfrac12\), and \(\beta = \tfrac12\). The constant agrees because \(\Gamma(\tfrac12) = \sqrt{\pi}\). So \(\chi^{2}_{1} = \text{Gamma}(\tfrac12, \tfrac12)\).
Statement. If \(Z_1, \ldots, Z_n\) are independent \(N(0,1)\) then \(\sum_{i} Z_i^{2} \sim \chi^{2}_{n}\), with density
\[ f(y) = \frac{1}{2^{n/2}\,\Gamma(n/2)} \, y^{n/2 - 1} e^{-y/2}, \qquad y > 0. \]Step 1 — the MGF of one square. From Step 5 above, \(Z^{2} \sim \text{Gamma}(\tfrac12, \tfrac12)\), whose MGF is
\[ M_{Z^{2}}(t) = \left(\frac{1/2}{1/2 - t}\right)^{1/2} = (1 - 2t)^{-1/2}, \qquad t < \tfrac12. \]Step 2 — multiply, using independence. The MGF of a sum of independent variables is the product of their MGFs, so
\[ M_{\sum Z_i^{2}}(t) = \left[(1-2t)^{-1/2}\right]^{n} = (1 - 2t)^{-n/2}. \]Step 3 — identify. \((1-2t)^{-n/2}\) is the MGF of Gamma\((n/2, 1/2)\), whose density is the one displayed. By the uniqueness of the MGF the distribution is determined. \(\blacksquare\)
Additivity follows at once. If \(U \sim \chi^{2}_{m}\) and \(V \sim \chi^{2}_{n}\) independently, then \(M_{U+V}(t) = (1-2t)^{-m/2}(1-2t)^{-n/2} = (1-2t)^{-(m+n)/2}\), so \(U + V \sim \chi^{2}_{m+n}\).
The skewness \(\sqrt{8/n} \to 0\), so the chi-square becomes more symmetric as the degrees of freedom grow — visible in the figure below, where \(\chi^{2}_{2}\) is a decreasing exponential and \(\chi^{2}_{9}\) is already nearly bell-shaped.
Statement. Let \(X_1, \ldots, X_n\) be i.i.d. \(N(\mu, \sigma^{2})\), and
\[ \bar X = \frac{1}{n}\sum_{i=1}^{n} X_i, \qquad S^{2} = \frac{1}{n-1}\sum_{i=1}^{n} \left(X_i - \bar X\right)^{2}. \]Then (i) \(\bar X \sim N(\mu, \sigma^{2}/n)\); (ii) \(\dfrac{(n-1)S^{2}}{\sigma^{2}} \sim \chi^{2}_{n-1}\); and (iii) \(\bar X\) and \(S^{2}\) are independent.
Proof of (i). A linear combination of independent normals is normal, with \(E(\bar X) = \mu\) and \(\operatorname{Var}(\bar X) = \dfrac{1}{n^{2}} \cdot n\sigma^{2} = \dfrac{\sigma^{2}}{n}\).
Proof of (ii) and (iii), by the Helmert transformation.
Step 1. Standardise: put \(Z_i = (X_i - \mu)/\sigma\), so the \(Z_i\) are independent \(N(0,1)\), and the claims become statements about \(\sum_i (Z_i - \bar Z)^{2}\) and \(\bar Z\).
Step 2. Apply the orthogonal Helmert matrix \(H\), whose first row is \((1/\sqrt{n}, \ldots, 1/\sqrt{n})\) and whose \(k\)-th row (\(k = 2, \ldots, n\)) is
\[ \frac{1}{\sqrt{k(k-1)}}\big(\underbrace{1, 1, \ldots, 1}_{k-1}, \; -(k-1), \; 0, \ldots, 0\big). \]Each row has unit length and any two rows are orthogonal — which is exactly what makes \(H\) orthogonal, \(H'H = I\). Set \(Y = HZ\).
Step 3. An orthogonal transformation of independent standard normals gives independent standard normals. The joint density depends on \(Z\) only through \(Z'Z\), which is unchanged since \(Y'Y = Z'H'HZ = Z'Z\), and the Jacobian is \(|\det H| = 1\). So \(Y_1, \ldots, Y_n\) are independent \(N(0,1)\).
Step 4. Read off the first component: \(Y_1 = \dfrac{1}{\sqrt{n}}\sum_i Z_i = \sqrt{n}\,\bar Z\).
Step 5. Use the length-preserving property on the remaining components:
\[ \sum_{k=2}^{n} Y_k^{2} = \sum_{i=1}^{n} Z_i^{2} - Y_1^{2} = \sum_{i=1}^{n} Z_i^{2} - n\bar Z^{2} = \sum_{i=1}^{n}\left(Z_i - \bar Z\right)^{2}, \]the last step being the usual identity \(\sum (z_i - \bar z)^2 = \sum z_i^2 - n\bar z^2\).
Step 6. Conclude. The left-hand side of Step 5 is a sum of \(n-1\) independent squared standard normals, so it is \(\chi^{2}_{n-1}\); and since
\[ \frac{(n-1)S^{2}}{\sigma^{2}} = \sum_{i=1}^{n}\left(Z_i - \bar Z\right)^{2} = \sum_{k=2}^{n} Y_k^{2}, \]claim (ii) holds. Claim (iii) holds because \(\bar X\) is a function of \(Y_1\) alone while \(S^{2}\) is a function of \(Y_2, \ldots, Y_n\) alone, and those are independent by Step 3. \(\blacksquare\)
Why (iii) matters. The \(t\) statistic puts \(\bar X - \mu\) over \(S/\sqrt{n}\). Its distribution is derived in section 3 from a normal divided by an independent chi-square. Without independence there is no \(t\) distribution, and every small-sample test on a normal mean loses its justification.
Given. A random sample of \(n = 10\) from \(N(\mu, \sigma^{2})\).
Asked. Find \(E(S^{2})\), \(\operatorname{Var}(S^{2})\), and \(P(S^{2} > 1.8\sigma^{2})\).
Step 1 — name the pivotal quantity. By part (ii), \(W = \dfrac{(n-1)S^{2}}{\sigma^{2}} = \dfrac{9 S^{2}}{\sigma^{2}} \sim \chi^{2}_{9}\).
Step 2 — the mean of \(S^{2}\). \(E(W) = 9\), so
\[ E\left(\frac{9S^{2}}{\sigma^{2}}\right) = 9 \;\Rightarrow\; E(S^{2}) = \sigma^{2}. \]The divisor \(n-1\) is exactly what makes \(S^{2}\) unbiased.
Step 3 — the variance of \(S^{2}\). \(\operatorname{Var}(W) = 2 \times 9 = 18\), and \(S^{2} = \dfrac{\sigma^{2}}{9} W\), so
\[ \operatorname{Var}(S^{2}) = \left(\frac{\sigma^{2}}{9}\right)^{2} \times 18 = \frac{18\sigma^{4}}{81} = \frac{2\sigma^{4}}{9} = 0.222222\,\sigma^{4}. \]The standard deviation of \(S^{2}\) is \(\sqrt{0.222222}\,\sigma^{2} = 0.4714\sigma^{2}\) — that is \(47\%\) of what it is estimating.
Step 4 — the tail probability. Translate the event into \(W\):
\[ P\left(S^{2} > 1.8\sigma^{2}\right) = P\left(\frac{9S^{2}}{\sigma^{2}} > 9 \times 1.8\right) = P\left(\chi^{2}_{9} > 16.2\right). \]Step 5 — evaluate.
\[ P\left(\chi^{2}_{9} > 16.2\right) = 0.062821. \]For orientation, the tabulated \(5\%\) point of \(\chi^{2}_{9}\) is \(16.919\), which confirms the value must be a little above \(0.05\), and it is.
Interpretation. With only ten observations there is a \(6.3\%\) chance that the sample variance overstates the true variance by \(80\%\) or more. A single small-sample variance is a weak estimate, which is why a confidence interval for \(\sigma^{2}\) built from the \(\chi^{2}_{9}\) percentage points \(3.325\) and \(16.919\) is so wide.
Definition. If \(Z \sim N(0,1)\) and \(V \sim \chi^{2}_{n}\) are independent, then
\[ T = \frac{Z}{\sqrt{V/n}} \sim t_n. \]Step 1 — write the joint density. By independence it is the product
\[ f_{Z,V}(z,v) = \frac{1}{\sqrt{2\pi}} e^{-z^{2}/2} \cdot \frac{1}{2^{n/2}\Gamma(n/2)} v^{n/2-1} e^{-v/2}. \]Step 2 — transform. Put \(t = z/\sqrt{v/n}\) and keep \(u = v\). The inverse is \(z = t\sqrt{u/n}\), \(v = u\), and the Jacobian is
\[ J = \begin{vmatrix} \sqrt{u/n} & \ast \\ 0 & 1 \end{vmatrix} = \sqrt{u/n}, \]the lower-left entry being \(\partial v/\partial t = 0\), so the determinant is the product of the diagonal.
Step 3 — substitute.
\[ f_{T,U}(t,u) = \frac{1}{\sqrt{2\pi}\,2^{n/2}\Gamma(n/2)} \, u^{n/2-1} \exp\left\{-\frac{u}{2}\left(1 + \frac{t^{2}}{n}\right)\right\}\sqrt{\frac{u}{n}}. \]Step 4 — integrate \(u\) out. Collecting powers, \(u^{n/2-1}\sqrt{u} = u^{(n+1)/2 - 1}\), so the \(u\)-integral is a gamma integral with shape \((n+1)/2\) and rate \(\tfrac12\left(1 + t^{2}/n\right)\):
\[ \int_{0}^{\infty} u^{(n+1)/2-1} e^{-u\beta} du = \frac{\Gamma\!\left(\frac{n+1}{2}\right)}{\beta^{(n+1)/2}}, \qquad \beta = \tfrac12\left(1 + \tfrac{t^{2}}{n}\right). \]Step 5 — collect the constants. The powers of 2 cancel against \(2^{n/2}\) and \(\beta^{(n+1)/2}\), leaving
\[ f_T(t) = \frac{\Gamma\!\left(\frac{n+1}{2}\right)}{\sqrt{n\pi}\,\Gamma\!\left(\frac{n}{2}\right)} \left(1 + \frac{t^{2}}{n}\right)^{-(n+1)/2}, \qquad t \in \mathbb{R}. \quad \blacksquare \]At \(n = 1\) the \(t\) density is \(\dfrac{1}{\pi(1+t^{2})}\), exactly the Cauchy of Unit 1, section 6 — which is why \(t_1\) has no mean. The limit to \(N(0,1)\) follows from \(\left(1 + t^{2}/n\right)^{-(n+1)/2} \to e^{-t^{2}/2}\).
Given. \(n = 16\) observations from a normal population, \(\bar x = 102.5\), \(s = 5\). Test \(H_0: \mu = 100\) against \(H_1: \mu \ne 100\).
Asked. The statistic, its null distribution, and the significance probability.
Step 1 — why a \(t\) and not a \(z\). \(\sigma\) is unknown and replaced by \(s\). By the theorem of section 2, \(\bar X\) and \(S^{2}\) are independent, so
\[ T = \frac{\bar X - \mu}{S/\sqrt{n}} = \frac{(\bar X - \mu)/(\sigma/\sqrt{n})}{\sqrt{\frac{(n-1)S^{2}/\sigma^{2}}{n-1}}} = \frac{Z}{\sqrt{V/(n-1)}} \sim t_{n-1}, \]which is exactly the structure required by the definition above, with \(n - 1 = 15\) degrees of freedom.
Step 2 — the standard error.
\[ \frac{s}{\sqrt{n}} = \frac{5}{\sqrt{16}} = \frac{5}{4} = 1.25. \]Step 3 — the statistic.
\[ t = \frac{102.5 - 100}{1.25} = \frac{2.5}{1.25} = 2.000. \]Step 4 — the significance probability.
\[ P\left(|t_{15}| > 2.000\right) = 0.063945. \]Step 5 — decide. \(0.063945 > 0.05\), so \(H_0\) is not rejected at the \(5\%\) level. The two-sided \(5\%\) point of \(t_{15}\) is \(2.131\), and \(2.000 < 2.131\), which agrees.
Interpretation. Had \(\sigma = 5\) been known, the statistic would be the same \(2.000\) but referred to \(N(0,1)\), giving \(P(|Z| > 2) = 0.045500\) and rejection at \(5\%\). The opposite conclusion, from the same data, purely because estimating \(\sigma\) from 16 observations costs precision. That cost is what the heavier tails of the \(t\) distribution encode.
If \(U \sim \chi^{2}_{m}\) and \(V \sim \chi^{2}_{n}\) are independent, then
\[ F = \frac{U/m}{V/n} \sim F_{m,n}, \qquad f(x) = \frac{\left(\frac{m}{n}\right)^{m/2}}{B\!\left(\frac{m}{2}, \frac{n}{2}\right)} \cdot \frac{x^{m/2-1}}{\left(1 + \frac{m}{n}x\right)^{(m+n)/2}}, \quad x > 0, \]derived by the same two-step route as the \(t\): form the joint density, transform to \((F, V)\) with Jacobian, and integrate \(V\) out against a gamma integral.
The reciprocal relation is what lets a one-tailed table serve both tails. Note that \(E(F) = n/(n-2) > 1\) always, so an \(F\) statistic is centred slightly above 1 even when the null hypothesis of equal variances is true.
Given. Independent normal samples: \(n_1 = 6\) with \(s_1^{2} = 24\), and \(n_2 = 11\) with \(s_2^{2} = 7.2\). Test \(H_0: \sigma_1^{2} = \sigma_2^{2}\) against \(H_1: \sigma_1^{2} > \sigma_2^{2}\).
Step 1 — the degrees of freedom. \(m = n_1 - 1 = 5\) and \(n = n_2 - 1 = 10\).
Step 2 — why the ratio is an \(F\). By section 2, \((n_i - 1)S_i^{2}/\sigma_i^{2} \sim \chi^{2}_{n_i - 1}\), independently. Under \(H_0\) the common \(\sigma^{2}\) cancels:
\[ \frac{S_1^{2}}{S_2^{2}} = \frac{\left[\frac{5 S_1^{2}}{\sigma^{2}}\right] / 5}{\left[\frac{10 S_2^{2}}{\sigma^{2}}\right] / 10} = \frac{U/5}{V/10} \sim F_{5,10}. \]Step 3 — the statistic.
\[ F = \frac{24}{7.2} = 3.333333. \]Step 4 — the significance probability.
\[ P\left(F_{5,10} > 3.333333\right) = 0.049697. \]Step 5 — decide. \(0.049697 < 0.05\), so \(H_0\) is rejected at the \(5\%\) level, but only just: the tabulated \(5\%\) point of \(F_{5,10}\) is \(3.33\), and the statistic is \(3.3333\).
Interpretation. A ratio of sample variances of \(3.33\) is only marginally significant with five and ten degrees of freedom — and note \(E(F_{5,10}) = 10/8 = 1.25\), so an \(F\) above 1 is expected even under \(H_0\). Variance comparisons need large samples before they say much, which is one reason the \(F\) test for equality of variances is a poor screening test before a \(t\) test.
Each central distribution has a non-central counterpart obtained by giving the underlying normals non-zero means. The extra parameter is the non-centrality parameter \(\lambda\), and \(\lambda = 0\) recovers the central case in every instance.
Non-central chi-square. If \(Z_i \sim N(\mu_i, 1)\) independently, then \(\sum_{i=1}^{n} Z_i^{2} \sim \chi'^{2}_{n}(\lambda)\) with \(\lambda = \sum_{i} \mu_i^{2}\), and
\[ E = n + \lambda, \qquad \operatorname{Var} = 2n + 4\lambda. \]Non-central t. \(T' = \dfrac{Z + \delta}{\sqrt{V/n}}\) with \(Z \sim N(0,1)\), \(V \sim \chi^{2}_{n}\) independent, has \(n\) degrees of freedom and non-centrality \(\delta\). It is not symmetric unless \(\delta = 0\).
Non-central F. \(F' = \dfrac{U/m}{V/n}\) with \(U \sim \chi'^{2}_{m}(\lambda)\) and \(V \sim \chi^{2}_{n}\) central and independent.
What they are for. A central distribution is the null distribution of a test statistic; the corresponding non-central distribution is its distribution under the alternative. So the power of a \(t\) test is a probability computed from the non-central \(t\), and the power of an \(F\) test in analysis of variance from the non-central \(F\). Any sample-size calculation is a non-centrality calculation in disguise.