Topics Covered
Contents
- 1. Multivariate Normal Distribution
- 2. Estimation of Mean Vector and Covariance Matrix
- 3. Distribution of Sample Mean Vector
- 4. Wishart Distribution
- 5. Simple, Partial & Multiple Correlation
- 6. Hotelling's $T^2$ Statistic
- 7. Discriminant Analysis
- 8. Principal Component Analysis
- 9. Canonical Correlation Analysis
Topic Overview — What & Why
Unit VIII generalises univariate methods to several jointly observed variables. The multivariate normal distribution plays the same central role here that the univariate normal plays in basic inference.
- Multivariate normal distribution: generalises $N(\mu,\sigma^2)$; characterised by mean vector and covariance matrix; conditional distributions remain normal — foundation of regression and classification.
- Estimation of mean vector & covariance matrix: $(\bar{\mathbf X},\mathbf S)$ are jointly sufficient and complete under MVN; Basu-style independence used in deriving sampling distributions.
- Distribution of sample mean vector: $\bar{\mathbf X}\sim N_p(\boldsymbol\mu,\Sigma/n)$ — basis of confidence ellipsoids and Hotelling's tests.
- Wishart distribution: multivariate generalisation of $\chi^2$; describes the distribution of sample covariance matrices.
- Simple, partial, multiple correlation: three measures of association — pairwise, conditional, and one-against-many.
- Hotelling's $T^2$: multivariate analogue of the $t$-statistic; tests hypotheses about mean vectors.
- Discriminant analysis: finds linear combinations that best separate two or more populations; basis of Fisher's classifier.
- Principal component analysis (PCA): reduces dimensionality by finding orthogonal directions of maximum variance — ubiquitous in data compression and exploratory analysis.
- Canonical correlation analysis: finds the most-correlated linear combinations between two sets of variables — generalises multiple correlation.
1. Multivariate Normal Distribution
Why this section? Almost every multivariate technique in the syllabus assumes data come from $N_p(\boldsymbol\mu,\Sigma)$. Mastery of this distribution and its conditional / marginal properties is non-negotiable.
Properties
- $E(\mathbf X)=\boldsymbol\mu,\,\text{Var}(\mathbf X)=\Sigma.$
- Linear combinations: $\mathbf a^T\mathbf X\sim N(\mathbf a^T\boldsymbol\mu,\mathbf a^T\Sigma\mathbf a).$
- Affine: $\mathbf{AX}+\mathbf b\sim N_q(\mathbf A\boldsymbol\mu+\mathbf b,\mathbf A\Sigma\mathbf A^T).$
- Marginals are normal; subvectors $\mathbf X_1\sim N_{p_1}(\boldsymbol\mu_1,\Sigma_{11}).$
- Independence iff $\Sigma_{12}=0.$
- Conditional: $\mathbf X_1\mid \mathbf X_2=\mathbf x_2\sim N(\boldsymbol\mu_1+\Sigma_{12}\Sigma_{22}^{-1}(\mathbf x_2-\boldsymbol\mu_2),\,\Sigma_{11}-\Sigma_{12}\Sigma_{22}^{-1}\Sigma_{21}).$
- $(\mathbf X-\boldsymbol\mu)^T\Sigma^{-1}(\mathbf X-\boldsymbol\mu)\sim\chi^2_p.$
- MGF: $M(\mathbf t)=\exp\{\boldsymbol\mu^T\mathbf t+\tfrac{1}{2}\mathbf t^T\Sigma\mathbf t\}.$
🌍 Where it's used in real life
- Modelling correlated test scores.
- Returns of several assets in a portfolio.
- Sensor arrays with correlated noise.
- Height, weight and age modelled jointly.
- The base for many statistics and ML methods.
2. Estimation of Mean Vector and Covariance Matrix
Sample $\mathbf X_1,\ldots,\mathbf X_n$ iid $N_p(\boldsymbol\mu,\Sigma).$ MLEs: $$\hat{\boldsymbol\mu}=\bar{\mathbf X}=\tfrac{1}{n}\sum\mathbf X_i,\quad \hat\Sigma=\tfrac{1}{n}\sum(\mathbf X_i-\bar{\mathbf X})(\mathbf X_i-\bar{\mathbf X})^T.$$ Unbiased: $\mathbf S=\tfrac{1}{n-1}\sum(\mathbf X_i-\bar{\mathbf X})(\mathbf X_i-\bar{\mathbf X})^T.$
$\bar{\mathbf X}$ and $\mathbf S$ (or $\hat\Sigma$) are independent. $(\bar{\mathbf X},\mathbf S)$ is jointly complete and sufficient.
🌍 Where it's used in real life
- Estimating asset means and covariances in finance.
- Summarising multi-feature datasets.
- Risk matrices in portfolio management.
- Calibrating multivariate models.
- Covariance estimation for sensor fusion.
3. Distribution of Sample Mean Vector
If $\mathbf X_i\overset{\text{iid}}{\sim}N_p(\boldsymbol\mu,\Sigma)$, then $$\bar{\mathbf X}\sim N_p(\boldsymbol\mu,\Sigma/n).$$ Standardised quadratic form: $n(\bar{\mathbf X}-\boldsymbol\mu)^T\Sigma^{-1}(\bar{\mathbf X}-\boldsymbol\mu)\sim\chi^2_p.$
🌍 Where it's used in real life
- Confidence regions for several means at once.
- Quality control of several measurements together.
- Multivariate polling estimates.
- Batch testing of multi-spec products.
- Foundation for Hotelling's T² tests.
4. Wishart Distribution
Intuition. The Wishart is the multivariate generalization of the $\chi^2$: just as $\sum Z_i^2\sim\sigma^2\chi^2_n$ governs the sample variance in one dimension, $\sum\mathbf Z_i\mathbf Z_i^T$ governs the sample covariance matrix. It is therefore the distribution sitting behind Hotelling's $T^2$, MANOVA, and every inference about $\Sigma$.
Properties
- If $\mathbf X_i\overset{\text{iid}}{\sim}N_p(\boldsymbol\mu,\Sigma)$, then $(n-1)\mathbf S\sim W_p(n-1,\Sigma).$
- $E(\mathbf W)=n\Sigma.$
- Additive: $W_1+W_2\sim W_p(n_1+n_2,\Sigma)$ if independent.
- $|\mathbf W|/|\Sigma|=\prod_{i=1}^p \chi^2_{n-i+1}$ (independent chi-squares).
- If $p=1,\Sigma=\sigma^2$: reduces to $\sigma^2\chi^2_n.$
Wilks' Lambda
$\Lambda=|\mathbf W_1|/|\mathbf W_1+\mathbf W_2|$ — used in MANOVA, equals $\prod\frac{1}{1+\lambda_j}.$🌍 Where it's used in real life
- Inference about a covariance matrix.
- Bayesian priors for covariances.
- Uncertainty in a portfolio risk matrix.
- Distributions behind MANOVA tests.
- Random-matrix models.
5. Simple, Partial & Multiple Correlation
Simple (Pearson) Correlation
$\rho_{ij}=\sigma_{ij}/\sqrt{\sigma_{ii}\sigma_{jj}}.$ Sample: $r_{ij}.$Partial Correlation
Correlation between $X_i,X_j$ removing the linear effect of $X_k$: $$\rho_{ij\cdot k}=\frac{\rho_{ij}-\rho_{ik}\rho_{jk}}{\sqrt{(1-\rho_{ik}^2)(1-\rho_{jk}^2)}}.$$Multiple Correlation
$\rho_{1\cdot 2,3,\ldots,p}$ = max correlation between $X_1$ and any linear combination of $X_2,\ldots,X_p:$ $$R_{1\cdot 2,\ldots,p}^2=\frac{\boldsymbol\sigma_{12}^T\Sigma_{22}^{-1}\boldsymbol\sigma_{12}}{\sigma_{11}}.$$Tests
- Test $H_0:\rho=0$: $t=r\sqrt{n-2}/\sqrt{1-r^2}\sim t_{n-2}.$
- Test $H_0:\rho=\rho_0$ (general): use Fisher's $z$-transform $\tfrac{1}{2}\log\tfrac{1+r}{1-r}\approx N(\tfrac{1}{2}\log\tfrac{1+\rho_0}{1-\rho_0},\tfrac{1}{n-3}).$
- Test $H_0:R^2=0$ in multiple correlation: $F=\frac{R^2/(p-1)}{(1-R^2)/(n-p)}\sim F_{p-1,n-p}.$
🌍 Where it's used in real life
- The link between two variables (simple).
- Study effect on marks, controlling for IQ (partial).
- Predicting one variable from many (multiple).
- Removing confounders in research.
- Ranking feature relevance in ML.
6. Hotelling's $T^2$ Statistic
Two-Sample $T^2$
$T^2=\frac{n_1 n_2}{n_1+n_2}(\bar{\mathbf X}-\bar{\mathbf Y})^T\mathbf S_p^{-1}(\bar{\mathbf X}-\bar{\mathbf Y})$ with pooled covariance. $\frac{n_1+n_2-p-1}{(n_1+n_2-2)p}T^2\sim F_{p,n_1+n_2-p-1}.$Confidence Region
$\{\boldsymbol\mu:n(\bar{\mathbf x}-\boldsymbol\mu)^T\mathbf S^{-1}(\bar{\mathbf x}-\boldsymbol\mu)\le\frac{p(n-1)}{n-p}F_{p,n-p,\alpha}\}.$🌍 Where it's used in real life
- Comparing two groups on many measures at once.
- Multivariate quality control.
- Before/after checks on several health metrics.
- Comparing full product profiles.
- Testing mean vectors in research.
7. Discriminant Analysis
Fisher's Linear Discriminant (Two Groups)
Find $\mathbf a$ to maximize between-group / within-group variance: $$\mathbf a=\Sigma^{-1}(\boldsymbol\mu_1-\boldsymbol\mu_2).$$ Classification rule: assign $\mathbf x$ to group 1 if $$\mathbf a^T\mathbf x\ge\tfrac{1}{2}\mathbf a^T(\boldsymbol\mu_1+\boldsymbol\mu_2).$$ Sample version uses $\bar{\mathbf X}_1,\bar{\mathbf X}_2,\mathbf S_p.$Mahalanobis Distance
$D^2=(\boldsymbol\mu_1-\boldsymbol\mu_2)^T\Sigma^{-1}(\boldsymbol\mu_1-\boldsymbol\mu_2).$ Larger $D^2\Rightarrow$ better separability.Probability of Misclassification
For equal priors and costs: $P(\text{error})=\Phi(-D/2).$Multiple Groups
Generalised to $k$ groups: maximise $|\mathbf B|/|\mathbf W|$ — leads to canonical discriminant axes.🌍 Where it's used in real life
- Classifying loan applicants (default or not).
- Medical diagnosis from several test results.
- Classifying flower or plant species.
- Basics of face and handwriting recognition.
- Grouping customers by credit risk.
8. Principal Component Analysis (PCA)
Reduces dimensionality by finding orthogonal directions of maximum variance.
Population PCA
$\Sigma$ has eigenvalues $\lambda_1\ge\lambda_2\ge\cdots\ge\lambda_p\ge 0$ with eigenvectors $\mathbf e_1,\ldots,\mathbf e_p.$ The $i$-th principal component: $$Y_i=\mathbf e_i^T\mathbf X,\quad \text{Var}(Y_i)=\lambda_i,\quad \text{Cov}(Y_i,Y_j)=0.$$Properties
- $\sum\text{Var}(Y_i)=\sum\lambda_i=\text{tr}(\Sigma)=\sum\sigma_{ii}.$
- Proportion of variance explained by first $k$: $\sum_{i=1}^k\lambda_i/\sum_{i=1}^p\lambda_i.$
- Correlation between $Y_i$ and $X_j$: $\rho_{Y_i,X_j}=e_{ij}\sqrt{\lambda_i}/\sqrt{\sigma_{jj}}.$
Standardisation
If variables on different scales, use correlation matrix instead of covariance.🌍 Where it's used in real life
- Reducing many features to a few.
- Compressing images.
- Visualising high-dimensional data.
- Combining indicators into a single index.
- Noise reduction in signals.
9. Canonical Correlation Analysis (CCA)
For two random vectors $\mathbf X(p\times 1)$ and $\mathbf Y(q\times 1)$ with covariance partitioned $\Sigma=\begin{pmatrix}\Sigma_{11}&\Sigma_{12}\\\Sigma_{21}&\Sigma_{22}\end{pmatrix}.$
Find $\mathbf a,\mathbf b$ to maximize $\text{Corr}(\mathbf a^T\mathbf X,\mathbf b^T\mathbf Y)$ subject to unit variance. Solution from eigenvalues of $$\Sigma_{11}^{-1}\Sigma_{12}\Sigma_{22}^{-1}\Sigma_{21}.$$
Properties
- Eigenvalues $\rho_1^2\ge\rho_2^2\ge\cdots\ge\rho_k^2\ge 0$ where $k=\min(p,q).$
- $\rho_i$ = $i$-th canonical correlation.
- Successive canonical variates uncorrelated within and across each side except for matched pair.
- Generalises multiple correlation ($q=1$ case).
Test
$H_0:\rho_1=\cdots=\rho_k=0$ uses Wilks' $\Lambda=\prod(1-\hat\rho_i^2).$🌍 Where it's used in real life
- Relating aptitude tests to job performance.
- Linking diet variables to health variables.
- Marketing-mix variables vs sales outcomes.
- Genes vs traits associations.
- Relating two blocks of survey questions.