Skip to the content

Topics Covered

Wilks' Lambda Simple (Pearson) Correlation Partial Correlation Multiple Correlation Two-Sample T^2 Confidence Region Fisher's Linear Discriminant (Two Groups) Mahalanobis Distance Probability of Misclassification Multiple Groups Population PCA Standardisation

Topic Overview — What & Why

Unit VIII generalises univariate methods to several jointly observed variables. The multivariate normal distribution plays the same central role here that the univariate normal plays in basic inference.

  • Multivariate normal distribution: generalises $N(\mu,\sigma^2)$; characterised by mean vector and covariance matrix; conditional distributions remain normal — foundation of regression and classification.
  • Estimation of mean vector & covariance matrix: $(\bar{\mathbf X},\mathbf S)$ are jointly sufficient and complete under MVN; Basu-style independence used in deriving sampling distributions.
  • Distribution of sample mean vector: $\bar{\mathbf X}\sim N_p(\boldsymbol\mu,\Sigma/n)$ — basis of confidence ellipsoids and Hotelling's tests.
  • Wishart distribution: multivariate generalisation of $\chi^2$; describes the distribution of sample covariance matrices.
  • Simple, partial, multiple correlation: three measures of association — pairwise, conditional, and one-against-many.
  • Hotelling's $T^2$: multivariate analogue of the $t$-statistic; tests hypotheses about mean vectors.
  • Discriminant analysis: finds linear combinations that best separate two or more populations; basis of Fisher's classifier.
  • Principal component analysis (PCA): reduces dimensionality by finding orthogonal directions of maximum variance — ubiquitous in data compression and exploratory analysis.
  • Canonical correlation analysis: finds the most-correlated linear combinations between two sets of variables — generalises multiple correlation.

1. Multivariate Normal Distribution

Why this section? Almost every multivariate technique in the syllabus assumes data come from $N_p(\boldsymbol\mu,\Sigma)$. Mastery of this distribution and its conditional / marginal properties is non-negotiable.

$\mathbf X=(X_1,\ldots,X_p)^T\sim N_p(\boldsymbol\mu,\Sigma)$ has pdf $$f(\mathbf x)=\frac{1}{(2\pi)^{p/2}|\Sigma|^{1/2}}\exp\left\{-\tfrac{1}{2}(\mathbf x-\boldsymbol\mu)^T\Sigma^{-1}(\mathbf x-\boldsymbol\mu)\right\},\,\Sigma\succ 0.$$

Properties

EXAMPLE 1 $\mathbf X\sim N_2(\mathbf 0,\Sigma)$ with $\Sigma=\begin{pmatrix}4&2\\2&3\end{pmatrix}.$ $X_1+X_2\sim N(0,4+2(2)+3)=N(0,11).$
EXAMPLE 2 $\mathbf X\sim N_3$ with $\Sigma=\text{diag}(1,4,9).$ Components are independent normals; $X_1\sim N(0,1),\,X_2\sim N(0,4),\,X_3\sim N(0,9).$

🌍 Where it's used in real life

  1. Modelling correlated test scores.
  2. Returns of several assets in a portfolio.
  3. Sensor arrays with correlated noise.
  4. Height, weight and age modelled jointly.
  5. The base for many statistics and ML methods.

2. Estimation of Mean Vector and Covariance Matrix

Sample $\mathbf X_1,\ldots,\mathbf X_n$ iid $N_p(\boldsymbol\mu,\Sigma).$ MLEs: $$\hat{\boldsymbol\mu}=\bar{\mathbf X}=\tfrac{1}{n}\sum\mathbf X_i,\quad \hat\Sigma=\tfrac{1}{n}\sum(\mathbf X_i-\bar{\mathbf X})(\mathbf X_i-\bar{\mathbf X})^T.$$ Unbiased: $\mathbf S=\tfrac{1}{n-1}\sum(\mathbf X_i-\bar{\mathbf X})(\mathbf X_i-\bar{\mathbf X})^T.$

$\bar{\mathbf X}$ and $\mathbf S$ (or $\hat\Sigma$) are independent. $(\bar{\mathbf X},\mathbf S)$ is jointly complete and sufficient.

EXAMPLE 1 $n=5$ bivariate observations; compute $\bar{\mathbf X}=(\bar X_1,\bar X_2)^T$ and $\mathbf S$ as 2×2 matrix with sample variances on diagonal and sample covariance off-diagonal.
EXAMPLE 2 For $p=3,n=20$: $(n-1)\mathbf S\sim W_3(19,\Sigma).$ Rank of $\mathbf S$ equals 3 with probability 1 when $n>p.$

🌍 Where it's used in real life

  1. Estimating asset means and covariances in finance.
  2. Summarising multi-feature datasets.
  3. Risk matrices in portfolio management.
  4. Calibrating multivariate models.
  5. Covariance estimation for sensor fusion.

3. Distribution of Sample Mean Vector

If $\mathbf X_i\overset{\text{iid}}{\sim}N_p(\boldsymbol\mu,\Sigma)$, then $$\bar{\mathbf X}\sim N_p(\boldsymbol\mu,\Sigma/n).$$ Standardised quadratic form: $n(\bar{\mathbf X}-\boldsymbol\mu)^T\Sigma^{-1}(\bar{\mathbf X}-\boldsymbol\mu)\sim\chi^2_p.$

EXAMPLE 1 $p=2,n=25,\Sigma=\begin{pmatrix}1&0.5\\0.5&1\end{pmatrix}.$ $\bar{\mathbf X}\sim N_2(\boldsymbol\mu,\Sigma/25).$
EXAMPLE 2 For known $\Sigma$, 100(1-$\alpha$)% confidence ellipsoid for $\boldsymbol\mu$: $n(\bar{\mathbf x}-\boldsymbol\mu)^T\Sigma^{-1}(\bar{\mathbf x}-\boldsymbol\mu)\le\chi^2_{p,\alpha}.$

🌍 Where it's used in real life

  1. Confidence regions for several means at once.
  2. Quality control of several measurements together.
  3. Multivariate polling estimates.
  4. Batch testing of multi-spec products.
  5. Foundation for Hotelling's T² tests.

4. Wishart Distribution

If $\mathbf Z_1,\ldots,\mathbf Z_n\overset{\text{iid}}{\sim}N_p(\mathbf 0,\Sigma)$, then $\mathbf W=\sum\mathbf Z_i\mathbf Z_i^T\sim W_p(n,\Sigma).$

Intuition. The Wishart is the multivariate generalization of the $\chi^2$: just as $\sum Z_i^2\sim\sigma^2\chi^2_n$ governs the sample variance in one dimension, $\sum\mathbf Z_i\mathbf Z_i^T$ governs the sample covariance matrix. It is therefore the distribution sitting behind Hotelling's $T^2$, MANOVA, and every inference about $\Sigma$.

Properties

Wilks' Lambda

$\Lambda=|\mathbf W_1|/|\mathbf W_1+\mathbf W_2|$ — used in MANOVA, equals $\prod\frac{1}{1+\lambda_j}.$
EXAMPLE 1 $p=1,n=10,\Sigma=4$: $W\sim 4\chi^2_{10}$ — sum of squares scaled.
EXAMPLE 2 For $p=2,n=20$: $\mathbf W\sim W_2(20,\Sigma).$ Determinant has distribution $|\Sigma|\chi^2_{20}\chi^2_{19}.$

🌍 Where it's used in real life

  1. Inference about a covariance matrix.
  2. Bayesian priors for covariances.
  3. Uncertainty in a portfolio risk matrix.
  4. Distributions behind MANOVA tests.
  5. Random-matrix models.

5. Simple, Partial & Multiple Correlation

Simple (Pearson) Correlation

$\rho_{ij}=\sigma_{ij}/\sqrt{\sigma_{ii}\sigma_{jj}}.$ Sample: $r_{ij}.$

Partial Correlation

Correlation between $X_i,X_j$ removing the linear effect of $X_k$: $$\rho_{ij\cdot k}=\frac{\rho_{ij}-\rho_{ik}\rho_{jk}}{\sqrt{(1-\rho_{ik}^2)(1-\rho_{jk}^2)}}.$$

Multiple Correlation

$\rho_{1\cdot 2,3,\ldots,p}$ = max correlation between $X_1$ and any linear combination of $X_2,\ldots,X_p:$ $$R_{1\cdot 2,\ldots,p}^2=\frac{\boldsymbol\sigma_{12}^T\Sigma_{22}^{-1}\boldsymbol\sigma_{12}}{\sigma_{11}}.$$

Tests

EXAMPLE 1 $\rho_{12}=0.6,\rho_{13}=0.5,\rho_{23}=0.4.$ $\rho_{12\cdot 3}=\frac{0.6-0.5\cdot 0.4}{\sqrt{(1-0.25)(1-0.16)}}=\frac{0.4}{0.794}=0.504.$
EXAMPLE 2 $R^2=0.7$ in 3-variable model, $n=30.$ $F=\frac{0.7/2}{0.3/27}=31.5,$ $p$-value$<0.001.$

🌍 Where it's used in real life

  1. The link between two variables (simple).
  2. Study effect on marks, controlling for IQ (partial).
  3. Predicting one variable from many (multiple).
  4. Removing confounders in research.
  5. Ranking feature relevance in ML.

6. Hotelling's $T^2$ Statistic

For $\bar{\mathbf X}\sim N_p(\boldsymbol\mu,\Sigma/n)$ with $\Sigma$ unknown estimated by $\mathbf S$: $$T^2=n(\bar{\mathbf X}-\boldsymbol\mu_0)^T\mathbf S^{-1}(\bar{\mathbf X}-\boldsymbol\mu_0).$$ $$\frac{n-p}{p(n-1)}T^2\sim F_{p,n-p}.$$

Two-Sample $T^2$

$T^2=\frac{n_1 n_2}{n_1+n_2}(\bar{\mathbf X}-\bar{\mathbf Y})^T\mathbf S_p^{-1}(\bar{\mathbf X}-\bar{\mathbf Y})$ with pooled covariance. $\frac{n_1+n_2-p-1}{(n_1+n_2-2)p}T^2\sim F_{p,n_1+n_2-p-1}.$

Confidence Region

$\{\boldsymbol\mu:n(\bar{\mathbf x}-\boldsymbol\mu)^T\mathbf S^{-1}(\bar{\mathbf x}-\boldsymbol\mu)\le\frac{p(n-1)}{n-p}F_{p,n-p,\alpha}\}.$
EXAMPLE 1 $p=3,n=20,T^2=15.$ $F=(17/57)\cdot 15=4.47$; critical $F_{3,17,0.05}=3.20.$ Reject $H_0:\boldsymbol\mu=\boldsymbol\mu_0.$
EXAMPLE 2 For $p=1,$ $T^2=t^2$ — recovers univariate $t$-test.

🌍 Where it's used in real life

  1. Comparing two groups on many measures at once.
  2. Multivariate quality control.
  3. Before/after checks on several health metrics.
  4. Comparing full product profiles.
  5. Testing mean vectors in research.

7. Discriminant Analysis

Discriminant analysis finds a rule — usually a linear combination of the measured variables — that best separates two or more known groups, and then uses it to classify a new observation into its most likely group.

Fisher's Linear Discriminant (Two Groups)

Find $\mathbf a$ to maximize between-group / within-group variance: $$\mathbf a=\Sigma^{-1}(\boldsymbol\mu_1-\boldsymbol\mu_2).$$ Classification rule: assign $\mathbf x$ to group 1 if $$\mathbf a^T\mathbf x\ge\tfrac{1}{2}\mathbf a^T(\boldsymbol\mu_1+\boldsymbol\mu_2).$$ Sample version uses $\bar{\mathbf X}_1,\bar{\mathbf X}_2,\mathbf S_p.$

Mahalanobis Distance

$D^2=(\boldsymbol\mu_1-\boldsymbol\mu_2)^T\Sigma^{-1}(\boldsymbol\mu_1-\boldsymbol\mu_2).$ Larger $D^2\Rightarrow$ better separability.

Probability of Misclassification

For equal priors and costs: $P(\text{error})=\Phi(-D/2).$

Multiple Groups

Generalised to $k$ groups: maximise $|\mathbf B|/|\mathbf W|$ — leads to canonical discriminant axes.
EXAMPLE 1 $D^2=4\Rightarrow$ optimal misclassification probability $=\Phi(-1)=0.1587.$
EXAMPLE 2 Two species of iris with $p=4$ measurements: Fisher's classic dataset; LDA gives near-perfect separation.

🌍 Where it's used in real life

  1. Classifying loan applicants (default or not).
  2. Medical diagnosis from several test results.
  3. Classifying flower or plant species.
  4. Basics of face and handwriting recognition.
  5. Grouping customers by credit risk.

8. Principal Component Analysis (PCA)

Reduces dimensionality by finding orthogonal directions of maximum variance.

Population PCA

$\Sigma$ has eigenvalues $\lambda_1\ge\lambda_2\ge\cdots\ge\lambda_p\ge 0$ with eigenvectors $\mathbf e_1,\ldots,\mathbf e_p.$ The $i$-th principal component: $$Y_i=\mathbf e_i^T\mathbf X,\quad \text{Var}(Y_i)=\lambda_i,\quad \text{Cov}(Y_i,Y_j)=0.$$

Properties

Standardisation

If variables on different scales, use correlation matrix instead of covariance.
EXAMPLE 1 $\Sigma=\begin{pmatrix}5&2\\2&2\end{pmatrix}$. Eigenvalues 6,1; eigenvectors $(2,1)^T/\sqrt 5,(1,-2)^T/\sqrt 5.$ PC1 explains $6/7=85.7\%$ variance.
variable X₁ variable X₂ PC1PC2 PCA rotates to axes of maximum variance
Principal components. PC1 (orange) points along the direction of greatest spread ($\lambda_1=6$); PC2 (green) is orthogonal to it with the leftover variance ($\lambda_2=1$). Rotating the data onto these axes decorrelates it and concentrates variance in the first few components.
EXAMPLE 2 For 4 standardised variables (correlation matrix has trace 4): if first 2 eigenvalues sum to 3.2, retain 2 PCs covering 80% variance.

🌍 Where it's used in real life

  1. Reducing many features to a few.
  2. Compressing images.
  3. Visualising high-dimensional data.
  4. Combining indicators into a single index.
  5. Noise reduction in signals.

9. Canonical Correlation Analysis (CCA)

For two random vectors $\mathbf X(p\times 1)$ and $\mathbf Y(q\times 1)$ with covariance partitioned $\Sigma=\begin{pmatrix}\Sigma_{11}&\Sigma_{12}\\\Sigma_{21}&\Sigma_{22}\end{pmatrix}.$

Find $\mathbf a,\mathbf b$ to maximize $\text{Corr}(\mathbf a^T\mathbf X,\mathbf b^T\mathbf Y)$ subject to unit variance. Solution from eigenvalues of $$\Sigma_{11}^{-1}\Sigma_{12}\Sigma_{22}^{-1}\Sigma_{21}.$$

Properties

Test

$H_0:\rho_1=\cdots=\rho_k=0$ uses Wilks' $\Lambda=\prod(1-\hat\rho_i^2).$
EXAMPLE 1 $X$ = academic scores (math, verbal); $Y$ = job performance, salary. CCA finds linear combinations on each side maximally correlated.
EXAMPLE 2 If $p=q=1$: canonical correlation = absolute value of simple correlation $|\rho|.$

🌍 Where it's used in real life

  1. Relating aptitude tests to job performance.
  2. Linking diet variables to health variables.
  3. Marketing-mix variables vs sales outcomes.
  4. Genes vs traits associations.
  5. Relating two blocks of survey questions.