Skip to the content

Useful for UGC NET · ASRB NET · ISS

Welcome

This is the complete study package for Multivariate Analysis (STS-202) together with Section B of STS-205, Estimation Theory and Multivariate Analysis (Conventional). It is the course in which the whole Foundation course is done again with \(p\) variables at once — a mean becomes a vector, a variance becomes a matrix, and the algebra of that matrix turns out to carry nearly all of the statistics.

What is assumed, and where to revise it. This course has no Foundation counterpart on the site, because there is no Foundation multivariate course. What it does lean on is all here: Each unit names what it borrows and links it, then goes past it. Nothing already proved elsewhere on this site is proved a second time.

Course Objectives

  1. To understand the extension from univariate to multivariate probability distributions, and its importance, necessity and real-time applications.
  2. To derive the basic properties, applications and problems of the multivariate probability distributions (multinomial and multivariate normal) and the multivariate sampling distributions (Wishart, Hotelling's \(T^{2}\) and Wilks' lambda).
  3. To learn the basic concepts of the multivariate statistical techniques used for data analysis — classification, identification, testing, clustering and the rest.

Course Outcomes

  1. Able to solve and derive the common and special properties of the multinomial and multivariate normal distributions and of the multivariate sampling distributions.
  2. Able to carry out the analysis of any multivariate data set using the multivariate tools — linear discriminant analysis, principal components, multidimensional scaling, factor analysis and cluster analysis.
  3. Able to identify the real-time applications of the multivariate statistical tools.

Units in this Course

UNIT 1

The Multinomial and Multivariate Normal Distributions

Mean vectors and dispersion matrices with the three rules used everywhere afterwards; the multinomial, and why every one of its covariances is negative; the multivariate normal density read factor by factor, its moment generating function, and the partition theorem for marginals and conditionals proved in full; maximum likelihood for \(\boldsymbol\mu\) and \(\boldsymbol\Sigma\).

UNIT 2

Wishart Distribution and the Distributions of Correlations

The Wishart as the chi-square in matrix form, with four properties each traced back to its univariate ancestor; why \(m \ge p\) is a hard requirement; the generalized variance as a product of independent chi-squares and the bias that follows; the null distributions of simple, multiple and partial correlations; and inference for regression coefficients.

UNIT 3

Hotelling's \(T^{2}\), Wilks' \(\Lambda\) and Discrimination

What goes wrong with \(p\) separate tests; \(T^{2}\) as the likelihood ratio test, with its invariance and its \(F\) transformation; Mahalanobis \(D^{2}\) and the two-sample test; contrasts for equality of components; Wilks' \(\Lambda\) and Bartlett's approximation; and Fisher's discriminant function with its misclassification probability.

UNIT 4

Principal Components, Clustering, Scaling and Factor Analysis

Principal components as the solution of a Rayleigh quotient, solved exactly for \(2 \times 2\) and from the characteristic cubic for \(3 \times 3\); canonical correlations; three linkages compared on one data set and \(K\)-means worked to convergence; classical scaling that recovers a triangle from its distances; the orthogonal factor model in closed form; and path, correspondence and conjoint analysis.

PRACTICAL

STS-205 Section B — Conventional

All thirteen prescribed experiments: eleven worked in full in the units on data chosen so that each result checks another, plus the two that are not — writing a \(p\)-variate normal density from its parameters and reading the parameters back out of one, and the one-sample Mahalanobis \(D^{2}\).

REFERENCE

Official Syllabus

The prescribed unit-wise outline for STS-202 and the Section B practical list for STS-205, as printed, with the objectives, outcomes and the reading list.

How This Course Connects to the Others

What is built hereWhere it is used
Mean vectors, dispersion matrices and the rule \(\operatorname{Cov}(\mathbf{AX}) = \mathbf{A}\boldsymbol\Sigma\mathbf{A}'\) Every derivation in this course, and the variance of any estimator that is a linear function of the data — including the least squares estimator of Linear Algebra & Linear Models, Unit 4
Maximum likelihood for \(\boldsymbol\mu\) and \(\boldsymbol\Sigma\) The same sufficiency and completeness argument as Estimation Theory, Unit 2, applied to a matrix parameter
The Wishart distribution Pooling covariance matrices; the exact distribution of every statistic in Unit 3; the matrix version of the chi-square results of Distribution Theory, Unit 3
Hotelling's \(T^{2}\) and Wilks' \(\Lambda\) Multivariate analysis of variance, and therefore the multivariate versions of the designs in Design and Analysis of Experiments
Mahalanobis \(D^{2}\) and discriminant analysis Classification of any kind; outlier detection; the nearest-centroid rules underlying machine learning classifiers
Principal components and factor analysis Dimension reduction before any modelling; the standard remedy for the multicollinearity diagnosed in Linear Algebra & Linear Models, Unit 4
Clustering and scaling Exploratory analysis of any data set with no response variable — the unsupervised half of the data science material

Next course in learning order: Applied Statistics Applied statistics