Topics Covered
Contents
- 1. Finite, Countable and Uncountable Sets
- 2. Sequences of Real Numbers
- 3. Series of Real Numbers
- 4. Power Series and Radius of Convergence
- 5. Functions of a Real Variable
- 6. Differentiability and Mean Value Theorems
- 7. Riemann Integration and Improper Integrals
- 8. Functions of Two Real Variables
- 9. Vector Spaces and Linear Independence
- 10. Rank, Nullity, RREF
- 11. Trace, Determinant, Inverse
- 12. Systems of Linear Equations
- 13. Gram–Schmidt Orthogonalization
- 14. Characteristic Roots and Vectors
- 15. Cayley–Hamilton Theorem
- 16. Symmetric, Skew-symmetric, Orthogonal Matrices
- 17. Positive Definite Matrices and Quadratic Forms
Topic Overview — What & Why
Unit II provides the analytical and algebraic tools used throughout statistics. Real analysis underpins probability theory and the convergence proofs of estimators; matrix algebra is indispensable for multivariate analysis, linear models, and design theory.
- Finite, countable, uncountable sets: classification of "sizes" of infinity. Tells us why some integrals can be computed term-by-term and why sample spaces of continuous variables behave differently from discrete ones.
- Sequences & their convergence: bedrock of limits; Cauchy criterion certifies convergence without knowing the limit.
- Series & convergence tests: determine when infinite sums make sense — needed for moment generating functions, characteristic functions, Taylor expansions.
- Power series: functions represented as polynomials of infinite degree (e.g., $e^x,\sin x$). Radius of convergence tells us where the representation is valid.
- Limits, continuity & differentiability: calculus on $\mathbb R$. Mean value theorems power Taylor expansions used in delta-method and asymptotic theory.
- Riemann integration & improper integrals: integrals are expectations; we must know when they exist and how to compute them.
- Functions of two variables, Lagrange multipliers: needed for joint distributions, constrained optimisation in MLE/MoM, Neyman allocation.
- Vector spaces, basis, dimension: the language for talking about contrasts, design subspaces, and parameter identifiability.
- Rank, nullity, RREF, determinants: tell us when systems of equations have unique solutions — e.g., when normal equations $X'X\beta=X'y$ are solvable.
- Gram-Schmidt: constructs orthonormal bases — the foundation of QR decomposition and orthogonal contrasts in DOE.
- Eigenvalues & Cayley-Hamilton: heart of PCA, stationarity of Markov chains, stability of AR models, diagonalisation.
- Symmetric / orthogonal matrices & quadratic forms: covariance matrices are symmetric PSD; rotations are orthogonal; quadratic forms appear in Hotelling's $T^2$, $\chi^2$ tests, ANOVA decompositions.
1. Finite, Countable and Uncountable Sets
Why this section? Probability spaces can be finite (a die), countably infinite (number of arrivals), or uncountable (a continuous time interval). Different machinery applies to each.
Key Results
- $\mathbb{Z}$, $\mathbb{Q}$, finite Cartesian products of countable sets are countable.
- Countable union of countable sets is countable.
- $\mathbb{R}$, $[0,1]$, $\mathcal{P}(\mathbb{N})$ are uncountable (Cantor's diagonal argument).
🌍 Where it's used in real life
- Knowing why discrete and continuous data need different tools.
- Computer science — what is and isn't computable.
- Database keys drawn from countable sets.
- Foundations of probability (measure theory).
- Digital vs analog signal representation.
2. Sequences of Real Numbers
Key Concepts
- Bounded: $\exists M$ s.t. $|x_n|\le M$ for all $n$.
- Monotonic: non-decreasing or non-increasing.
- Monotone Convergence Theorem: a bounded monotonic sequence converges.
- Bolzano–Weierstrass: every bounded sequence has a convergent subsequence.
- Cauchy Criterion: $(x_n)$ converges in $\mathbb{R}$ $\iff$ it is Cauchy ($\forall\varepsilon>0\,\exists N: |x_m-x_n|<\varepsilon\,\forall m,n\ge N$).
Limit Algebra
$$\lim(x_n+y_n)=\lim x_n+\lim y_n,\quad \lim(x_n y_n)=(\lim x_n)(\lim y_n),\quad \lim\frac{x_n}{y_n}=\frac{\lim x_n}{\lim y_n}\text{ if }\lim y_n\ne 0.$$🌍 Where it's used in real life
- Compound-interest balances year by year.
- Iterative algorithms homing in on a solution.
- Population growth over successive years.
- Drug concentration falling between doses.
- Successive approximations in engineering design.
3. Series of Real Numbers
Tests of Convergence
- Necessary: $a_n\to 0.$ (Not sufficient — harmonic series diverges.)
- Comparison: If $0\le a_n\le b_n$ and $\sum b_n$ converges, $\sum a_n$ converges.
- Ratio (D'Alembert): Let $L=\lim|a_{n+1}/a_n|$. If $L<1$ converges; $L>1$ diverges.
- Root (Cauchy): $L=\limsup|a_n|^{1/n}$. If $L<1$ converges; $L>1$ diverges.
- Integral test: If $f$ is positive, decreasing, $\sum f(n)$ and $\int_1^\infty f$ both converge or diverge.
- Leibniz (Alternating): $\sum(-1)^n a_n$ converges if $a_n\downarrow 0.$
- Absolute convergence: $\sum|a_n|$ converges $\Rightarrow \sum a_n$ converges. Conditional: convergent but not absolutely.
p-series
$\sum 1/n^p$ converges iff $p>1.$🌍 Where it's used in real life
- Present value of a loan or annuity (geometric series).
- Total distance of a bouncing ball.
- Calculators evaluating sin, exp, log (Taylor series).
- Signal reconstruction with Fourier series.
- Valuing a perpetuity in finance.
4. Power Series and Radius of Convergence
Series converges absolutely for $|x-x_0|<R$, diverges for $|x-x_0|>R$. Boundary $|x-x_0|=R$ requires separate analysis.
🌍 Where it's used in real life
- Calculators and computers evaluating functions.
- Approximating hard functions in physics.
- Option-pricing expansions in finance.
- Solving differential equations by series.
- Error analysis in numerical methods.
5. Functions of a Real Variable
Limits
$\lim_{x\to a}f(x)=L$ if $\forall\varepsilon>0\,\exists\delta>0: 0<|x-a|<\delta\Rightarrow|f(x)-L|<\varepsilon.$Continuity
$f$ continuous at $a$ if $\lim_{x\to a}f(x)=f(a).$Uniform Continuity
$f$ is uniformly continuous on $E$ if $\forall\varepsilon>0\,\exists\delta>0: |x-y|<\delta\Rightarrow|f(x)-f(y)|<\varepsilon$ for all $x,y\in E$ (one $\delta$ works everywhere).Monotone Functions
A monotone function on $[a,b]$ has at most countably many discontinuities (all jump-type).🌍 Where it's used in real life
- Modelling cost or revenue as a function of quantity.
- Position as a smooth function of time in physics.
- Sensor calibration curves.
- Smooth, jerk-free motion in robotics.
- Temperature changing continuously over a day.
6. Differentiability and Mean Value Theorems
Rolle's Theorem
If $f$ is continuous on $[a,b]$, differentiable on $(a,b)$, and $f(a)=f(b)$, then $\exists c\in(a,b): f'(c)=0.$Lagrange's Mean Value Theorem
$\exists c\in(a,b): f(b)-f(a)=f'(c)(b-a).$Cauchy's MVT
$\exists c\in(a,b): \frac{f(b)-f(a)}{g(b)-g(a)}=\frac{f'(c)}{g'(c)}$ (suitable $g$).Taylor's Theorem
$$f(x)=\sum_{k=0}^{n}\frac{f^{(k)}(a)}{k!}(x-a)^k+R_n(x),\quad R_n(x)=\frac{f^{(n+1)}(\xi)}{(n+1)!}(x-a)^{n+1}.$$L'Hôpital's Rule
For $0/0$ or $\infty/\infty$ forms, $\lim\frac{f}{g}=\lim\frac{f'}{g'}$ when the latter exists.🌍 Where it's used in real life
- Speed as the rate of change of distance.
- Marginal cost and revenue in economics.
- Maximising profit or minimising cost.
- Rate a drug leaves the body.
- Gradient descent training in machine learning.
7. Riemann Integration and Improper Integrals
Properties
- Continuous functions on $[a,b]$ are Riemann integrable.
- Monotone bounded functions are Riemann integrable.
- FTC: If $f$ continuous, $F(x)=\int_a^x f(t)dt$ is differentiable with $F'=f.$
- $|\int f|\le\int|f|.$
Improper Integrals
- Type I (infinite limits): $\int_a^\infty f=\lim_{b\to\infty}\int_a^b f.$
- Type II (unbounded integrand): $\int_a^b f=\lim_{c\to a^+}\int_c^b f$ if $f\to\pm\infty$ at $a$.
🌍 Where it's used in real life
- Distance travelled from a speed curve (area under it).
- Total revenue from a demand curve.
- Work done by a varying force.
- Total rainfall from an intensity curve.
- Probabilities as areas under a density curve.
8. Functions of Two Real Variables
Limits and Continuity
$\lim_{(x,y)\to(a,b)}f(x,y)=L$ if $\forall\varepsilon>0\,\exists\delta>0: \|(x,y)-(a,b)\|<\delta\Rightarrow|f-L|<\varepsilon.$Partial and Total Derivative
- $f_x=\partial f/\partial x$: holding $y$ fixed.
- $f$ is differentiable at $(a,b)$ if $f(a+h,b+k)=f(a,b)+f_x h+f_y k+o(\sqrt{h^2+k^2}).$
- Differentiability $\Rightarrow$ continuity and existence of partials. Converse not true.
Maxima/Minima
At interior critical point $(f_x=f_y=0)$, set $H=\det\begin{pmatrix}f_{xx}&f_{xy}\\f_{xy}&f_{yy}\end{pmatrix}=f_{xx}f_{yy}-f_{xy}^2.$- $H>0$, $f_{xx}>0$: local minimum.
- $H>0$, $f_{xx}<0$: local maximum.
- $H<0$: saddle point.
- $H=0$: inconclusive.
Lagrange Multipliers
To extremize $f(x,y)$ subject to $g(x,y)=0$: solve $\nabla f=\lambda\nabla g,\,g=0.$Multiple Integrals
$$\iint_R f(x,y)\,dA,\qquad \iiint_V f(x,y,z)\,dV,$$ with change of variables: $\iint f(x,y)dx\,dy=\iint f(x(u,v),y(u,v))|J|\,du\,dv.$🌍 Where it's used in real life
- Profit that depends on price and advertising together.
- Least-cost mix of inputs in production (Lagrange).
- Heat maps and terrain surfaces.
- Return-vs-risk trade-offs for portfolios.
- Loss surfaces over two model parameters in ML.
9. Vector Spaces, Span, Linear Independence, Basis
Subspaces
$W\subseteq V$ is a subspace if it is closed under addition and scalar multiplication and contains $\mathbf 0$.Span and Linear Independence
- $\text{span}(S)=\{c_1 v_1+\cdots+c_k v_k:v_i\in S,c_i\in F\}.$
- $\{v_1,\ldots,v_k\}$ linearly independent if $\sum c_i v_i=0\Rightarrow c_i=0\,\forall i.$
- Basis: linearly independent spanning set; size = dimension.
🌍 Where it's used in real life
- 3D graphics and game engines.
- Documents as word vectors in text mining.
- Force and velocity vectors in physics.
- RGB colour space in imaging.
- Feature vectors in machine learning.
10. Rank, Nullity, Row Reduced Echelon Form
Row/Column Space
Row space = span of rows; column space = span of columns. $\dim(\text{row space})=\dim(\text{col space})=\text{rank}.$Rank–Nullity Theorem
$$\text{rank}(A)+\text{nullity}(A)=\text{number of columns}.$$RREF Procedure
Row operations to reach: each pivot is 1, only nonzero entry in its column.🌍 Where it's used in real life
- Checking whether equations have a solution.
- Spotting redundant sensors or variables.
- Low-rank image compression.
- Independent loops in circuit analysis.
- Detecting multicollinearity in regression.
11. Trace, Determinant, Inverse
Trace
$\text{tr}(A)=\sum a_{ii}.$ Properties: $\text{tr}(A+B)=\text{tr}A+\text{tr}B,\,\text{tr}(AB)=\text{tr}(BA),\,\text{tr}(P^{-1}AP)=\text{tr}(A).$ Also $\text{tr}(A)=\sum\lambda_i.$Determinant
For $n\times n$ matrix: $\det(A)=\sum_{\sigma\in S_n}\text{sgn}(\sigma)\prod a_{i,\sigma(i)}.$- $\det(AB)=\det A\cdot\det B$, $\det(A^T)=\det A.$
- $\det(cA)=c^n\det A.$
- Row swap multiplies det by $-1$; row scaling by $c$ multiplies by $c$; row addition leaves det unchanged.
- $A$ invertible $\iff \det A\ne 0.$
- $\det A=\prod\lambda_i.$
Inverse
$A^{-1}=\frac{1}{\det A}\,\text{adj}(A)$. $(AB)^{-1}=B^{-1}A^{-1}.$🌍 Where it's used in real life
- Solving small linear systems (Cramer's rule).
- Area and volume scaling in graphics.
- Checking a transformation is invertible.
- Covariance-matrix determinant in statistics.
- Matrix-based encryption (Hill cipher).
12. Systems of Linear Equations
$Ax=b$ where $A$ is $m\times n$.
- Consistent: $\text{rank}(A)=\text{rank}([A|b]).$
- Unique solution: $\text{rank}(A)=n.$
- Infinite solutions: $\text{rank}(A)=\text{rank}([A|b])<n.$
- No solution: $\text{rank}(A)<\text{rank}([A|b]).$
Cramer's Rule
For $n\times n$ nonsingular $A$: $x_i=\det(A_i)/\det(A)$, where $A_i$ replaces column $i$ of $A$ by $b$.🌍 Where it's used in real life
- Balancing chemical equations.
- Traffic-flow and network analysis.
- Finding currents in electrical circuits.
- Input–output models in economics.
- Fitting a curve through data points.
13. Gram–Schmidt Orthogonalization
🌍 Where it's used in real life
- QR decomposition for stable regression.
- Orthogonal signal bases in communications.
- Removing correlation among predictors.
- Orthonormal camera frames in graphics.
- Building orthogonal polynomials for fitting.
14. Characteristic Roots and Vectors
Properties
- $\sum\lambda_i=\text{tr}(A),\,\prod\lambda_i=\det(A).$
- If $\lambda$ is eigenvalue of $A$, $\lambda^k$ is eigenvalue of $A^k$, $1/\lambda$ of $A^{-1}$, $\lambda+c$ of $A+cI.$
- $A$ and $A^T$ have same eigenvalues.
- Eigenvectors of distinct eigenvalues are linearly independent.
🌍 Where it's used in real life
- Google PageRank ranking web pages.
- Principal component analysis for data reduction.
- Vibration and resonance modes of bridges.
- Stability of control systems.
- Energy levels in quantum mechanics.
15. Cayley–Hamilton Theorem
Useful for computing $A^{-1}$ and high powers of $A$.
🌍 Where it's used in real life
- Computing high powers of a matrix efficiently.
- Finding matrix inverses in control systems.
- Solving systems of differential equations.
- Digital-filter design.
- Simplifying transition-matrix calculations.
16. Symmetric, Skew-symmetric, Orthogonal Matrices
| Type | Definition | Eigenvalues |
|---|---|---|
| Symmetric | $A=A^T$ | All real; orthogonal eigenvectors |
| Skew-symmetric | $A=-A^T$ | Purely imaginary or zero; det is 0 if $n$ odd |
| Orthogonal | $A^TA=I=AA^T$ | $|\lambda|=1$; $\det=\pm 1$ |
| Idempotent | $A^2=A$ | Eigenvalues 0 or 1; rank = trace |
| Nilpotent | $A^k=0$ for some $k$ | All eigenvalues 0 |
Spectral Theorem
A real symmetric $A$ can be written $A=Q\Lambda Q^T$ with $Q$ orthogonal and $\Lambda$ diagonal of eigenvalues.🌍 Where it's used in real life
- Rotations in graphics and robotics (orthogonal).
- Covariance matrices in statistics (symmetric).
- Stress and strain tensors in engineering.
- Whitening and PCA transforms in ML.
- Coordinate rotations in GPS.
17. Positive Definite Matrices and Quadratic Forms
Quadratic Form
$Q(x)=x^T A x$ where $A$ symmetric. Categories:- Positive definite (PD): $x^T A x>0\,\forall x\ne 0\iff$ all eigenvalues $>0$ $\iff$ all leading principal minors positive (Sylvester).
- Positive semi-definite (PSD): $x^T A x\ge 0$ $\iff$ all eigenvalues $\ge 0.$
- Negative definite: all eigenvalues $<0$.
- Indefinite: both positive and negative eigenvalues.
Properties of PD Matrices
- $A=B^T B$ for some $B$ (Cholesky).
- Determinant and trace positive.
- $A^{-1}$ also PD.
- Sum of two PD is PD.
Intuition. A positive-definite quadratic form $x^TAx$ is a genuine “bowl” opening upward, curving up in every direction with its minimum at the origin. This geometry is why definiteness matters throughout statistics: a covariance matrix must be PSD because “every variance is non-negative,” and a Hessian being PD certifies that a critical point is a local minimum.
Reduction to Canonical Form
Every quadratic form can be written as $\sum\lambda_i y_i^2$ via orthogonal transformation $x=Qy$ where $Q$ diagonalizes $A.$🌍 Where it's used in real life
- Confirming a cost surface has a true minimum.
- Valid covariance matrices for portfolios.
- Stability (Lyapunov) analysis in control.
- Uniqueness of least-squares solutions.
- Energy functions in physics.