Skip to the content

Topics Covered

Limit Algebra Tests of Convergence p-series Uniform Continuity Monotone Functions Rolle's Theorem Lagrange's Mean Value Theorem Cauchy's MVT Taylor's Theorem L'Hôpital's Rule Improper Integrals Limits and Continuity

Topic Overview — What & Why

Unit II provides the analytical and algebraic tools used throughout statistics. Real analysis underpins probability theory and the convergence proofs of estimators; matrix algebra is indispensable for multivariate analysis, linear models, and design theory.

  • Finite, countable, uncountable sets: classification of "sizes" of infinity. Tells us why some integrals can be computed term-by-term and why sample spaces of continuous variables behave differently from discrete ones.
  • Sequences & their convergence: bedrock of limits; Cauchy criterion certifies convergence without knowing the limit.
  • Series & convergence tests: determine when infinite sums make sense — needed for moment generating functions, characteristic functions, Taylor expansions.
  • Power series: functions represented as polynomials of infinite degree (e.g., $e^x,\sin x$). Radius of convergence tells us where the representation is valid.
  • Limits, continuity & differentiability: calculus on $\mathbb R$. Mean value theorems power Taylor expansions used in delta-method and asymptotic theory.
  • Riemann integration & improper integrals: integrals are expectations; we must know when they exist and how to compute them.
  • Functions of two variables, Lagrange multipliers: needed for joint distributions, constrained optimisation in MLE/MoM, Neyman allocation.
  • Vector spaces, basis, dimension: the language for talking about contrasts, design subspaces, and parameter identifiability.
  • Rank, nullity, RREF, determinants: tell us when systems of equations have unique solutions — e.g., when normal equations $X'X\beta=X'y$ are solvable.
  • Gram-Schmidt: constructs orthonormal bases — the foundation of QR decomposition and orthogonal contrasts in DOE.
  • Eigenvalues & Cayley-Hamilton: heart of PCA, stationarity of Markov chains, stability of AR models, diagonalisation.
  • Symmetric / orthogonal matrices & quadratic forms: covariance matrices are symmetric PSD; rotations are orthogonal; quadratic forms appear in Hotelling's $T^2$, $\chi^2$ tests, ANOVA decompositions.

1. Finite, Countable and Uncountable Sets

Why this section? Probability spaces can be finite (a die), countably infinite (number of arrivals), or uncountable (a continuous time interval). Different machinery applies to each.

A set $A$ is finite if there is a bijection $A\leftrightarrow\{1,2,\ldots,n\}$ for some $n$. Countable if a bijection exists with $\mathbb{N}$. Uncountable if it is infinite but not countable.

Key Results

EXAMPLE 1 $\mathbb{Q}$ is countable: list rationals $p/q$ in increasing order of $|p|+q$, skipping repeats. Bijection with $\mathbb{N}$ exists.
EXAMPLE 2 Cantor's argument: assume $[0,1]$ is countable, list as $0.a_{11}a_{12}\ldots$, $0.a_{21}a_{22}\ldots$. Construct $b=0.b_1b_2\ldots$ where $b_k\ne a_{kk}$. Then $b\in[0,1]$ is not in the list — contradiction.

🌍 Where it's used in real life

  1. Knowing why discrete and continuous data need different tools.
  2. Computer science — what is and isn't computable.
  3. Database keys drawn from countable sets.
  4. Foundations of probability (measure theory).
  5. Digital vs analog signal representation.

2. Sequences of Real Numbers

A sequence $(x_n)$ converges to $L$ if $\forall\varepsilon>0,\exists N$ such that $|x_n-L|<\varepsilon$ for all $n\ge N.$

Key Concepts

Limit Algebra

$$\lim(x_n+y_n)=\lim x_n+\lim y_n,\quad \lim(x_n y_n)=(\lim x_n)(\lim y_n),\quad \lim\frac{x_n}{y_n}=\frac{\lim x_n}{\lim y_n}\text{ if }\lim y_n\ne 0.$$
EXAMPLE 1 $x_n=\left(1+\frac{1}{n}\right)^n$ — increasing and bounded above by 3, hence converges. Limit $=e\approx 2.71828.$
EXAMPLE 2 $x_n=\frac{(-1)^n}{n}$. Cauchy: $|x_m-x_n|\le 1/m+1/n\to 0$. Hence converges; limit $=0.$

🌍 Where it's used in real life

  1. Compound-interest balances year by year.
  2. Iterative algorithms homing in on a solution.
  3. Population growth over successive years.
  4. Drug concentration falling between doses.
  5. Successive approximations in engineering design.

3. Series of Real Numbers

$\sum_{n=1}^\infty a_n$ converges if partial sums $S_n=\sum_{k=1}^n a_k$ converge.

Tests of Convergence

p-series

$\sum 1/n^p$ converges iff $p>1.$
EXAMPLE 1 $\sum n!/n^n$. By ratio: $\frac{(n+1)!/(n+1)^{n+1}}{n!/n^n}=\frac{n^n}{(n+1)^n}=\frac{1}{(1+1/n)^n}\to 1/e<1$, converges.
EXAMPLE 2 $\sum (-1)^n/\sqrt n$ — converges by Leibniz, but $\sum 1/\sqrt n$ diverges (p-series, $p=1/2$). Conditionally convergent.

🌍 Where it's used in real life

  1. Present value of a loan or annuity (geometric series).
  2. Total distance of a bouncing ball.
  3. Calculators evaluating sin, exp, log (Taylor series).
  4. Signal reconstruction with Fourier series.
  5. Valuing a perpetuity in finance.

4. Power Series and Radius of Convergence

A power series $\sum_{n=0}^\infty a_n(x-x_0)^n$ has radius of convergence $R$ given by Cauchy–Hadamard: $$\frac{1}{R}=\limsup_{n\to\infty}|a_n|^{1/n}.$$

Series converges absolutely for $|x-x_0|<R$, diverges for $|x-x_0|>R$. Boundary $|x-x_0|=R$ requires separate analysis.

EXAMPLE 1 $\sum x^n/n!$: $|a_n|^{1/n}=(1/n!)^{1/n}\to 0\Rightarrow R=\infty$ (this is $e^x$).
EXAMPLE 2 $\sum n!\,x^n$: $|a_n|^{1/n}\to\infty\Rightarrow R=0$ — converges only at $x=0.$

🌍 Where it's used in real life

  1. Calculators and computers evaluating functions.
  2. Approximating hard functions in physics.
  3. Option-pricing expansions in finance.
  4. Solving differential equations by series.
  5. Error analysis in numerical methods.

5. Functions of a Real Variable

Limits

$\lim_{x\to a}f(x)=L$ if $\forall\varepsilon>0\,\exists\delta>0: 0<|x-a|<\delta\Rightarrow|f(x)-L|<\varepsilon.$

Continuity

$f$ continuous at $a$ if $\lim_{x\to a}f(x)=f(a).$

Uniform Continuity

$f$ is uniformly continuous on $E$ if $\forall\varepsilon>0\,\exists\delta>0: |x-y|<\delta\Rightarrow|f(x)-f(y)|<\varepsilon$ for all $x,y\in E$ (one $\delta$ works everywhere).
A continuous function on a closed bounded interval $[a,b]$ is uniformly continuous (Heine–Cantor) and attains its supremum and infimum (Extreme Value Theorem).

Monotone Functions

A monotone function on $[a,b]$ has at most countably many discontinuities (all jump-type).
EXAMPLE 1 $f(x)=1/x$ on $(0,1)$ is continuous but NOT uniformly continuous: for $x_n=1/n,\,y_n=1/(n+1)$, $|x_n-y_n|\to 0$ but $|f(x_n)-f(y_n)|=1.$
EXAMPLE 2 $f(x)=x^2$ on $\mathbb{R}$ is continuous but not uniformly continuous; restricted to $[0,M]$ it is uniformly continuous.

🌍 Where it's used in real life

  1. Modelling cost or revenue as a function of quantity.
  2. Position as a smooth function of time in physics.
  3. Sensor calibration curves.
  4. Smooth, jerk-free motion in robotics.
  5. Temperature changing continuously over a day.

6. Differentiability and Mean Value Theorems

$f$ is differentiable at $a$ if $f'(a)=\lim_{h\to 0}\frac{f(a+h)-f(a)}{h}$ exists.

Rolle's Theorem

If $f$ is continuous on $[a,b]$, differentiable on $(a,b)$, and $f(a)=f(b)$, then $\exists c\in(a,b): f'(c)=0.$

Lagrange's Mean Value Theorem

$\exists c\in(a,b): f(b)-f(a)=f'(c)(b-a).$

Cauchy's MVT

$\exists c\in(a,b): \frac{f(b)-f(a)}{g(b)-g(a)}=\frac{f'(c)}{g'(c)}$ (suitable $g$).

Taylor's Theorem

$$f(x)=\sum_{k=0}^{n}\frac{f^{(k)}(a)}{k!}(x-a)^k+R_n(x),\quad R_n(x)=\frac{f^{(n+1)}(\xi)}{(n+1)!}(x-a)^{n+1}.$$

L'Hôpital's Rule

For $0/0$ or $\infty/\infty$ forms, $\lim\frac{f}{g}=\lim\frac{f'}{g'}$ when the latter exists.
EXAMPLE 1 For $f(x)=x^3-x$ on $[-1,1]$, $f(-1)=f(1)=0$, Rolle gives $c=\pm 1/\sqrt3$ where $f'(c)=3c^2-1=0.$
EXAMPLE 2 $\lim_{x\to 0}\frac{\sin x}{x}\overset{L'H}{=}\lim_{x\to 0}\frac{\cos x}{1}=1.$ Or by Taylor $\sin x=x-x^3/6+\ldots$

🌍 Where it's used in real life

  1. Speed as the rate of change of distance.
  2. Marginal cost and revenue in economics.
  3. Maximising profit or minimising cost.
  4. Rate a drug leaves the body.
  5. Gradient descent training in machine learning.

7. Riemann Integration and Improper Integrals

A bounded $f$ on $[a,b]$ is Riemann integrable if $\sup L(P,f)=\inf U(P,f)$ over partitions $P$, denoted $\int_a^b f.$

Properties

Improper Integrals

EXAMPLE 1 $\int_0^1 x^2 dx=1/3$ via partitions or FTC.
EXAMPLE 2 $\int_1^\infty \frac{1}{x^p}dx$ converges iff $p>1$ (gives $1/(p-1)$). $\int_0^1 \frac{1}{x^p}dx$ converges iff $p<1.$

🌍 Where it's used in real life

  1. Distance travelled from a speed curve (area under it).
  2. Total revenue from a demand curve.
  3. Work done by a varying force.
  4. Total rainfall from an intensity curve.
  5. Probabilities as areas under a density curve.

8. Functions of Two Real Variables

Limits and Continuity

$\lim_{(x,y)\to(a,b)}f(x,y)=L$ if $\forall\varepsilon>0\,\exists\delta>0: \|(x,y)-(a,b)\|<\delta\Rightarrow|f-L|<\varepsilon.$

Partial and Total Derivative

Maxima/Minima

At interior critical point $(f_x=f_y=0)$, set $H=\det\begin{pmatrix}f_{xx}&f_{xy}\\f_{xy}&f_{yy}\end{pmatrix}=f_{xx}f_{yy}-f_{xy}^2.$

Lagrange Multipliers

To extremize $f(x,y)$ subject to $g(x,y)=0$: solve $\nabla f=\lambda\nabla g,\,g=0.$

Multiple Integrals

$$\iint_R f(x,y)\,dA,\qquad \iiint_V f(x,y,z)\,dV,$$ with change of variables: $\iint f(x,y)dx\,dy=\iint f(x(u,v),y(u,v))|J|\,du\,dv.$
EXAMPLE 1 $f(x,y)=x^2+y^2-2x-4y+5$. $f_x=2x-2=0,f_y=2y-4=0\Rightarrow(1,2).$ $H=4>0,\,f_{xx}=2>0\Rightarrow$ local minimum, value $0$.
EXAMPLE 2 Maximize $f=xy$ subject to $x+y=10$. $\nabla f=(y,x)=\lambda(1,1)\Rightarrow x=y=5$; max value $25.$

🌍 Where it's used in real life

  1. Profit that depends on price and advertising together.
  2. Least-cost mix of inputs in production (Lagrange).
  3. Heat maps and terrain surfaces.
  4. Return-vs-risk trade-offs for portfolios.
  5. Loss surfaces over two model parameters in ML.

9. Vector Spaces, Span, Linear Independence, Basis

A vector space $V$ over field $F$ is a set with operations of addition and scalar multiplication satisfying 8 axioms (associativity, commutativity, identity, inverses, distributivity).

Subspaces

$W\subseteq V$ is a subspace if it is closed under addition and scalar multiplication and contains $\mathbf 0$.

Span and Linear Independence

EXAMPLE 1 $\{(1,0,0),(0,1,0),(0,0,1)\}$ is the standard basis of $\mathbb{R}^3$, $\dim=3.$
EXAMPLE 2 $\{(1,1),(2,2)\}$ is linearly dependent in $\mathbb{R}^2$; spans only the line $y=x.$

🌍 Where it's used in real life

  1. 3D graphics and game engines.
  2. Documents as word vectors in text mining.
  3. Force and velocity vectors in physics.
  4. RGB colour space in imaging.
  5. Feature vectors in machine learning.

10. Rank, Nullity, Row Reduced Echelon Form

Row/Column Space

Row space = span of rows; column space = span of columns. $\dim(\text{row space})=\dim(\text{col space})=\text{rank}.$

Rank–Nullity Theorem

$$\text{rank}(A)+\text{nullity}(A)=\text{number of columns}.$$

RREF Procedure

Row operations to reach: each pivot is 1, only nonzero entry in its column.
EXAMPLE 1 $A=\begin{pmatrix}1&2&3\\2&4&6\\1&1&1\end{pmatrix}.$ R2 - 2R1, R3 - R1 gives $\begin{pmatrix}1&2&3\\0&0&0\\0&-1&-2\end{pmatrix}.$ Rank = 2, nullity = 1.
EXAMPLE 2 $A=\begin{pmatrix}1&0\\0&1\\1&1\end{pmatrix}.$ Columns linearly independent $\Rightarrow$ rank = 2, nullity = 0.

🌍 Where it's used in real life

  1. Checking whether equations have a solution.
  2. Spotting redundant sensors or variables.
  3. Low-rank image compression.
  4. Independent loops in circuit analysis.
  5. Detecting multicollinearity in regression.

11. Trace, Determinant, Inverse

Trace

$\text{tr}(A)=\sum a_{ii}.$ Properties: $\text{tr}(A+B)=\text{tr}A+\text{tr}B,\,\text{tr}(AB)=\text{tr}(BA),\,\text{tr}(P^{-1}AP)=\text{tr}(A).$ Also $\text{tr}(A)=\sum\lambda_i.$

Determinant

For $n\times n$ matrix: $\det(A)=\sum_{\sigma\in S_n}\text{sgn}(\sigma)\prod a_{i,\sigma(i)}.$

Inverse

$A^{-1}=\frac{1}{\det A}\,\text{adj}(A)$. $(AB)^{-1}=B^{-1}A^{-1}.$
EXAMPLE 1 $A=\begin{pmatrix}1&2\\3&4\end{pmatrix}$. $\det A=4-6=-2$. $A^{-1}=\frac{1}{-2}\begin{pmatrix}4&-2\\-3&1\end{pmatrix}=\begin{pmatrix}-2&1\\1.5&-0.5\end{pmatrix}.$
EXAMPLE 2 For diagonal $D=\text{diag}(d_1,\ldots,d_n)$, $\det D=\prod d_i,\,\text{tr}(D)=\sum d_i,\,D^{-1}=\text{diag}(1/d_i)$ when all nonzero.

🌍 Where it's used in real life

  1. Solving small linear systems (Cramer's rule).
  2. Area and volume scaling in graphics.
  3. Checking a transformation is invertible.
  4. Covariance-matrix determinant in statistics.
  5. Matrix-based encryption (Hill cipher).

12. Systems of Linear Equations

$Ax=b$ where $A$ is $m\times n$.

Cramer's Rule

For $n\times n$ nonsingular $A$: $x_i=\det(A_i)/\det(A)$, where $A_i$ replaces column $i$ of $A$ by $b$.
EXAMPLE 1 $x+y=3,\,2x+2y=6$. $\text{rank}(A)=\text{rank}([A|b])=1<2$ — infinite solutions; line $x+y=3$.
EXAMPLE 2 $x+y=2,\,x+y=3$. Inconsistent; no solution since $\text{rank}(A)=1<\text{rank}([A|b])=2.$

🌍 Where it's used in real life

  1. Balancing chemical equations.
  2. Traffic-flow and network analysis.
  3. Finding currents in electrical circuits.
  4. Input–output models in economics.
  5. Fitting a curve through data points.

13. Gram–Schmidt Orthogonalization

Given linearly independent $\{v_1,\ldots,v_k\}$ in inner product space, the orthogonal set $\{u_1,\ldots,u_k\}$ is constructed by: $$u_1=v_1,\quad u_j=v_j-\sum_{i=1}^{j-1}\frac{\langle v_j,u_i\rangle}{\langle u_i,u_i\rangle}u_i.$$ Normalize $e_j=u_j/\|u_j\|$ to get an orthonormal basis.
EXAMPLE 1 $v_1=(1,1,0),\,v_2=(1,0,1).$ $u_1=(1,1,0).$ $\langle v_2,u_1\rangle=1,\,\|u_1\|^2=2.$ $u_2=(1,0,1)-(1/2)(1,1,0)=(1/2,-1/2,1).$
EXAMPLE 2 In $\mathbb{R}^2$, $v_1=(3,4),\,v_2=(1,0)$. $u_1=(3,4)$, $u_2=(1,0)-\frac{3}{25}(3,4)=(16/25,-12/25)$. Normalising gives $e_1=(3/5,4/5),\,e_2=(4/5,-3/5).$

🌍 Where it's used in real life

  1. QR decomposition for stable regression.
  2. Orthogonal signal bases in communications.
  3. Removing correlation among predictors.
  4. Orthonormal camera frames in graphics.
  5. Building orthogonal polynomials for fitting.

14. Characteristic Roots and Vectors

$\lambda$ is a characteristic root (eigenvalue) of $A$ if $Av=\lambda v$ for some $v\ne 0$ (eigenvector). Solve $\det(A-\lambda I)=0$ — the characteristic polynomial.

Properties

EXAMPLE 1 $A=\begin{pmatrix}2&1\\1&2\end{pmatrix}.$ $\det(A-\lambda I)=(2-\lambda)^2-1=\lambda^2-4\lambda+3=0\Rightarrow\lambda=1,3.$ Eigenvectors $(1,-1)^T$ and $(1,1)^T.$
EXAMPLE 2 For $A=\text{diag}(5,-2,3)$, eigenvalues are 5, -2, 3 with standard basis vectors as eigenvectors.

🌍 Where it's used in real life

  1. Google PageRank ranking web pages.
  2. Principal component analysis for data reduction.
  3. Vibration and resonance modes of bridges.
  4. Stability of control systems.
  5. Energy levels in quantum mechanics.

15. Cayley–Hamilton Theorem

Every square matrix satisfies its own characteristic equation: if $p(\lambda)=\det(\lambda I-A)=\lambda^n+c_{n-1}\lambda^{n-1}+\cdots+c_0$, then $$A^n+c_{n-1}A^{n-1}+\cdots+c_0 I=0.$$

Useful for computing $A^{-1}$ and high powers of $A$.

EXAMPLE 1 $A=\begin{pmatrix}1&2\\3&4\end{pmatrix}$, $p(\lambda)=\lambda^2-5\lambda-2.$ So $A^2=5A+2I$ and $A^{-1}=\frac{1}{2}(A-5I).$
EXAMPLE 2 For $A=\begin{pmatrix}2&0\\0&3\end{pmatrix}$, $p(\lambda)=(\lambda-2)(\lambda-3).$ Verify $(A-2I)(A-3I)=\mathbf{0}.$

🌍 Where it's used in real life

  1. Computing high powers of a matrix efficiently.
  2. Finding matrix inverses in control systems.
  3. Solving systems of differential equations.
  4. Digital-filter design.
  5. Simplifying transition-matrix calculations.

16. Symmetric, Skew-symmetric, Orthogonal Matrices

TypeDefinitionEigenvalues
Symmetric$A=A^T$All real; orthogonal eigenvectors
Skew-symmetric$A=-A^T$Purely imaginary or zero; det is 0 if $n$ odd
Orthogonal$A^TA=I=AA^T$$|\lambda|=1$; $\det=\pm 1$
Idempotent$A^2=A$Eigenvalues 0 or 1; rank = trace
Nilpotent$A^k=0$ for some $k$All eigenvalues 0

Spectral Theorem

A real symmetric $A$ can be written $A=Q\Lambda Q^T$ with $Q$ orthogonal and $\Lambda$ diagonal of eigenvalues.
EXAMPLE 1 Rotation matrix $R(\theta)=\begin{pmatrix}\cos\theta&-\sin\theta\\\sin\theta&\cos\theta\end{pmatrix}$ is orthogonal with $\det=1.$
EXAMPLE 2 $A=\begin{pmatrix}0&1\\-1&0\end{pmatrix}$ skew-symmetric; eigenvalues $\pm i$; $\det=1.$

🌍 Where it's used in real life

  1. Rotations in graphics and robotics (orthogonal).
  2. Covariance matrices in statistics (symmetric).
  3. Stress and strain tensors in engineering.
  4. Whitening and PCA transforms in ML.
  5. Coordinate rotations in GPS.

17. Positive Definite Matrices and Quadratic Forms

Quadratic Form

$Q(x)=x^T A x$ where $A$ symmetric. Categories:

Properties of PD Matrices

Intuition. A positive-definite quadratic form $x^TAx$ is a genuine “bowl” opening upward, curving up in every direction with its minimum at the origin. This geometry is why definiteness matters throughout statistics: a covariance matrix must be PSD because “every variance is non-negative,” and a Hessian being PD certifies that a critical point is a local minimum.

Reduction to Canonical Form

Every quadratic form can be written as $\sum\lambda_i y_i^2$ via orthogonal transformation $x=Qy$ where $Q$ diagonalizes $A.$
EXAMPLE 1 $Q(x_1,x_2)=2x_1^2+2x_2^2+2x_1 x_2.$ Matrix $A=\begin{pmatrix}2&1\\1&2\end{pmatrix}$, eigenvalues 1,3 — all positive, so PD. Leading minors: 2 and 3, both positive.
EXAMPLE 2 $Q=x_1^2-x_2^2$ with $A=\text{diag}(1,-1).$ Eigenvalues $1,-1$ — indefinite (saddle).

🌍 Where it's used in real life

  1. Confirming a cost surface has a true minimum.
  2. Valid covariance matrices for portfolios.
  3. Stability (Lyapunov) analysis in control.
  4. Uniqueness of least-squares solutions.
  5. Energy functions in physics.