Fig 4.1 — Bell shape; symmetric; inflexion points at μ ± σ
2. Properties of the Normal Curve
Bell-shaped and symmetric about \(\mu\).
Mean = Median = Mode = \(\mu\).
The curve has its maximum at \(x = \mu\), value \(\dfrac{1}{\sigma \sqrt{2\pi}}\).
Asymptotic to the X-axis as \(|x - \mu| \to \infty\).
Total area under curve = 1.
Two points of inflexion at \(x = \mu \pm \sigma\) (where curvature changes).
About 68.27 % of area in \(\mu \pm \sigma\), 95.45 % in \(\mu \pm 2\sigma\), 99.73 % in \(\mu \pm 3\sigma\).
Quartile deviation = \(\dfrac{2}{3}\sigma\); Mean deviation = \(\dfrac{4}{5}\sigma\) (approximately).
Fig 4.2 — The empirical rule: about 68 % of the area lies within one
standard deviation of the mean, 95 % within two, and 99.7 % within
three. This is why \(\mu \pm 3\sigma\) is treated as the practical range of a normal variable — and the
basis of 3-sigma control limits.
3. Importance of the Normal Distribution
Many natural phenomena are approximately normal: heights, weights, IQ, errors of measurement.
Central Limit Theorem — sums and averages tend to normal regardless of original distribution.
Limiting form of Binomial, Poisson, NB and Gamma for large parameters.
Basis of most parametric inference: \(z\)-tests, \(t\)-tests, regression, ANOVA, control charts.
Two-parameter family — only mean & variance need to be specified.
4. MGF, CF, CGF
MGF
\[
M_X(t) \;=\; \exp\!\left(\mu t + \dfrac{\sigma^2 t^2}{2}\right).
\]
\[
\phi_X(t) = \exp\!\left(i\mu t - \dfrac{\sigma^2 t^2}{2}\right), \qquad
K_X(t) = \mu t + \dfrac{\sigma^2 t^2}{2}.
\]
From CGF, only \(k_1 = \mu\) and \(k_2 = \sigma^2\) are non-zero — all higher cumulants vanish.
5. Mean = Median = Mode
By symmetry of \(f\) about \(x = \mu\), the median and mode are also \(\mu\). At the mode, \(f'(\mu) = 0\) and \(f(\mu) = 1/(\sigma\sqrt{2\pi})\).
Proof. We use the moment generating function (MGF) and its uniqueness.
MGF of each variable. A normal variable \(N(\mu,\sigma^2)\) has MGF
\(M_X(t) = e^{\mu t + \sigma^2 t^2/2}\), so
\(M_{X_1}(t) = e^{\mu_1 t + \sigma_1^2 t^2/2}\) and \(M_{X_2}(t) = e^{\mu_2 t + \sigma_2^2 t^2/2}\).
MGF of the sum. By independence the MGF of the sum is the product of the MGFs:
\[ M_{X_1+X_2}(t) = M_{X_1}(t)\,M_{X_2}(t) = e^{\mu_1 t + \sigma_1^2 t^2/2}\cdot e^{\mu_2 t + \sigma_2^2 t^2/2}. \]
Add and regroup the exponents. Adding the exponents, then collecting the \(t\) terms and the
\(t^2\) terms separately,
\[ M_{X_1+X_2}(t) = e^{(\mu_1 t + \sigma_1^2 t^2/2) + (\mu_2 t + \sigma_2^2 t^2/2)}
= e^{(\mu_1+\mu_2)t + (\sigma_1^2+\sigma_2^2)t^2/2}. \]
Identify the distribution. This is the MGF of \(N(\mu_1+\mu_2,\ \sigma_1^2+\sigma_2^2)\).
By the uniqueness of the MGF, \(X_1+X_2 \sim N(\mu_1+\mu_2,\ \sigma_1^2+\sigma_2^2)\). \(\blacksquare\)
9. Linear Combination of Normal Variates
If \(X_i \sim N(\mu_i, \sigma_i^2)\) are independent and \(a_1, \ldots, a_n\) are constants, then
From \(f''(x) = 0\), the curvature of the normal PDF changes sign at
\[
x \;=\; \mu \pm \sigma.
\]
The curve is concave between \(\mu - \sigma\) and \(\mu + \sigma\), and convex outside.
11. Cauchy Distribution
The Cauchy distribution is bell-shaped like the normal but has such heavy tails
that none of its moments exist (not even the mean). With location \(\mu\) and scale
\(\lambda > 0\):
Its median and mode are \(\mu\); its quartiles are \(\mu \pm \lambda\). The standard Cauchy
(\(\mu=0,\lambda=1\)) arises as the ratio \(Z_1/Z_2\) of two independent standard normals, and equals
Student's \(t\) with 1 degree of freedom. Because the mean does not exist, the sample mean of Cauchy
data does not settle down — the CLT does not apply.
EXAMPLE
For the standard Cauchy, \(P(|X| < 1) = \dfrac{1}{\pi}\big[\arctan(1) - \arctan(-1)\big]
= \dfrac{1}{\pi}\cdot\dfrac{\pi}{2} = 0.5\): half the probability lies within one scale unit of the
centre.
Fig 11.1 — Both curves peak at the centre, but the Cauchy (red) sits lower and its tails stay well above the normal's, so extreme values are far more likely — the reason the Cauchy has no finite mean or variance. Dashed lines mark the quartiles at \(\mu\pm\lambda\).
12. Lognormal Distribution
A positive variable \(X\) is lognormal if \(\ln X \sim N(\mu, \sigma^{2})\). It
models positively-skewed quantities (incomes, share prices, particle sizes). Its density and moments
are
Because \(E(X) > \) median \(>\) mode, the distribution is right-skewed.
EXAMPLE
If \(\ln X \sim N(0, 0.25)\) (so \(\sigma = 0.5\)): median \(= e^{0} = 1\) and mean
\(= e^{0 + 0.125} = 1.133\) — the mean exceeds the median, confirming right skew.
Fig 12.1 — With \(\mu=0\), a larger \(\sigma\) makes the lognormal rise more steeply near 0 and stretch its right tail further. The median stays at \(e^{\mu}=1\) while the mean \(e^{\mu+\sigma^2/2}\) drifts rightward, producing the positive skew.
13. Bivariate Normal Distribution
The pair \((X, Y)\) is bivariate normal with means \(\mu_X, \mu_Y\), variances
\(\sigma_X^2, \sigma_Y^2\) and correlation \(\rho\) if
The conditional mean is the population regression line of \(Y\) on \(X\). When
\(X\) and \(Y\) are jointly (bivariate) normal, zero correlation implies independence
(\(\rho = 0 \Leftrightarrow X, Y\) independent). This needs the joint distribution to be normal: two variables
that are each normal on their own can be uncorrelated and still dependent.
EXAMPLE
With \(\mu_X=10,\mu_Y=20,\sigma_X=2,\sigma_Y=3,\rho=0.6\): at \(x = 12\),
\(E(Y\mid X=12) = 20 + 0.6\cdot\dfrac{3}{2}(12-10) = 21.8\) and
\(\text{Var}(Y\mid X=12) = 9(1-0.36) = 5.76\).
Fig 13.1 — Contours of equal density are concentric ellipses tilted by the correlation \(\rho=0.6\); the dashed red line is the regression \(E(Y\mid X)=\mu_Y+\rho\tfrac{\sigma_Y}{\sigma_X}(X-\mu_X)\). A larger \(|\rho|\) makes the ellipses narrower and more steeply tilted; \(\rho=0\) gives axis-aligned circles/ellipses (independence).
14. Worked Examples
EXAMPLE 1 (Probability via standardization)
The marks of 1000 students follow \(N(60, 100)\). Find percentage of students scoring (i) above 75, (ii) between 50 and 70.