Skip to the content

Topics Covered

PDF Properties MGF / CF / CGF Mean = Median = Mode Moments Skewness & Kurtosis Additive Property Linear Combination Points of Inflexion
On this page
  1. 1. Definition & PDF
  2. 2. Properties of the Normal Curve
  3. 3. Importance of the Normal Distribution
  4. 4. MGF, CF, CGF
  5. 5. Mean = Median = Mode
  6. 6. Moments about the Mean
  7. 7. Skewness and Kurtosis
  8. 8. Additive Property
  9. 9. Linear Combination of Normal Variates
  10. 10. Points of Inflexion
  11. 11. Cauchy Distribution
  12. 12. Lognormal Distribution
  13. 13. Bivariate Normal Distribution
  14. 14. Worked Examples
  15. Key Take-aways

1. Definition & PDF

DEFINITION

A continuous r.v. \(X\) follows a Normal distribution with mean \(\mu\) and variance \(\sigma^2\) if its PDF is

\[ f(x) \;=\; \dfrac{1}{\sigma \sqrt{2\pi}}\, \exp\!\left[-\dfrac{(x - \mu)^2}{2 \sigma^2}\right], \quad -\infty < x < \infty. \]

Notation: \(X \sim N(\mu, \sigma^2)\).

Mean = Median = Mode μ−σ μ μ+σ inflexion inflexion Normal Probability Curve
Fig 4.1 — Bell shape; symmetric; inflexion points at μ ± σ

2. Properties of the Normal Curve

  1. Bell-shaped and symmetric about \(\mu\).
  2. Mean = Median = Mode = \(\mu\).
  3. The curve has its maximum at \(x = \mu\), value \(\dfrac{1}{\sigma \sqrt{2\pi}}\).
  4. Asymptotic to the X-axis as \(|x - \mu| \to \infty\).
  5. Total area under curve = 1.
  6. Two points of inflexion at \(x = \mu \pm \sigma\) (where curvature changes).
  7. About 68.27 % of area in \(\mu \pm \sigma\), 95.45 % in \(\mu \pm 2\sigma\), 99.73 % in \(\mu \pm 3\sigma\).
  8. Quartile deviation = \(\dfrac{2}{3}\sigma\); Mean deviation = \(\dfrac{4}{5}\sigma\) (approximately).
The 68–95–99.7 (empirical) rule μ−3σμ−2σμ−σ μμ+σμ+2σμ+3σ 68.27% 95.45% within ±2σ  ·  99.73% within ±3σ
Fig 4.2 — The empirical rule: about 68 % of the area lies within one standard deviation of the mean, 95 % within two, and 99.7 % within three. This is why \(\mu \pm 3\sigma\) is treated as the practical range of a normal variable — and the basis of 3-sigma control limits.

3. Importance of the Normal Distribution

  1. Many natural phenomena are approximately normal: heights, weights, IQ, errors of measurement.
  2. Central Limit Theorem — sums and averages tend to normal regardless of original distribution.
  3. Limiting form of Binomial, Poisson, NB and Gamma for large parameters.
  4. Basis of most parametric inference: \(z\)-tests, \(t\)-tests, regression, ANOVA, control charts.
  5. Two-parameter family — only mean & variance need to be specified.

4. MGF, CF, CGF

MGF \[ M_X(t) \;=\; \exp\!\left(\mu t + \dfrac{\sigma^2 t^2}{2}\right). \]
\[ \phi_X(t) = \exp\!\left(i\mu t - \dfrac{\sigma^2 t^2}{2}\right), \qquad K_X(t) = \mu t + \dfrac{\sigma^2 t^2}{2}. \]

From CGF, only \(k_1 = \mu\) and \(k_2 = \sigma^2\) are non-zero — all higher cumulants vanish.

5. Mean = Median = Mode

By symmetry of \(f\) about \(x = \mu\), the median and mode are also \(\mu\). At the mode, \(f'(\mu) = 0\) and \(f(\mu) = 1/(\sigma\sqrt{2\pi})\).

6. Moments about the Mean

Even moments \[ \mu_{2r} \;=\; (2r - 1)(2r - 3) \cdots 3 \cdot 1 \cdot \sigma^{2r} \;=\; \dfrac{(2r)!}{2^r\, r!}\sigma^{2r}. \] Odd moments \[ \mu_{2r+1} \;=\; 0 \quad \text{for every } r \ge 0. \]

So \(\mu_2 = \sigma^2,\;\mu_4 = 3\sigma^4,\;\mu_6 = 15\sigma^6\), etc.

7. Skewness and Kurtosis

\[ \beta_1 = 0, \qquad \beta_2 = \dfrac{\mu_4}{\mu_2^2} = \dfrac{3\sigma^4}{\sigma^4} = 3. \] \[ \gamma_1 = 0, \qquad \gamma_2 = 0. \]

Normal is the benchmark for symmetric, mesokurtic shape.

8. Additive Property

If \(X_1 \sim N(\mu_1, \sigma_1^2)\) and \(X_2 \sim N(\mu_2, \sigma_2^2)\) are independent, then

\[ X_1 + X_2 \;\sim\; N(\mu_1 + \mu_2, \;\sigma_1^2 + \sigma_2^2). \]

Proof. We use the moment generating function (MGF) and its uniqueness.

  1. MGF of each variable. A normal variable \(N(\mu,\sigma^2)\) has MGF \(M_X(t) = e^{\mu t + \sigma^2 t^2/2}\), so \(M_{X_1}(t) = e^{\mu_1 t + \sigma_1^2 t^2/2}\) and \(M_{X_2}(t) = e^{\mu_2 t + \sigma_2^2 t^2/2}\).
  2. MGF of the sum. By independence the MGF of the sum is the product of the MGFs: \[ M_{X_1+X_2}(t) = M_{X_1}(t)\,M_{X_2}(t) = e^{\mu_1 t + \sigma_1^2 t^2/2}\cdot e^{\mu_2 t + \sigma_2^2 t^2/2}. \]
  3. Add and regroup the exponents. Adding the exponents, then collecting the \(t\) terms and the \(t^2\) terms separately, \[ M_{X_1+X_2}(t) = e^{(\mu_1 t + \sigma_1^2 t^2/2) + (\mu_2 t + \sigma_2^2 t^2/2)} = e^{(\mu_1+\mu_2)t + (\sigma_1^2+\sigma_2^2)t^2/2}. \]
  4. Identify the distribution. This is the MGF of \(N(\mu_1+\mu_2,\ \sigma_1^2+\sigma_2^2)\). By the uniqueness of the MGF, \(X_1+X_2 \sim N(\mu_1+\mu_2,\ \sigma_1^2+\sigma_2^2)\). \(\blacksquare\)

9. Linear Combination of Normal Variates

If \(X_i \sim N(\mu_i, \sigma_i^2)\) are independent and \(a_1, \ldots, a_n\) are constants, then

\[ Y \;=\; \sum a_i X_i \;\sim\; N\!\left(\sum a_i \mu_i,\; \sum a_i^2 \sigma_i^2\right). \]

10. Points of Inflexion

From \(f''(x) = 0\), the curvature of the normal PDF changes sign at

\[ x \;=\; \mu \pm \sigma. \]

The curve is concave between \(\mu - \sigma\) and \(\mu + \sigma\), and convex outside.

11. Cauchy Distribution

The Cauchy distribution is bell-shaped like the normal but has such heavy tails that none of its moments exist (not even the mean). With location \(\mu\) and scale \(\lambda > 0\):

\[ f(x) = \dfrac{1}{\pi\lambda\left[1 + \left(\dfrac{x-\mu}{\lambda}\right)^{2}\right]}, \quad -\infty < x < \infty. \]

Its median and mode are \(\mu\); its quartiles are \(\mu \pm \lambda\). The standard Cauchy (\(\mu=0,\lambda=1\)) arises as the ratio \(Z_1/Z_2\) of two independent standard normals, and equals Student's \(t\) with 1 degree of freedom. Because the mean does not exist, the sample mean of Cauchy data does not settle down — the CLT does not apply.

EXAMPLE

For the standard Cauchy, \(P(|X| < 1) = \dfrac{1}{\pi}\big[\arctan(1) - \arctan(-1)\big] = \dfrac{1}{\pi}\cdot\dfrac{\pi}{2} = 0.5\): half the probability lies within one scale unit of the centre.

Cauchy vs. standard normal -5 -4 -3 -2 -1 0 1 2 3 4 5 quartiles μ±λ heavy tails Normal N(0,1) Cauchy(0,1)
Fig 11.1 — Both curves peak at the centre, but the Cauchy (red) sits lower and its tails stay well above the normal's, so extreme values are far more likely — the reason the Cauchy has no finite mean or variance. Dashed lines mark the quartiles at \(\mu\pm\lambda\).

12. Lognormal Distribution

A positive variable \(X\) is lognormal if \(\ln X \sim N(\mu, \sigma^{2})\). It models positively-skewed quantities (incomes, share prices, particle sizes). Its density and moments are

\[ f(x) = \dfrac{1}{x\sigma\sqrt{2\pi}}\,\exp\!\left[-\dfrac{(\ln x - \mu)^2}{2\sigma^2}\right],\ x>0; \qquad E(X) = e^{\mu + \sigma^2/2}, \quad \text{median} = e^{\mu}. \]

Because \(E(X) > \) median \(>\) mode, the distribution is right-skewed.

EXAMPLE

If \(\ln X \sim N(0, 0.25)\) (so \(\sigma = 0.5\)): median \(= e^{0} = 1\) and mean \(= e^{0 + 0.125} = 1.133\) — the mean exceeds the median, confirming right skew.

Lognormal density (μ = 0): right-skew grows with σ 0 1 2 3 4 5 median e^μ=1 σ = 0.25 σ = 0.5 σ = 1.0
Fig 12.1 — With \(\mu=0\), a larger \(\sigma\) makes the lognormal rise more steeply near 0 and stretch its right tail further. The median stays at \(e^{\mu}=1\) while the mean \(e^{\mu+\sigma^2/2}\) drifts rightward, producing the positive skew.

13. Bivariate Normal Distribution

The pair \((X, Y)\) is bivariate normal with means \(\mu_X, \mu_Y\), variances \(\sigma_X^2, \sigma_Y^2\) and correlation \(\rho\) if

\[ f(x,y) = \dfrac{1}{2\pi\sigma_X\sigma_Y\sqrt{1-\rho^2}} \exp\!\left\{-\dfrac{1}{2(1-\rho^2)}\!\left[\dfrac{(x-\mu_X)^2}{\sigma_X^2} - \dfrac{2\rho(x-\mu_X)(y-\mu_Y)}{\sigma_X\sigma_Y} + \dfrac{(y-\mu_Y)^2}{\sigma_Y^2}\right]\right\}. \]

Key properties: both marginals are normal; each conditional distribution is normal with

\[ E(Y\mid X=x) = \mu_Y + \rho\dfrac{\sigma_Y}{\sigma_X}(x - \mu_X), \qquad \text{Var}(Y\mid X=x) = \sigma_Y^2(1 - \rho^2). \]

The conditional mean is the population regression line of \(Y\) on \(X\). When \(X\) and \(Y\) are jointly (bivariate) normal, zero correlation implies independence (\(\rho = 0 \Leftrightarrow X, Y\) independent). This needs the joint distribution to be normal: two variables that are each normal on their own can be uncorrelated and still dependent.

EXAMPLE

With \(\mu_X=10,\mu_Y=20,\sigma_X=2,\sigma_Y=3,\rho=0.6\): at \(x = 12\), \(E(Y\mid X=12) = 20 + 0.6\cdot\dfrac{3}{2}(12-10) = 21.8\) and \(\text{Var}(Y\mid X=12) = 9(1-0.36) = 5.76\).

Bivariate normal contours (ρ = 0.6) 4 6 8 10 12 14 16 8 12 16 20 24 28 32 X Y (μ_X, μ_Y) E(Y|X) regression line
Fig 13.1 — Contours of equal density are concentric ellipses tilted by the correlation \(\rho=0.6\); the dashed red line is the regression \(E(Y\mid X)=\mu_Y+\rho\tfrac{\sigma_Y}{\sigma_X}(X-\mu_X)\). A larger \(|\rho|\) makes the ellipses narrower and more steeply tilted; \(\rho=0\) gives axis-aligned circles/ellipses (independence).

14. Worked Examples

EXAMPLE 1 (Probability via standardization)

The marks of 1000 students follow \(N(60, 100)\). Find percentage of students scoring (i) above 75, (ii) between 50 and 70.

(i) \(z = (75 - 60)/10 = 1.5;\;\) \(P(Z > 1.5) = 0.0668\) ≈ 6.68 %.

(ii) \(z_1 = -1, z_2 = 1;\;\) \(P(-1 < Z < 1) = 0.6827\) ≈ 68.27 %.

EXAMPLE 2 (Linear combination)

Heights of men \(\sim N(170, 36)\) and of women \(\sim N(160, 25)\). Find distribution of the difference (man − woman) of one randomly chosen pair.

Difference \(\sim N(170 - 160, 36 + 25) = N(10, 61)\). SD \(= \sqrt{61} \approx 7.81\).

Key Take-aways