Skip to the content

Topics Covered

Review of Basic Distributions Lognormal Weibull & Hazard Rate Pareto Laplace Cauchy
On this page
  1. 1. Brief Review of the Basic Distributions
  2. 2. The Lognormal Distribution
  3. 3. The Weibull Distribution
  4. 4. The Pareto Distribution
  5. 5. The Laplace Distribution
  6. 6. The Cauchy Distribution
  7. 7. The Five Distributions Side by Side
  8. Key Take-aways
Where this unit starts. The syllabus names a long pre-requisite — the discrete uniform, Bernoulli, binomial, Poisson, negative binomial, geometric and hypergeometric distributions, and the continuous uniform, normal, exponential, gamma, both kinds of beta and the Cauchy — together with univariate and bivariate transformations. All of it is taught in the Foundation notes and is not re-derived here: Theoretical Discrete Distributions covers the first seven, Theoretical Continuous Distributions the rest, and Theory of Probability, Unit 3 the bivariate transformations. This unit adds the five distributions that are new at postgraduate level.

1. Brief Review of the Basic Distributions

WHAT IS ASSUMED, IN ONE TABLE

Each row links the page where the distribution is derived in full. Nothing in this table is proved again on this page.

DistributionMeanVarianceDerived in
Binomial \((n,p)\)\(np\)\(npq\) Discrete, Unit 1
Poisson \((\lambda)\)\(\lambda\)\(\lambda\) Discrete, Unit 2
Negative binomial\(rq/p\)\(rq/p^{2}\) Discrete, Unit 3
Geometric \((p)\)\(q/p\)\(q/p^{2}\) Discrete, Unit 4
Hypergeometric\(nK/N\)\(n\frac{K}{N}\frac{N-K}{N}\frac{N-n}{N-1}\) Discrete, Unit 5
Uniform \((a,b)\)\(\frac{a+b}{2}\)\(\frac{(b-a)^{2}}{12}\) Continuous, Unit 1
Exponential \((\theta)\)\(1/\theta\)\(1/\theta^{2}\) Continuous, Unit 2
Gamma, Betavia \(\Gamma(\cdot)\) and \(B(\cdot,\cdot)\) Continuous, Unit 3
Normal \((\mu,\sigma^{2})\)\(\mu\)\(\sigma^{2}\) Continuous, Unit 4

2. The Lognormal Distribution

DEFINITION

\(X\) is lognormal with parameters \(\mu\) and \(\sigma^{2}\), written \(X \sim LN(\mu, \sigma^{2})\), if \(Y = \ln X \sim N(\mu, \sigma^{2})\). Equivalently \(X = e^{Y}\). Its density on \(x > 0\) is

\[ f(x) = \frac{1}{x \sigma \sqrt{2\pi}} \exp\left\{ -\frac{(\ln x - \mu)^{2}}{2\sigma^{2}} \right\}. \]

Note carefully: \(\mu\) and \(\sigma^{2}\) are the mean and variance of \(\ln X\), not of \(X\). That is the commonest error with this distribution.

DERIVING THE MOMENTS FROM THE NORMAL MGF

Statement. \(E(X^{r}) = \exp\left(r\mu + \tfrac12 r^{2}\sigma^{2}\right)\).

Step 1 — turn the moment into an MGF. With \(Y = \ln X\) we have \(X = e^{Y}\), so \(X^{r} = e^{rY}\) and

\[ E(X^{r}) = E\left(e^{rY}\right) = M_Y(r), \]

the moment generating function of \(Y\) evaluated at \(r\).

Step 2 — substitute the normal MGF. For \(Y \sim N(\mu, \sigma^{2})\), \(M_Y(t) = \exp\left(\mu t + \tfrac12 \sigma^{2} t^{2}\right)\), a result derived in Continuous Distributions, Unit 4. Putting \(t = r\),

\[ E(X^{r}) = \exp\left(r\mu + \tfrac12 r^{2}\sigma^{2}\right). \]

Step 3 — the first two moments. At \(r = 1\) and \(r = 2\),

\[ E(X) = e^{\mu + \sigma^{2}/2}, \qquad E(X^{2}) = e^{2\mu + 2\sigma^{2}}. \]

Step 4 — the variance.

\[ \operatorname{Var}(X) = e^{2\mu + 2\sigma^{2}} - \left(e^{\mu + \sigma^{2}/2}\right)^{2} = e^{2\mu + 2\sigma^{2}} - e^{2\mu + \sigma^{2}} = \left(e^{\sigma^{2}} - 1\right) e^{2\mu + \sigma^{2}}. \quad \blacksquare \]

No MGF of its own. The lognormal has finite moments of every order, yet \(E(e^{tX}) = \infty\) for every \(t > 0\). It is the standard example showing that a sequence of moments need not determine a distribution, and that a distribution can have all moments and still have no moment generating function.

THE THREE CENTRES, AND WHY THEY DIFFER \[ \text{mode} = e^{\mu - \sigma^{2}}, \qquad \text{median} = e^{\mu}, \qquad \text{mean} = e^{\mu + \sigma^{2}/2}. \]

Since \(\sigma^{2} > 0\), these are always in the order mode < median < mean, which is the signature of right skewness. The median is \(e^{\mu}\) because \(\ln X\) is symmetric about \(\mu\) and the exponential is increasing, so it carries the median across unchanged.

EXAMPLE 1.1 — A LOGNORMAL INCOME DISTRIBUTION

Given. Annual income \(X\) is modelled as \(LN(\mu = 10, \sigma = 0.5)\), in currency units.

Asked. Find the mode, median, mean and standard deviation, and \(P(X > 30{,}000)\).

Step 1 — median. \( \text{median} = e^{\mu} = e^{10} = 22{,}026.47.\)

Step 2 — mean. \(\mu + \sigma^{2}/2 = 10 + 0.25/2 = 10.125\), so

\[ E(X) = e^{10.125} = 24{,}959.26. \]

Step 3 — mode. \(\mu - \sigma^{2} = 10 - 0.25 = 9.75\), so

\[ \text{mode} = e^{9.75} = 17{,}154.23. \]

Step 4 — check the ordering. \(17{,}154.23 < 22{,}026.47 < 24{,}959.26\), so mode < median < mean as the formulae require. \(\checkmark\)

Step 5 — variance and standard deviation.

\[ \operatorname{Var}(X) = \left(e^{0.25} - 1\right) e^{20 + 0.25} = (1.284025 - 1)\, e^{20.25} = 176{,}937{,}735.28, \] \[ \text{s.d.} = \sqrt{176{,}937{,}735.28} = 13{,}301.79. \]

Step 6 — the tail probability. Work on the log scale, where the distribution is normal:

\[ P(X > 30{,}000) = P(\ln X > \ln 30{,}000) = P\left(Z > \frac{10.308953 - 10}{0.5}\right) = P(Z > 0.6179), \] \[ = 1 - \Phi(0.6179) = 1 - 0.731681 = 0.268319. \]

Interpretation. About \(26.8\%\) of incomes exceed \(30{,}000\), yet the median is only \(22{,}026\) — half the population is below that. The mean exceeds the median by the factor \(e^{\sigma^{2}/2} = 1.133148\), so quoting the mean overstates the typical income by more than \(13\%\). That gap is why income and wealth are reported by median rather than mean, and it is the lognormal's defining feature rather than an artefact of these particular numbers.

Lognormal density of Example 1.1, with its three centres mode 17,154 median 22,026 mean 24,959 mean exceeds median by 13.3% 0 20,000 40,000 60,000 the long right tail is what pulls the mean above the median vertical scale is density, peak value 4.31 × 10⁻⁵
Fig 1.1 — Every point of the curve and all three markers are computed from the density, not sketched.

3. The Weibull Distribution

DEFINITION AND THE HAZARD RATE

\(X \sim\) Weibull\((k, \lambda)\) with shape \(k > 0\) and scale \(\lambda > 0\) has, for \(x \ge 0\),

\[ f(x) = \frac{k}{\lambda}\left(\frac{x}{\lambda}\right)^{k-1} e^{-(x/\lambda)^{k}}, \qquad F(x) = 1 - e^{-(x/\lambda)^{k}}, \qquad R(x) = e^{-(x/\lambda)^{k}}, \]

where \(R = 1 - F\) is the reliability or survival function. The hazard rate is

\[ h(x) = \frac{f(x)}{R(x)} = \frac{k}{\lambda^{k}} x^{k-1}, \]

obtained by dividing the density by the survival function: the two exponential factors cancel exactly, which is what makes this distribution so convenient in reliability work.

WHAT THE SHAPE PARAMETER MEANS \[ \begin{aligned} k < 1 :\quad & h(x) \text{ decreasing} && \text{infant mortality — early failures dominate} \\ k = 1 :\quad & h(x) = 1/\lambda \text{ constant} && \text{the exponential distribution; memoryless} \\ k > 1 :\quad & h(x) \text{ increasing} && \text{wear-out — failure becomes more likely with age} \end{aligned} \]

One family therefore covers all three regions of the classical bathtub curve, and \(k = 1\) recovers the exponential of Continuous Distributions, Unit 2 exactly.

MOMENTS, BY SUBSTITUTION INTO THE GAMMA INTEGRAL

Statement. \(E(X^{r}) = \lambda^{r}\,\Gamma\!\left(1 + \dfrac{r}{k}\right)\).

Step 1 — write the integral.

\[ E(X^{r}) = \int_{0}^{\infty} x^{r} \frac{k}{\lambda}\left(\frac{x}{\lambda}\right)^{k-1} e^{-(x/\lambda)^{k}} dx. \]

Step 2 — substitute \(u = (x/\lambda)^{k}\). Then \(x = \lambda u^{1/k}\) and, differentiating, \(du = \dfrac{k}{\lambda}\left(\dfrac{x}{\lambda}\right)^{k-1} dx\) — which is exactly the factor already standing beside \(dx\) in the integral. The limits carry over: \(x = 0\) gives \(u = 0\), and \(x \to \infty\) gives \(u \to \infty\). So

\[ E(X^{r}) = \int_{0}^{\infty} \left(\lambda u^{1/k}\right)^{r} e^{-u} \, du = \lambda^{r} \int_{0}^{\infty} u^{r/k} e^{-u} \, du. \]

Step 3 — recognise the gamma integral. By definition \(\Gamma(s) = \int_0^\infty u^{s-1} e^{-u} du\). Matching exponents, \(s - 1 = r/k\), so \(s = 1 + r/k\) and

\[ E(X^{r}) = \lambda^{r} \, \Gamma\!\left(1 + \frac{r}{k}\right). \quad \blacksquare \]
EXAMPLE 1.2 — A COMPONENT WITH WEAR-OUT FAILURE

Given. Time to failure, in hours, is Weibull with \(k = 2\) and \(\lambda = 1000\).

Asked. Find the reliability at 500 hours, the mean and standard deviation of life, the median life, and compare the hazard at 500 and 1500 hours.

Step 1 — reliability at 500 hours.

\[ R(500) = e^{-(500/1000)^{2}} = e^{-(0.5)^{2}} = e^{-0.25} = 0.778801. \]

So \(77.88\%\) of components survive to 500 hours, and \(1 - 0.778801 = 0.221199\), about \(22.12\%\), fail before then.

Step 2 — the mean. Using the formula with \(r = 1\) and \(k = 2\),

\[ E(X) = 1000 \times \Gamma\!\left(1 + \tfrac12\right) = 1000 \times \Gamma(1.5). \]

Now \(\Gamma(1.5) = \tfrac12 \Gamma(0.5) = \tfrac12 \sqrt{\pi} = 0.8862269\), using \(\Gamma(s+1) = s\,\Gamma(s)\) and \(\Gamma(1/2) = \sqrt{\pi}\). Hence

\[ E(X) = 886.23 \text{ hours}. \]

Step 3 — the second moment and the variance. With \(r = 2\),

\[ E(X^{2}) = 1000^{2}\,\Gamma(1 + 1) = 10^{6} \times \Gamma(2) = 10^{6} \times 1 = 10^{6}, \] \[ \operatorname{Var}(X) = 10^{6} - (886.2269)^{2} = 10^{6}\left(1 - \tfrac{\pi}{4}\right) = 10^{6} (1 - 0.7853982) = 214{,}601.84, \]

where \(\Gamma(1.5)^{2} = (\sqrt{\pi}/2)^{2} = \pi/4\) exactly. So

\[ \text{s.d.} = \sqrt{214{,}601.84} = 463.25 \text{ hours}. \]

Step 4 — the median. Solve \(R(m) = 0.5\):

\[ e^{-(m/1000)^{2}} = 0.5 \;\Rightarrow\; \left(\frac{m}{1000}\right)^{2} = \ln 2 \;\Rightarrow\; m = 1000\sqrt{\ln 2} = 1000\sqrt{0.693147} = 832.55 \text{ hours}. \]

Median \(<\) mean, so this Weibull is right-skewed too.

Step 5 — the hazard, then and later. With \(k = 2\), \(h(x) = 2x/10^{6}\):

\[ h(500) = \frac{2 \times 500}{10^{6}} = 0.001000 \text{ per hour}, \qquad h(1500) = \frac{2 \times 1500}{10^{6}} = 0.003000 \text{ per hour}. \]

Interpretation. The instantaneous failure rate trebles between 500 and 1500 hours, which is what \(k = 2 > 1\) means: the component wears out. An exponential model would have forced a constant rate and therefore predicted that a 1500-hour-old unit is as good as new — the memoryless property — which for a mechanical part is plainly false. Choosing \(k\) is choosing whether ageing is allowed into the model at all.

4. The Pareto Distribution

DEFINITION

\(X \sim\) Pareto\((\alpha, x_m)\) with shape \(\alpha > 0\) and minimum \(x_m > 0\) has, for \(x \ge x_m\),

\[ f(x) = \frac{\alpha x_m^{\alpha}}{x^{\alpha+1}}, \qquad F(x) = 1 - \left(\frac{x_m}{x}\right)^{\alpha}, \qquad P(X > x) = \left(\frac{x_m}{x}\right)^{\alpha}. \]

The survival function is a pure power of \(x\), which is why the Pareto is called a power-law or heavy-tailed distribution: its tail decays polynomially, not exponentially.

MOMENTS, AND EXACTLY WHEN THEY EXIST

Statement. \(E(X^{r})\) is finite if and only if \(r < \alpha\), and then \(E(X^{r}) = \dfrac{\alpha x_m^{r}}{\alpha - r}\).

Step 1 — set up the integral.

\[ E(X^{r}) = \int_{x_m}^{\infty} x^{r} \frac{\alpha x_m^{\alpha}}{x^{\alpha+1}} \, dx = \alpha x_m^{\alpha} \int_{x_m}^{\infty} x^{r - \alpha - 1} \, dx. \]

Step 2 — decide convergence before integrating. The integral \(\int_{x_m}^{\infty} x^{p} dx\) converges at the upper limit precisely when \(p < -1\). Here \(p = r - \alpha - 1\), so the condition is \(r - \alpha - 1 < -1\), that is \(r < \alpha\). If \(r \ge \alpha\) the integral diverges and the moment does not exist.

Step 3 — integrate when it does converge.

\[ \int_{x_m}^{\infty} x^{r-\alpha-1} dx = \left[\frac{x^{r-\alpha}}{r - \alpha}\right]_{x_m}^{\infty} = 0 - \frac{x_m^{r-\alpha}}{r-\alpha} = \frac{x_m^{r-\alpha}}{\alpha - r}, \]

the upper limit vanishing because \(r - \alpha < 0\) makes \(x^{r-\alpha} \to 0\).

Step 4 — assemble.

\[ E(X^{r}) = \alpha x_m^{\alpha} \cdot \frac{x_m^{r-\alpha}}{\alpha - r} = \frac{\alpha x_m^{r}}{\alpha - r}. \quad \blacksquare \]

Consequences. The mean exists only for \(\alpha > 1\); the variance only for \(\alpha > 2\). A fitted \(\alpha\) below 2 means the sample variance is estimating something that does not exist, and it will fail to settle however large the sample — a direct consequence of the strong law's necessity condition in Probability Theory, Unit 4.

EXAMPLE 1.3 — A PARETO CLAIM-SIZE MODEL

Given. Insurance claims above a retention of \(x_m = 20{,}000\) follow a Pareto with \(\alpha = 2.5\).

Asked. Find the mean, the standard deviation, the median and \(P(X > 50{,}000)\).

Step 1 — check what exists. \(\alpha = 2.5 > 2\), so both the mean and the variance are finite. The third moment (\(r = 3 > 2.5\)) is not, so skewness is undefined here — worth knowing before anyone reports one.

Step 2 — the mean.

\[ E(X) = \frac{\alpha x_m}{\alpha - 1} = \frac{2.5 \times 20{,}000}{1.5} = \frac{50{,}000}{1.5} = 33{,}333.33. \]

Step 3 — the second moment and variance.

\[ E(X^{2}) = \frac{\alpha x_m^{2}}{\alpha - 2} = \frac{2.5 \times 4 \times 10^{8}}{0.5} = 2 \times 10^{9}, \] \[ \operatorname{Var}(X) = 2 \times 10^{9} - (33{,}333.33)^{2} = 2 \times 10^{9} - 1.111111 \times 10^{9} = 888{,}888{,}888.89, \] \[ \text{s.d.} = \sqrt{888{,}888{,}888.89} = 29{,}814.24. \]

Step 4 — the median. Solve \(F(m) = 0.5\):

\[ \left(\frac{20{,}000}{m}\right)^{2.5} = 0.5 \;\Rightarrow\; m = 20{,}000 \times 2^{1/2.5} = 20{,}000 \times 2^{0.4} = 26{,}390.16. \]

Step 5 — the tail probability.

\[ P(X > 50{,}000) = \left(\frac{20{,}000}{50{,}000}\right)^{2.5} = (0.4)^{2.5}. \]

Evaluate through logarithms: \(\ln 0.4 = -0.9162907\), so \(2.5 \ln 0.4 = -2.2907268\) and

\[ (0.4)^{2.5} = e^{-2.2907268} = 0.101193. \]

Interpretation. The standard deviation, \(29{,}814\), is almost as large as the mean itself — a coefficient of variation of \(0.894\). Just over one claim in ten exceeds two and a half times the retention. This is what heavy-tailed means in practice: the average is a poor summary because a small number of very large claims carry much of the total, which is why reinsurance is priced from the tail rather than from the mean.

5. The Laplace Distribution

DEFINITION AND MOMENTS

\(X \sim\) Laplace\((\mu, b)\), the double exponential, has density on all of \(\mathbb{R}\)

\[ f(x) = \frac{1}{2b} \exp\left(-\frac{|x - \mu|}{b}\right), \qquad b > 0. \]

It is two exponential densities placed back to back at \(\mu\), each carrying half the probability — which is exactly what the factor \(\tfrac12\) is doing.

Mean. The density is symmetric about \(\mu\), so \(E(X) = \mu\) whenever the mean exists; and it does, since \(\int |x| e^{-|x|/b} dx\) converges.

Variance. Write \(Z = (X - \mu)/b\), which has density \(\tfrac12 e^{-|z|}\). Then

\[ E(Z^{2}) = \int_{-\infty}^{\infty} z^{2} \tfrac12 e^{-|z|} dz = 2 \int_{0}^{\infty} z^{2} \tfrac12 e^{-z} dz = \int_{0}^{\infty} z^{2} e^{-z} dz = \Gamma(3) = 2! = 2, \]

the factor 2 coming from symmetry of \(z^{2}e^{-|z|}\) about zero. Hence

\[ \operatorname{Var}(X) = b^{2} E(Z^{2}) = 2b^{2}. \]

Characteristic function. For \(\mu = 0, b = 1\), \(\phi_X(t) = \dfrac{1}{1 + t^{2}}\) — the function identified as a characteristic function in Probability Theory, Unit 2, Example 2.3(b).

Why it is used. Its tails are heavier than the normal's (decaying like \(e^{-|x|}\) rather than \(e^{-x^{2}}\)), so it tolerates outliers. Maximising a Laplace likelihood is equivalent to minimising \(\sum |x_i - \mu|\), which is why the median is the maximum likelihood estimator under a Laplace model, just as the mean is under a normal one.

6. The Cauchy Distribution

DEFINITION

\(X \sim\) Cauchy\((\theta, \lambda)\) has density

\[ f(x) = \frac{1}{\pi \lambda \left[1 + \left(\frac{x - \theta}{\lambda}\right)^{2}\right]}, \qquad x \in \mathbb{R}, \]

with \(\theta\) the location (its median and mode) and \(\lambda > 0\) the scale. The standard case \(\theta = 0, \lambda = 1\) gives \(f(x) = \dfrac{1}{\pi(1 + x^{2})}\), and is met in Continuous Distributions, Unit 4 as the ratio of two independent standard normals.

PROOF THAT THE MEAN DOES NOT EXIST

Statement. For the standard Cauchy, \(E|X| = \infty\), so \(E(X)\) does not exist — and this is a stronger statement than "the integral is symmetric so it is zero".

Step 1 — write the absolute first moment.

\[ E|X| = \int_{-\infty}^{\infty} \frac{|x|}{\pi(1 + x^{2})} dx = \frac{2}{\pi} \int_{0}^{\infty} \frac{x}{1 + x^{2}} dx, \]

using symmetry of the integrand about zero.

Step 2 — integrate. Substitute \(u = 1 + x^{2}\), so \(du = 2x \, dx\) and \(x \, dx = \tfrac12 du\):

\[ \int_{0}^{\infty} \frac{x}{1+x^{2}} dx = \frac{1}{2}\int_{1}^{\infty} \frac{du}{u} = \frac{1}{2}\Big[\ln u\Big]_{1}^{\infty} = \infty. \]

Step 3 — conclude. \(E|X| = \infty\). The definition of expectation requires \(E|X| < \infty\), so \(E(X)\) is undefined, not zero. \(\blacksquare\)

Step 4 — why symmetry is not enough. One might argue that the two halves cancel. They do not, because \(\infty - \infty\) is not defined; the principal value exists and equals zero, but the expectation does not. The practical consequence is decisive: by Probability Theory, Unit 4, \(\bar{X}_n\) for a Cauchy sample has exactly the same distribution as a single observation, for every \(n\). Averaging accomplishes nothing at all.

7. The Five Distributions Side by Side

DistributionSupportMeanVarianceTypical use
Lognormal \((\mu,\sigma^{2})\)\(x > 0\) \(e^{\mu+\sigma^{2}/2}\)\(\left(e^{\sigma^{2}}-1\right)e^{2\mu+\sigma^{2}}\) incomes, particle sizes, anything multiplicative
Weibull \((k,\lambda)\)\(x \ge 0\) \(\lambda\Gamma(1+1/k)\) \(\lambda^{2}\left[\Gamma(1+2/k) - \Gamma(1+1/k)^{2}\right]\) lifetimes, reliability, wind speeds
Pareto \((\alpha,x_m)\)\(x \ge x_m\) \(\frac{\alpha x_m}{\alpha-1}\), \(\alpha > 1\) \(\frac{\alpha x_m^{2}}{(\alpha-1)^{2}(\alpha-2)}\), \(\alpha > 2\) claim sizes, city sizes, wealth
Laplace \((\mu,b)\)\(\mathbb{R}\) \(\mu\)\(2b^{2}\) robust models, errors with outliers
Cauchy \((\theta,\lambda)\)\(\mathbb{R}\) does not existdoes not exist ratios of normals; a standard counterexample

Key Take-aways