Skip to the content

Topics Covered

Jacobian Transformations Truncated Distributions Mixtures Exponential Family Power Series Family Compound Distributions
On this page
  1. 1. Functions of Random Variables: the Jacobian Rule
  2. 2. Truncated Distributions
  3. 3. Mixture Distributions
  4. 4. The Exponential Family
  5. 5. The Power Series Family
  6. 6. Compound Distributions
  7. Key Take-aways
Where this unit starts. The change-of-variable rule for a single random variable, and the bivariate case with its Jacobian, are set out in Theory of Probability, Unit 3. That rule is restated here in one box and then used; what is new is everything built on it — truncation, mixing, the two exponential-type families, and compounding.

1. Functions of Random Variables: the Jacobian Rule

THE RULE

Let \((X, Y)\) have joint density \(f_{X,Y}\) and let \(u = u(x,y)\), \(v = v(x,y)\) be a one-to-one transformation with inverse \(x = x(u,v)\), \(y = y(u,v)\). Then

\[ f_{U,V}(u,v) = f_{X,Y}\big(x(u,v),\, y(u,v)\big) \, |J|, \qquad J = \begin{vmatrix} \dfrac{\partial x}{\partial u} & \dfrac{\partial x}{\partial v} \\ \dfrac{\partial y}{\partial u} & \dfrac{\partial y}{\partial v} \end{vmatrix}. \]

Three things decide whether an answer is right: the inverse must be solved correctly, the Jacobian must be the determinant of the derivatives of the old variables with respect to the new ones, and the support must be transformed too. The third is the one most often dropped, and it is usually where the marks go.

EXAMPLE 2.1 — A SUM AND A RATIO FROM TWO EXPONENTIALS

Given. \(X\) and \(Y\) independent, each Exponential\((1)\), so

\[ f_{X,Y}(x,y) = e^{-x} e^{-y} = e^{-(x+y)}, \qquad x > 0,\; y > 0. \]

Asked. Find the joint density of \(U = X + Y\) and \(V = \dfrac{X}{X+Y}\), and identify the marginals.

Step 1 — invert the transformation. From \(V = X/U\) we get \(X = UV\), and then \(Y = U - X = U(1 - V)\). So

\[ x = uv, \qquad y = u(1-v). \]

Step 2 — the Jacobian. Differentiate the inverse:

\[ \frac{\partial x}{\partial u} = v, \quad \frac{\partial x}{\partial v} = u, \quad \frac{\partial y}{\partial u} = 1 - v, \quad \frac{\partial y}{\partial v} = -u, \] \[ J = \begin{vmatrix} v & u \\ 1-v & -u \end{vmatrix} = v(-u) - u(1-v) = -uv - u + uv = -u, \qquad |J| = u, \]

using \(u > 0\) in the last step.

Step 3 — transform the support. \(x > 0\) and \(y > 0\) require \(uv > 0\) and \(u(1-v) > 0\). Since \(u = x + y > 0\), both reduce to \(0 < v < 1\). So the new region is \(u > 0\), \(0 < v < 1\) — a rectangle, where the old region was a quadrant.

Step 4 — substitute. Note \(x + y = uv + u(1-v) = u\), so the exponent simplifies at once:

\[ f_{U,V}(u,v) = e^{-u} \cdot u = u e^{-u}, \qquad u > 0,\; 0 < v < 1. \]

Step 5 — read off the marginals. The density factors as

\[ f_{U,V}(u,v) = \underbrace{u e^{-u}}_{\text{depends only on } u} \times \underbrace{1}_{\text{depends only on } v}, \]

and the support is a rectangle, so \(U\) and \(V\) are independent, with

\[ f_U(u) = u e^{-u} \ (u > 0) \quad\text{i.e. } U \sim \text{Gamma}(2, 1), \qquad f_V(v) = 1 \ (0 < v < 1) \quad\text{i.e. } V \sim U(0,1). \]

Step 6 — check each integrates to 1. \(\int_0^\infty u e^{-u} du = \Gamma(2) = 1! = 1\), and \(\int_0^1 1 \, dv = 1\). \(\checkmark\)

Interpretation. The total \(X+Y\) and the share \(X/(X+Y)\) carry completely separate information: knowing the total tells you nothing about how it split, and the split is uniform. The rectangular support was essential — had the region been triangular, the density could factor and the variables still not be independent.

2. Truncated Distributions

DEFINITION

If \(X\) has density (or pmf) \(f\) and we observe \(X\) only when it falls in a set \(A\) with \(P(X \in A) > 0\), the truncated distribution is the conditional one:

\[ f_A(x) = \frac{f(x)}{P(X \in A)}, \qquad x \in A, \]

and zero outside \(A\). Dividing by \(P(X \in A)\) is what restores the total to 1.

Truncation is not a modelling choice made for convenience; it is forced by how the data arose. Counts of household size from a survey of households cannot contain a zero. Claims below a deductible are never reported. Fitting an untruncated model to such data biases every estimate.

THE THREE CASES NAMED IN THE SYLLABUS \[ \begin{aligned} \textbf{Truncated binomial}\ (x \ge 1):\quad & P(X = x) = \frac{\binom{n}{x} p^{x} q^{n-x}}{1 - q^{n}} \\ \textbf{Truncated Poisson}\ (x \ge 1):\quad & P(X = x) = \frac{e^{-\lambda}\lambda^{x}/x!}{1 - e^{-\lambda}} \\ \textbf{Truncated normal}\ (a < x < b):\quad & f(x) = \frac{\frac{1}{\sigma}\varphi\!\left(\frac{x-\mu}{\sigma}\right)} {\Phi\!\left(\frac{b-\mu}{\sigma}\right) - \Phi\!\left(\frac{a-\mu}{\sigma}\right)} \end{aligned} \]

where \(\varphi\) and \(\Phi\) are the standard normal density and distribution function. In each case the numerator is the original law and the denominator is the probability of the retained region.

EXAMPLE 2.2 — THE ZERO-TRUNCATED POISSON

Given. \(X \sim\) Poisson\((\lambda = 2)\), observed only when \(X \ge 1\).

Asked. Write the truncated pmf, find its mean and variance, and compare with the untruncated \(\lambda = 2\).

Step 1 — the truncation constant.

\[ P(X = 0) = e^{-2} = 0.1353353, \qquad P(X \ge 1) = 1 - e^{-2} = 0.8646647. \]

Step 2 — the truncated pmf.

\[ P^{*}(X = x) = \frac{e^{-2} 2^{x}/x!}{0.8646647}, \qquad x = 1, 2, 3, \ldots \]

Evaluating the first three:

\[ P^{*}(1) = \frac{0.2706706}{0.8646647} = 0.313035, \quad P^{*}(2) = \frac{0.2706706}{0.8646647} = 0.313035, \quad P^{*}(3) = \frac{0.1804470}{0.8646647} = 0.208690. \]

(\(P^{*}(1) = P^{*}(2)\) is not an error: for Poisson\((\lambda)\) the probabilities at \(\lambda - 1\) and \(\lambda\) are equal when \(\lambda\) is an integer, and truncation divides both by the same constant.)

Step 3 — the mean. The truncated mean is the original mean divided by the retained probability, because the term removed at \(x = 0\) contributes \(0 \times P(0) = 0\) to the sum:

\[ E^{*}(X) = \frac{\sum_{x \ge 1} x\, e^{-2} 2^{x}/x!}{1 - e^{-2}} = \frac{E(X)}{1 - e^{-2}} = \frac{2}{0.8646647} = 2.3130353. \]

Step 4 — the second moment. The same argument applies, since \(0^{2} \times P(0) = 0\):

\[ E^{*}(X^{2}) = \frac{E(X^{2})}{1 - e^{-2}} = \frac{\lambda + \lambda^{2}}{0.8646647} = \frac{2 + 4}{0.8646647} = \frac{6}{0.8646647} = 6.9391059. \]

Step 5 — the variance.

\[ \operatorname{Var}^{*}(X) = 6.9391059 - (2.3130353)^{2} = 6.9391059 - 5.3501322 = 1.5889736. \]

Step 6 — compare.

untruncated Poisson(2)zero-truncated
mean2.00000002.3130353
variance2.00000001.5889736
variance / mean1.00000000.6870

Interpretation. Removing the zeros raises the mean by \(15.7\%\) and lowers the variance by \(20.6\%\). The equality of mean and variance — the single property by which a Poisson is usually recognised — is destroyed: the truncated law is under-dispersed. So a dataset that cannot contain zeros will fail a Poisson dispersion test even when the underlying process is exactly Poisson, and the remedy is to fit the truncated model, not to abandon the Poisson.

3. Mixture Distributions

DEFINITION

A finite mixture of \(k\) component densities \(f_1, \ldots, f_k\) with weights \(w_i > 0\), \(\sum_i w_i = 1\), is

\[ f(x) = \sum_{i=1}^{k} w_i f_i(x). \]

The mechanism is two-stage: pick component \(i\) with probability \(w_i\), then draw from \(f_i\). Introducing the label \(I\) with \(P(I = i) = w_i\), the two identities of Probability Theory, Unit 2 give the moments immediately:

\[ E(X) = \sum_i w_i \mu_i, \qquad \operatorname{Var}(X) = \underbrace{\sum_i w_i \sigma_i^{2}}_{E[\operatorname{Var}(X \mid I)]} + \underbrace{\sum_i w_i (\mu_i - \bar\mu)^{2}}_{\operatorname{Var}[E(X \mid I)]}, \]

with \(\bar\mu = \sum_i w_i \mu_i\). The second term is why a mixture is always more variable than the average of its components.

A mixture is not a sum. Mixing \(N(0,1)\) and \(N(3,1)\) selects one of them; adding them would give \(N(3,2)\), a single normal. The mixture is not normal at all.

EXAMPLE 2.3 — A TWO-COMPONENT NORMAL MIXTURE

Given. \(f(x) = 0.6\,\varphi(x; 0, 1) + 0.4\,\varphi(x; 3, 1)\), where \(\varphi(x; \mu, \sigma^{2})\) is the \(N(\mu, \sigma^{2})\) density.

Asked. Find the mean and variance, and say whether the mixture is normal.

Step 1 — the mean.

\[ E(X) = 0.6 \times 0 + 0.4 \times 3 = 1.2. \]

Step 2 — the second moment. For each component, \(E(X^{2} \mid I = i) = \sigma_i^{2} + \mu_i^{2}\), so

\[ E(X^{2}) = 0.6\,(1 + 0^{2}) + 0.4\,(1 + 3^{2}) = 0.6(1) + 0.4(10) = 0.6 + 4.0 = 4.6. \]

Step 3 — the variance.

\[ \operatorname{Var}(X) = 4.6 - (1.2)^{2} = 4.6 - 1.44 = 3.16. \]

Step 4 — check it against the decomposition. Within-component variance is \(0.6(1) + 0.4(1) = 1\); between-component variance is

\[ 0.6(0 - 1.2)^{2} + 0.4(3 - 1.2)^{2} = 0.6(1.44) + 0.4(3.24) = 0.864 + 1.296 = 2.16, \] \[ 1 + 2.16 = 3.16. \checkmark \]

There is also a shortcut for two components: \(\operatorname{Var} = \sigma^{2} + w_1 w_2 (\mu_2 - \mu_1)^{2} = 1 + (0.6)(0.4)(3)^{2} = 1 + 2.16 = 3.16\).

Step 5 — is it normal? No. A normal density has exactly one mode; this one has two. Evaluating the density confirms it:

\[ f(0) = 0.241138, \qquad f(1.5) = 0.129518, \qquad f(3) = 0.162236. \]

The value at \(1.5\) is below the values on either side, so there is a genuine trough between two peaks — impossible for any normal density.

Interpretation. The variance \(3.16\) is more than three times either component's variance of \(1\), and all of that excess is the separation between the means. Reporting mean \(1.2\) and s.d. \(1.78\) would describe a population in which almost nobody sits near \(1.2\) — the summary is arithmetically correct and substantively misleading, which is the standard hazard of averaging over unrecognised groups.

The mixture (solid) is not normal: two modes, and a trough between them mean 1.2 sits in the trough 0.6 × N(0,1) 0.4 × N(3,1) -3 0 3 6 0.10 0.20 f(0) = 0.2411, f(1.5) = 0.1295, f(3) = 0.1622 — the middle value is the lowest variance 3.16 against 1 for either component, the excess being the gap between means
Fig 2.1 — The mixture of Example 2.3 with its two weighted components. All three curves are computed from the densities.

4. The Exponential Family

DEFINITION

A family \(\{f(x;\theta)\}\) belongs to the one-parameter exponential family if its density or pmf can be written

\[ f(x; \theta) = h(x)\, \exp\big\{ \eta(\theta)\, T(x) - A(\theta) \big\}, \]

with \(h\) and \(T\) free of \(\theta\), and the support free of \(\theta\). Here \(\eta\) is the natural parameter, \(T\) the natural sufficient statistic and \(A\) the log-normaliser. Writing \(\eta\) itself as the parameter gives the canonical form \(f(x;\eta) = h(x) e^{\eta T(x) - A(\eta)}\).

The support condition is not a technicality. \(U(0,\theta)\) cannot be written in this form, because its support \((0,\theta)\) moves with \(\theta\). That single fact is why the uniform is the standard counterexample to results proved for exponential families.

MOMENTS STRAIGHT FROM THE LOG-NORMALISER \[ E\big(T(X)\big) = A'(\eta), \qquad \operatorname{Var}\big(T(X)\big) = A''(\eta). \]

Two derivatives of one function replace two integrations. This is the reason the exponential family is the organising idea of Estimation Theory (STS-201): \(T(X)\) is sufficient, and under mild conditions complete, which is what makes the Lehmann–Scheffé theorem usable.

EXAMPLE 2.4 — THE BINOMIAL AS AN EXPONENTIAL FAMILY

Given. \(X \sim \text{Bin}(n, p)\) with \(n\) known.

Asked. Put it in exponential-family form, identify \(\eta\), \(T\) and \(A\), and recover the mean and variance by differentiating \(A\). Check with \(n = 10\), \(p = 0.3\).

Step 1 — start from the pmf and exponentiate the logarithm.

\[ P(X = x) = \binom{n}{x} p^{x}(1-p)^{n-x} = \binom{n}{x} \exp\big\{ x \ln p + (n - x)\ln(1-p) \big\}. \]

Step 2 — collect the terms in \(x\).

\[ x \ln p + (n-x)\ln(1-p) = x\big[\ln p - \ln(1-p)\big] + n \ln(1-p) = x \ln\!\left(\frac{p}{1-p}\right) + n\ln(1-p). \]

Step 3 — read off the pieces. Comparing with \(h(x)\exp\{\eta T(x) - A(\eta)\}\),

\[ h(x) = \binom{n}{x}, \qquad T(x) = x, \qquad \eta = \ln\!\left(\frac{p}{1-p}\right), \qquad A = -n\ln(1-p). \]

The natural parameter is the log-odds, which is exactly the link function of logistic regression — not a coincidence but this calculation.

Step 4 — write \(A\) in terms of \(\eta\). Inverting \(\eta = \ln(p/(1-p))\) gives \(p = e^{\eta}/(1 + e^{\eta})\), hence \(1 - p = 1/(1 + e^{\eta})\) and

\[ A(\eta) = -n \ln\!\left(\frac{1}{1 + e^{\eta}}\right) = n \ln\left(1 + e^{\eta}\right). \]

Step 5 — differentiate once.

\[ A'(\eta) = n \cdot \frac{e^{\eta}}{1 + e^{\eta}} = n p = E(X). \checkmark \]

Step 6 — differentiate twice. By the quotient rule,

\[ A''(\eta) = n \cdot \frac{e^{\eta}(1 + e^{\eta}) - e^{\eta} \cdot e^{\eta}}{(1+e^{\eta})^{2}} = n \cdot \frac{e^{\eta}}{(1+e^{\eta})^{2}} = n p (1-p) = npq = \operatorname{Var}(X). \checkmark \]

Step 7 — numerical check at \(n = 10\), \(p = 0.3\).

\[ \eta = \ln\!\left(\frac{0.3}{0.7}\right) = -0.8472979, \qquad A(\eta) = 10 \ln\left(1 + e^{-0.8472979}\right) = 3.5667494, \] \[ A'(\eta) = 3.0000000 = np = 10(0.3), \qquad A''(\eta) = 2.1000000 = npq = 10(0.3)(0.7). \checkmark \]

5. The Power Series Family

DEFINITION AND MOMENTS

A discrete distribution belongs to the power series family if

\[ P(X = x) = \frac{a_x \theta^{x}}{g(\theta)}, \qquad x = 0, 1, 2, \ldots, \quad \theta > 0, \]

where \(a_x \ge 0\) does not involve \(\theta\) and \(g(\theta) = \sum_{x} a_x \theta^{x}\) is the series that makes the probabilities sum to 1.

Mean and variance. Differentiating \(g(\theta) = \sum_x a_x \theta^{x}\) term by term,

\[ E(X) = \frac{\theta g'(\theta)}{g(\theta)}, \qquad \operatorname{Var}(X) = \theta \frac{d}{d\theta} E(X) = \theta \mu'(\theta). \]

Both follow from one differentiation of one series, with no summation over \(x\) at all.

Distribution\(a_x\)\(\theta\)\(g(\theta)\)MeanVariance
Binomial \((n,p)\)\(\binom{n}{x}\)\(p/q\)\((1+\theta)^{n}\)\(np\)\(npq\)
Poisson \((\lambda)\)\(1/x!\)\(\lambda\)\(e^{\theta}\)\(\lambda\)\(\lambda\)
Geometric \((p)\)\(1\)\(q\)\(1/(1-\theta)\)\(q/p\)\(q/p^{2}\)

Worked check, Poisson. With \(g(\theta) = e^{\theta}\) we get \(g'(\theta) = e^{\theta}\), so

\[ E(X) = \frac{\theta e^{\theta}}{e^{\theta}} = \theta = \lambda, \qquad \operatorname{Var}(X) = \theta \frac{d\theta}{d\theta} = \theta = \lambda, \]

recovering the equality of mean and variance in two lines rather than by summing \(\sum x^{2} e^{-\lambda}\lambda^{x}/x!\).

Worked check, geometric. \(g(\theta) = (1-\theta)^{-1}\) gives \(g'(\theta) = (1-\theta)^{-2}\), so

\[ E(X) = \frac{\theta (1-\theta)^{-2}}{(1-\theta)^{-1}} = \frac{\theta}{1-\theta} = \frac{q}{p}, \]

since \(\theta = q\) and \(1 - \theta = p\).

6. Compound Distributions

WHAT COMPOUNDING IS

A compound (or mixed) distribution arises when a parameter of one distribution is itself random:

\[ X \mid \theta \sim f(\cdot \mid \theta), \qquad \theta \sim \pi(\cdot). \]

The moments follow from the two identities of Probability Theory, Unit 2:

\[ E(X) = E\big[E(X \mid \theta)\big], \qquad \operatorname{Var}(X) = E\big[\operatorname{Var}(X \mid \theta)\big] + \operatorname{Var}\big[E(X \mid \theta)\big]. \]

The second term is always non-negative, so compounding can only increase variance relative to the conditional model. That is the mechanism of over-dispersion.

EXAMPLE 2.5 — THE BINOMIAL–POISSON COMPOUND

Given. \(N \sim\) Poisson\((\lambda = 10)\) customers arrive, and each independently buys with probability \(p = 0.3\), so \(X \mid N \sim \text{Bin}(N, 0.3)\).

Asked. Find the distribution, mean and variance of the number of buyers \(X\).

Step 1 — the conditional moments. \(E(X \mid N) = Np\) and \(\operatorname{Var}(X \mid N) = Npq\).

Step 2 — the mean.

\[ E(X) = E(Np) = p\,E(N) = 0.3 \times 10 = 3. \]

Step 3 — the variance, term by term.

\[ E\big[\operatorname{Var}(X \mid N)\big] = E(Npq) = pq\,E(N) = 0.3 \times 0.7 \times 10 = 2.1, \] \[ \operatorname{Var}\big[E(X \mid N)\big] = \operatorname{Var}(Np) = p^{2}\operatorname{Var}(N) = 0.09 \times 10 = 0.9, \] \[ \operatorname{Var}(X) = 2.1 + 0.9 = 3.0. \]

Step 4 — identify the distribution. Mean and variance are both 3, which points at a Poisson, and indeed the compound is exactly Poisson\((\lambda p)\) = Poisson\((3)\). Thinning a Poisson stream by independent selection leaves a Poisson stream.

Interpretation. Here compounding did not over-disperse, because the two components happen to add back to \(\lambda p\). That is special to the Poisson; the next example shows the general behaviour.

EXAMPLE 2.6 — THE POISSON–GAMMA COMPOUND IS A NEGATIVE BINOMIAL

Given. \(N \mid \lambda \sim\) Poisson\((\lambda)\), with \(\lambda \sim\) Gamma(shape \(a = 4\), rate \(b = 2\)).

Asked. Find \(E(N)\) and \(\operatorname{Var}(N)\), and identify the marginal distribution.

Step 1 — the moments of the mixing distribution. For Gamma(shape \(a\), rate \(b\)),

\[ E(\lambda) = \frac{a}{b} = \frac{4}{2} = 2, \qquad \operatorname{Var}(\lambda) = \frac{a}{b^{2}} = \frac{4}{4} = 1. \]

Step 2 — the mean. Since \(E(N \mid \lambda) = \lambda\),

\[ E(N) = E(\lambda) = 2. \]

Step 3 — the variance. Since \(\operatorname{Var}(N \mid \lambda) = \lambda\) as well,

\[ \operatorname{Var}(N) = E(\lambda) + \operatorname{Var}(\lambda) = 2 + 1 = 3. \]

Step 4 — identify the marginal. Integrating the Poisson pmf against the gamma density gives a negative binomial with

\[ r = a = 4, \qquad p = \frac{b}{1 + b} = \frac{2}{3} = 0.666667. \]

Check against the negative binomial formulae \(E = r(1-p)/p\) and \(\operatorname{Var} = r(1-p)/p^{2}\):

\[ E = \frac{4 \times \tfrac13}{\tfrac23} = \frac{4/3}{2/3} = 2, \qquad \operatorname{Var} = \frac{4 \times \tfrac13}{(2/3)^{2}} = \frac{4/3}{4/9} = 3. \checkmark \]

Step 5 — the index of dispersion.

\[ \frac{\operatorname{Var}(N)}{E(N)} = \frac{3}{2} = 1.5 > 1. \]

Interpretation. The data are over-dispersed relative to a Poisson by exactly the variance of the mixing distribution. This is the standard explanation for count data whose variance exceeds its mean: the rate is not constant across units. The negative binomial is therefore not an arbitrary alternative to the Poisson but the exact consequence of a gamma-distributed rate — which is also why it is the default model for over-dispersed counts.

Key Take-aways