Let \((X, Y)\) have joint density \(f_{X,Y}\) and let \(u = u(x,y)\), \(v = v(x,y)\) be a one-to-one transformation with inverse \(x = x(u,v)\), \(y = y(u,v)\). Then
\[ f_{U,V}(u,v) = f_{X,Y}\big(x(u,v),\, y(u,v)\big) \, |J|, \qquad J = \begin{vmatrix} \dfrac{\partial x}{\partial u} & \dfrac{\partial x}{\partial v} \\ \dfrac{\partial y}{\partial u} & \dfrac{\partial y}{\partial v} \end{vmatrix}. \]Three things decide whether an answer is right: the inverse must be solved correctly, the Jacobian must be the determinant of the derivatives of the old variables with respect to the new ones, and the support must be transformed too. The third is the one most often dropped, and it is usually where the marks go.
Given. \(X\) and \(Y\) independent, each Exponential\((1)\), so
\[ f_{X,Y}(x,y) = e^{-x} e^{-y} = e^{-(x+y)}, \qquad x > 0,\; y > 0. \]Asked. Find the joint density of \(U = X + Y\) and \(V = \dfrac{X}{X+Y}\), and identify the marginals.
Step 1 — invert the transformation. From \(V = X/U\) we get \(X = UV\), and then \(Y = U - X = U(1 - V)\). So
\[ x = uv, \qquad y = u(1-v). \]Step 2 — the Jacobian. Differentiate the inverse:
\[ \frac{\partial x}{\partial u} = v, \quad \frac{\partial x}{\partial v} = u, \quad \frac{\partial y}{\partial u} = 1 - v, \quad \frac{\partial y}{\partial v} = -u, \] \[ J = \begin{vmatrix} v & u \\ 1-v & -u \end{vmatrix} = v(-u) - u(1-v) = -uv - u + uv = -u, \qquad |J| = u, \]using \(u > 0\) in the last step.
Step 3 — transform the support. \(x > 0\) and \(y > 0\) require \(uv > 0\) and \(u(1-v) > 0\). Since \(u = x + y > 0\), both reduce to \(0 < v < 1\). So the new region is \(u > 0\), \(0 < v < 1\) — a rectangle, where the old region was a quadrant.
Step 4 — substitute. Note \(x + y = uv + u(1-v) = u\), so the exponent simplifies at once:
\[ f_{U,V}(u,v) = e^{-u} \cdot u = u e^{-u}, \qquad u > 0,\; 0 < v < 1. \]Step 5 — read off the marginals. The density factors as
\[ f_{U,V}(u,v) = \underbrace{u e^{-u}}_{\text{depends only on } u} \times \underbrace{1}_{\text{depends only on } v}, \]and the support is a rectangle, so \(U\) and \(V\) are independent, with
\[ f_U(u) = u e^{-u} \ (u > 0) \quad\text{i.e. } U \sim \text{Gamma}(2, 1), \qquad f_V(v) = 1 \ (0 < v < 1) \quad\text{i.e. } V \sim U(0,1). \]Step 6 — check each integrates to 1. \(\int_0^\infty u e^{-u} du = \Gamma(2) = 1! = 1\), and \(\int_0^1 1 \, dv = 1\). \(\checkmark\)
Interpretation. The total \(X+Y\) and the share \(X/(X+Y)\) carry completely separate information: knowing the total tells you nothing about how it split, and the split is uniform. The rectangular support was essential — had the region been triangular, the density could factor and the variables still not be independent.
If \(X\) has density (or pmf) \(f\) and we observe \(X\) only when it falls in a set \(A\) with \(P(X \in A) > 0\), the truncated distribution is the conditional one:
\[ f_A(x) = \frac{f(x)}{P(X \in A)}, \qquad x \in A, \]and zero outside \(A\). Dividing by \(P(X \in A)\) is what restores the total to 1.
Truncation is not a modelling choice made for convenience; it is forced by how the data arose. Counts of household size from a survey of households cannot contain a zero. Claims below a deductible are never reported. Fitting an untruncated model to such data biases every estimate.
where \(\varphi\) and \(\Phi\) are the standard normal density and distribution function. In each case the numerator is the original law and the denominator is the probability of the retained region.
Given. \(X \sim\) Poisson\((\lambda = 2)\), observed only when \(X \ge 1\).
Asked. Write the truncated pmf, find its mean and variance, and compare with the untruncated \(\lambda = 2\).
Step 1 — the truncation constant.
\[ P(X = 0) = e^{-2} = 0.1353353, \qquad P(X \ge 1) = 1 - e^{-2} = 0.8646647. \]Step 2 — the truncated pmf.
\[ P^{*}(X = x) = \frac{e^{-2} 2^{x}/x!}{0.8646647}, \qquad x = 1, 2, 3, \ldots \]Evaluating the first three:
\[ P^{*}(1) = \frac{0.2706706}{0.8646647} = 0.313035, \quad P^{*}(2) = \frac{0.2706706}{0.8646647} = 0.313035, \quad P^{*}(3) = \frac{0.1804470}{0.8646647} = 0.208690. \](\(P^{*}(1) = P^{*}(2)\) is not an error: for Poisson\((\lambda)\) the probabilities at \(\lambda - 1\) and \(\lambda\) are equal when \(\lambda\) is an integer, and truncation divides both by the same constant.)
Step 3 — the mean. The truncated mean is the original mean divided by the retained probability, because the term removed at \(x = 0\) contributes \(0 \times P(0) = 0\) to the sum:
\[ E^{*}(X) = \frac{\sum_{x \ge 1} x\, e^{-2} 2^{x}/x!}{1 - e^{-2}} = \frac{E(X)}{1 - e^{-2}} = \frac{2}{0.8646647} = 2.3130353. \]Step 4 — the second moment. The same argument applies, since \(0^{2} \times P(0) = 0\):
\[ E^{*}(X^{2}) = \frac{E(X^{2})}{1 - e^{-2}} = \frac{\lambda + \lambda^{2}}{0.8646647} = \frac{2 + 4}{0.8646647} = \frac{6}{0.8646647} = 6.9391059. \]Step 5 — the variance.
\[ \operatorname{Var}^{*}(X) = 6.9391059 - (2.3130353)^{2} = 6.9391059 - 5.3501322 = 1.5889736. \]Step 6 — compare.
| untruncated Poisson(2) | zero-truncated | |
|---|---|---|
| mean | 2.0000000 | 2.3130353 |
| variance | 2.0000000 | 1.5889736 |
| variance / mean | 1.0000000 | 0.6870 |
Interpretation. Removing the zeros raises the mean by \(15.7\%\) and lowers the variance by \(20.6\%\). The equality of mean and variance — the single property by which a Poisson is usually recognised — is destroyed: the truncated law is under-dispersed. So a dataset that cannot contain zeros will fail a Poisson dispersion test even when the underlying process is exactly Poisson, and the remedy is to fit the truncated model, not to abandon the Poisson.
A finite mixture of \(k\) component densities \(f_1, \ldots, f_k\) with weights \(w_i > 0\), \(\sum_i w_i = 1\), is
\[ f(x) = \sum_{i=1}^{k} w_i f_i(x). \]The mechanism is two-stage: pick component \(i\) with probability \(w_i\), then draw from \(f_i\). Introducing the label \(I\) with \(P(I = i) = w_i\), the two identities of Probability Theory, Unit 2 give the moments immediately:
\[ E(X) = \sum_i w_i \mu_i, \qquad \operatorname{Var}(X) = \underbrace{\sum_i w_i \sigma_i^{2}}_{E[\operatorname{Var}(X \mid I)]} + \underbrace{\sum_i w_i (\mu_i - \bar\mu)^{2}}_{\operatorname{Var}[E(X \mid I)]}, \]with \(\bar\mu = \sum_i w_i \mu_i\). The second term is why a mixture is always more variable than the average of its components.
A mixture is not a sum. Mixing \(N(0,1)\) and \(N(3,1)\) selects one of them; adding them would give \(N(3,2)\), a single normal. The mixture is not normal at all.
Given. \(f(x) = 0.6\,\varphi(x; 0, 1) + 0.4\,\varphi(x; 3, 1)\), where \(\varphi(x; \mu, \sigma^{2})\) is the \(N(\mu, \sigma^{2})\) density.
Asked. Find the mean and variance, and say whether the mixture is normal.
Step 1 — the mean.
\[ E(X) = 0.6 \times 0 + 0.4 \times 3 = 1.2. \]Step 2 — the second moment. For each component, \(E(X^{2} \mid I = i) = \sigma_i^{2} + \mu_i^{2}\), so
\[ E(X^{2}) = 0.6\,(1 + 0^{2}) + 0.4\,(1 + 3^{2}) = 0.6(1) + 0.4(10) = 0.6 + 4.0 = 4.6. \]Step 3 — the variance.
\[ \operatorname{Var}(X) = 4.6 - (1.2)^{2} = 4.6 - 1.44 = 3.16. \]Step 4 — check it against the decomposition. Within-component variance is \(0.6(1) + 0.4(1) = 1\); between-component variance is
\[ 0.6(0 - 1.2)^{2} + 0.4(3 - 1.2)^{2} = 0.6(1.44) + 0.4(3.24) = 0.864 + 1.296 = 2.16, \] \[ 1 + 2.16 = 3.16. \checkmark \]There is also a shortcut for two components: \(\operatorname{Var} = \sigma^{2} + w_1 w_2 (\mu_2 - \mu_1)^{2} = 1 + (0.6)(0.4)(3)^{2} = 1 + 2.16 = 3.16\).
Step 5 — is it normal? No. A normal density has exactly one mode; this one has two. Evaluating the density confirms it:
\[ f(0) = 0.241138, \qquad f(1.5) = 0.129518, \qquad f(3) = 0.162236. \]The value at \(1.5\) is below the values on either side, so there is a genuine trough between two peaks — impossible for any normal density.
Interpretation. The variance \(3.16\) is more than three times either component's variance of \(1\), and all of that excess is the separation between the means. Reporting mean \(1.2\) and s.d. \(1.78\) would describe a population in which almost nobody sits near \(1.2\) — the summary is arithmetically correct and substantively misleading, which is the standard hazard of averaging over unrecognised groups.
A family \(\{f(x;\theta)\}\) belongs to the one-parameter exponential family if its density or pmf can be written
\[ f(x; \theta) = h(x)\, \exp\big\{ \eta(\theta)\, T(x) - A(\theta) \big\}, \]with \(h\) and \(T\) free of \(\theta\), and the support free of \(\theta\). Here \(\eta\) is the natural parameter, \(T\) the natural sufficient statistic and \(A\) the log-normaliser. Writing \(\eta\) itself as the parameter gives the canonical form \(f(x;\eta) = h(x) e^{\eta T(x) - A(\eta)}\).
The support condition is not a technicality. \(U(0,\theta)\) cannot be written in this form, because its support \((0,\theta)\) moves with \(\theta\). That single fact is why the uniform is the standard counterexample to results proved for exponential families.
Two derivatives of one function replace two integrations. This is the reason the exponential family is the organising idea of Estimation Theory (STS-201): \(T(X)\) is sufficient, and under mild conditions complete, which is what makes the Lehmann–Scheffé theorem usable.
Given. \(X \sim \text{Bin}(n, p)\) with \(n\) known.
Asked. Put it in exponential-family form, identify \(\eta\), \(T\) and \(A\), and recover the mean and variance by differentiating \(A\). Check with \(n = 10\), \(p = 0.3\).
Step 1 — start from the pmf and exponentiate the logarithm.
\[ P(X = x) = \binom{n}{x} p^{x}(1-p)^{n-x} = \binom{n}{x} \exp\big\{ x \ln p + (n - x)\ln(1-p) \big\}. \]Step 2 — collect the terms in \(x\).
\[ x \ln p + (n-x)\ln(1-p) = x\big[\ln p - \ln(1-p)\big] + n \ln(1-p) = x \ln\!\left(\frac{p}{1-p}\right) + n\ln(1-p). \]Step 3 — read off the pieces. Comparing with \(h(x)\exp\{\eta T(x) - A(\eta)\}\),
\[ h(x) = \binom{n}{x}, \qquad T(x) = x, \qquad \eta = \ln\!\left(\frac{p}{1-p}\right), \qquad A = -n\ln(1-p). \]The natural parameter is the log-odds, which is exactly the link function of logistic regression — not a coincidence but this calculation.
Step 4 — write \(A\) in terms of \(\eta\). Inverting \(\eta = \ln(p/(1-p))\) gives \(p = e^{\eta}/(1 + e^{\eta})\), hence \(1 - p = 1/(1 + e^{\eta})\) and
\[ A(\eta) = -n \ln\!\left(\frac{1}{1 + e^{\eta}}\right) = n \ln\left(1 + e^{\eta}\right). \]Step 5 — differentiate once.
\[ A'(\eta) = n \cdot \frac{e^{\eta}}{1 + e^{\eta}} = n p = E(X). \checkmark \]Step 6 — differentiate twice. By the quotient rule,
\[ A''(\eta) = n \cdot \frac{e^{\eta}(1 + e^{\eta}) - e^{\eta} \cdot e^{\eta}}{(1+e^{\eta})^{2}} = n \cdot \frac{e^{\eta}}{(1+e^{\eta})^{2}} = n p (1-p) = npq = \operatorname{Var}(X). \checkmark \]Step 7 — numerical check at \(n = 10\), \(p = 0.3\).
\[ \eta = \ln\!\left(\frac{0.3}{0.7}\right) = -0.8472979, \qquad A(\eta) = 10 \ln\left(1 + e^{-0.8472979}\right) = 3.5667494, \] \[ A'(\eta) = 3.0000000 = np = 10(0.3), \qquad A''(\eta) = 2.1000000 = npq = 10(0.3)(0.7). \checkmark \]A discrete distribution belongs to the power series family if
\[ P(X = x) = \frac{a_x \theta^{x}}{g(\theta)}, \qquad x = 0, 1, 2, \ldots, \quad \theta > 0, \]where \(a_x \ge 0\) does not involve \(\theta\) and \(g(\theta) = \sum_{x} a_x \theta^{x}\) is the series that makes the probabilities sum to 1.
Mean and variance. Differentiating \(g(\theta) = \sum_x a_x \theta^{x}\) term by term,
\[ E(X) = \frac{\theta g'(\theta)}{g(\theta)}, \qquad \operatorname{Var}(X) = \theta \frac{d}{d\theta} E(X) = \theta \mu'(\theta). \]Both follow from one differentiation of one series, with no summation over \(x\) at all.
| Distribution | \(a_x\) | \(\theta\) | \(g(\theta)\) | Mean | Variance |
|---|---|---|---|---|---|
| Binomial \((n,p)\) | \(\binom{n}{x}\) | \(p/q\) | \((1+\theta)^{n}\) | \(np\) | \(npq\) |
| Poisson \((\lambda)\) | \(1/x!\) | \(\lambda\) | \(e^{\theta}\) | \(\lambda\) | \(\lambda\) |
| Geometric \((p)\) | \(1\) | \(q\) | \(1/(1-\theta)\) | \(q/p\) | \(q/p^{2}\) |
Worked check, Poisson. With \(g(\theta) = e^{\theta}\) we get \(g'(\theta) = e^{\theta}\), so
\[ E(X) = \frac{\theta e^{\theta}}{e^{\theta}} = \theta = \lambda, \qquad \operatorname{Var}(X) = \theta \frac{d\theta}{d\theta} = \theta = \lambda, \]recovering the equality of mean and variance in two lines rather than by summing \(\sum x^{2} e^{-\lambda}\lambda^{x}/x!\).
Worked check, geometric. \(g(\theta) = (1-\theta)^{-1}\) gives \(g'(\theta) = (1-\theta)^{-2}\), so
\[ E(X) = \frac{\theta (1-\theta)^{-2}}{(1-\theta)^{-1}} = \frac{\theta}{1-\theta} = \frac{q}{p}, \]since \(\theta = q\) and \(1 - \theta = p\).
A compound (or mixed) distribution arises when a parameter of one distribution is itself random:
\[ X \mid \theta \sim f(\cdot \mid \theta), \qquad \theta \sim \pi(\cdot). \]The moments follow from the two identities of Probability Theory, Unit 2:
\[ E(X) = E\big[E(X \mid \theta)\big], \qquad \operatorname{Var}(X) = E\big[\operatorname{Var}(X \mid \theta)\big] + \operatorname{Var}\big[E(X \mid \theta)\big]. \]The second term is always non-negative, so compounding can only increase variance relative to the conditional model. That is the mechanism of over-dispersion.
Given. \(N \sim\) Poisson\((\lambda = 10)\) customers arrive, and each independently buys with probability \(p = 0.3\), so \(X \mid N \sim \text{Bin}(N, 0.3)\).
Asked. Find the distribution, mean and variance of the number of buyers \(X\).
Step 1 — the conditional moments. \(E(X \mid N) = Np\) and \(\operatorname{Var}(X \mid N) = Npq\).
Step 2 — the mean.
\[ E(X) = E(Np) = p\,E(N) = 0.3 \times 10 = 3. \]Step 3 — the variance, term by term.
\[ E\big[\operatorname{Var}(X \mid N)\big] = E(Npq) = pq\,E(N) = 0.3 \times 0.7 \times 10 = 2.1, \] \[ \operatorname{Var}\big[E(X \mid N)\big] = \operatorname{Var}(Np) = p^{2}\operatorname{Var}(N) = 0.09 \times 10 = 0.9, \] \[ \operatorname{Var}(X) = 2.1 + 0.9 = 3.0. \]Step 4 — identify the distribution. Mean and variance are both 3, which points at a Poisson, and indeed the compound is exactly Poisson\((\lambda p)\) = Poisson\((3)\). Thinning a Poisson stream by independent selection leaves a Poisson stream.
Interpretation. Here compounding did not over-disperse, because the two components happen to add back to \(\lambda p\). That is special to the Poisson; the next example shows the general behaviour.
Given. \(N \mid \lambda \sim\) Poisson\((\lambda)\), with \(\lambda \sim\) Gamma(shape \(a = 4\), rate \(b = 2\)).
Asked. Find \(E(N)\) and \(\operatorname{Var}(N)\), and identify the marginal distribution.
Step 1 — the moments of the mixing distribution. For Gamma(shape \(a\), rate \(b\)),
\[ E(\lambda) = \frac{a}{b} = \frac{4}{2} = 2, \qquad \operatorname{Var}(\lambda) = \frac{a}{b^{2}} = \frac{4}{4} = 1. \]Step 2 — the mean. Since \(E(N \mid \lambda) = \lambda\),
\[ E(N) = E(\lambda) = 2. \]Step 3 — the variance. Since \(\operatorname{Var}(N \mid \lambda) = \lambda\) as well,
\[ \operatorname{Var}(N) = E(\lambda) + \operatorname{Var}(\lambda) = 2 + 1 = 3. \]Step 4 — identify the marginal. Integrating the Poisson pmf against the gamma density gives a negative binomial with
\[ r = a = 4, \qquad p = \frac{b}{1 + b} = \frac{2}{3} = 0.666667. \]Check against the negative binomial formulae \(E = r(1-p)/p\) and \(\operatorname{Var} = r(1-p)/p^{2}\):
\[ E = \frac{4 \times \tfrac13}{\tfrac23} = \frac{4/3}{2/3} = 2, \qquad \operatorname{Var} = \frac{4 \times \tfrac13}{(2/3)^{2}} = \frac{4/3}{4/9} = 3. \checkmark \]Step 5 — the index of dispersion.
\[ \frac{\operatorname{Var}(N)}{E(N)} = \frac{3}{2} = 1.5 > 1. \]Interpretation. The data are over-dispersed relative to a Poisson by exactly the variance of the mixing distribution. This is the standard explanation for count data whose variance exceeds its mean: the rate is not constant across units. The negative binomial is therefore not an arbitrary alternative to the Poisson but the exact consequence of a gamma-distributed rate — which is also why it is the default model for over-dispersed counts.