Eight coins are tossed and the number of heads \(X\) is recorded; the experiment is repeated 256 times:
| x | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|---|
| f | 2 | 6 | 30 | 52 | 67 | 56 | 32 | 10 | 1 |
Fit a Binomial distribution and obtain the expected frequencies.
To fit a Binomial distribution to observed data by the direct method and compare expected with observed frequencies.
Applying it:
Blank working table (compute p(x) and the expected frequencies):
| x | p(x) | Expected f = 256 p(x) | Observed f |
|---|---|---|---|
| 0 | 2 | ||
| 1 | 6 | ||
| 2 | 30 | ||
| 3 | 52 | ||
| 4 | 67 | ||
| 5 | 56 | ||
| 6 | 32 | ||
| 7 | 10 | ||
| 8 | 1 | ||
| Total | 1.000 | 256 | 256 |
\(N = 256\); \(\sum f x = 0+6+60+156+268+280+192+70+8 = 1040\); \(\bar x = 1040/256 = 4.0625\).
\(\hat p = 4.0625/8 = 0.5078\), \(\hat q = 0.4922\).
| x | p(x) | Expected f | Observed f |
|---|---|---|---|
| 0 | 0.0034 | 0.88 | 2 |
| 1 | 0.0284 | 7.28 | 6 |
| 2 | 0.1027 | 26.28 | 30 |
| 3 | 0.2118 | 54.23 | 52 |
| 4 | 0.2732 | 69.93 | 67 |
| 5 | 0.2255 | 57.72 | 56 |
| 6 | 0.1163 | 29.77 | 32 |
| 7 | 0.0343 | 8.78 | 10 |
| 8 | 0.0044 | 1.13 | 1 |
| Total | 1.000 | 256.0 | 256 |
The fitted distribution is \(B(8,\ 0.5078)\). The expected frequencies (0.9, 7.3, 26.3, 54.2, 69.9, 57.7, 29.8, 8.8, 1.1) are close to the observed values, so the Binomial model fits the data well.
For the same coin-tossing data as Experiment 1, obtain the Binomial probabilities using the recurrence relation instead of computing each \(\binom{n}{x}\) separately.
To fit a Binomial distribution using the recurrence relation, which avoids repeated factorial computation.
Applying it:
Blank working table (fill in each probability from the recurrence):
| x | Multiplier (n−x)/(x+1) · p/q | P(x) |
|---|---|---|
| 0 | — | |
| 1 | ||
| 2 | ||
| 3 | ||
| 4 |
\(P(0) = 0.4922^{8} = 0.0034\).
The recurrence reproduces exactly the probabilities of Experiment 1 (0.0034, 0.0284, 0.1027, 0.2118, 0.2732, …) with far less arithmetic, confirming the direct-method fit.
The number of accidents per day at a junction over 100 days:
| x | 0 | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|---|
| f | 46 | 38 | 12 | 3 | 1 | 0 |
Fit a Poisson distribution.
To fit a Poisson distribution to observed count data by the direct method.
Applying it:
Blank working table (compute p(x) and the expected frequencies):
| x | p(x) | Expected f = 100 p(x) | Observed f |
|---|---|---|---|
| 0 | 46 | ||
| 1 | 38 | ||
| 2 | 12 | ||
| 3 | 3 | ||
| 4 | 1 | ||
| 5+ | 0 | ||
| Total | 1.000 | 100 | 100 |
\(N = 100\); \(\sum f x = 0+38+24+9+4+0 = 75\); \(\hat\lambda = 0.75\); \(e^{-0.75} = 0.4724\).
| x | p(x) | Expected f | Observed f |
|---|---|---|---|
| 0 | 0.4724 | 47.24 | 46 |
| 1 | 0.3543 | 35.43 | 38 |
| 2 | 0.1329 | 13.29 | 12 |
| 3 | 0.0332 | 3.32 | 3 |
| 4 | 0.0062 | 0.62 | 1 |
| 5+ | 0.0010 | 0.10 | 0 |
| Total | 1.000 | 100.0 | 100 |
The fitted distribution is Poisson with \(\hat\lambda = 0.75\). Expected and observed frequencies agree closely, so the data follow a Poisson law.
For the same accident data as Experiment 3, obtain the Poisson probabilities using the recurrence relation.
To fit a Poisson distribution using the recurrence relation.
Applying it:
Blank working table (fill in each probability from the recurrence):
| x | Multiplier λ/(x+1) | P(x) |
|---|---|---|
| 0 | — | |
| 1 | ||
| 2 | ||
| 3 | ||
| 4 |
The recurrence gives the same probabilities as the direct method (0.4724, 0.3543, 0.1329, 0.0332, 0.0062), confirming the Poisson fit with \(\hat\lambda = 0.75\).
The number of claims \(X\) filed per policy-holder in a year, recorded over 200 policy-holders:
| x | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| f | 25 | 38 | 40 | 31 | 23 | 16 | 11 | 7 | 4 | 3 | 2 |
Fit a Negative Binomial distribution (\(X\) = number of failures before the \(r\)-th success).
To fit a Negative Binomial distribution by the method of moments, appropriate when the data are over-dispersed (variance > mean).
Applying it:
Blank working table (compute p(x) and the expected frequencies):
| x | p(x) | Expected f = 200 p(x) | Observed f |
|---|---|---|---|
| 0 | 25 | ||
| 1 | 38 | ||
| 2 | 40 | ||
| 3 | 31 | ||
| 4 | 23 | ||
| 5 | 16 | ||
| 6 | 11 | ||
| 7 | 7 | ||
| 8 | 4 | ||
| 9 | 3 | ||
| 10 | 2 |
\(N = 200\); \(\sum f x = 577\), so \(\bar x = 577/200 = 2.885\).
\(\sum f x^2 = 2683\), so \(s^2 = 2683/200 - 2.885^2 = 13.415 - 8.323 = 5.092\). Since \(s^2 = 5.09 > \bar x = 2.89\), the data are over-dispersed and NB is appropriate.
Method of moments: \(\hat p = 2.885/5.092 = 0.567\), \(\hat r = 2.885^2/(5.092 - 2.885) = 8.323/2.207 = 3.77\). Rounding, \(r = 4\); re-estimate \(p = 4/(4 + 2.885) = 0.581\), \(q = 0.419\).
| x | p(x) | Expected f | Observed f |
|---|---|---|---|
| 0 | 0.1139 | 22.79 | 25 |
| 1 | 0.1910 | 38.19 | 38 |
| 2 | 0.2000 | 40.01 | 40 |
| 3 | 0.1676 | 33.53 | 31 |
| 4 | 0.1229 | 24.59 | 23 |
| 5 | 0.0824 | 16.48 | 16 |
| 6 | 0.0518 | 10.36 | 11 |
| 7 | 0.0310 | 6.20 | 7 |
| 8 | 0.0179 | 3.57 | 4 |
| 9 | 0.0100 | 2.00 | 3 |
| 10 | 0.0054 | 1.09 | 2 |
(The remaining probability, \(P(X \ge 11) \approx 0.006\), accounts for the small difference from 200.)
The fitted distribution is Negative Binomial with \(r = 4,\ p = 0.581\). The expected frequencies track the observed values well, confirming that the over-dispersed claim data follow a Negative Binomial law.
For the claim data of Experiment 5 (fitted \(r = 4,\ p = 0.581,\ q = 0.419\)), obtain the Negative Binomial probabilities using the recurrence relation.
To fit a Negative Binomial distribution using the recurrence relation.
Applying it:
Blank working table (fill in each probability from the recurrence):
| x | Multiplier (r+x)q/(x+1) | P(x) |
|---|---|---|
| 0 | — | |
| 1 | ||
| 2 | ||
| 3 |
The recurrence reproduces the direct-method probabilities of Experiment 5 (0.1139, 0.1910, 0.2000, 0.1676, …), confirming the Negative Binomial fit with \(r = 4,\ p = 0.581\).
The number of failures \(X\) before the first success of a marketing call, over 100 call-sequences:
| x | 0 | 1 | 2 | 3 | 4 | 5+ |
|---|---|---|---|---|---|---|
| f | 40 | 24 | 14 | 10 | 7 | 5 |
Fit a Geometric distribution.
To fit a Geometric distribution (number of failures before the first success) to observed data.
Applying it:
Blank working table (compute p(x) and the expected frequencies):
| x | p(x) | Expected f = 100 p(x) | Observed f |
|---|---|---|---|
| 0 | 40 | ||
| 1 | 24 | ||
| 2 | 14 | ||
| 3 | 10 | ||
| 4 | 7 | ||
| 5+ | 5 |
\(N = 100\); \(\sum f x = 0+24+28+30+28+25 = 135\); \(\bar x = 1.35\).
\(\hat p = 1/(1 + 1.35) = 1/2.35 = 0.4255\), \(\hat q = 0.5745\).
| x | p(x) | Expected f | Observed f |
|---|---|---|---|
| 0 | 0.4255 | 42.55 | 40 |
| 1 | 0.2445 | 24.45 | 24 |
| 2 | 0.1404 | 14.04 | 14 |
| 3 | 0.0807 | 8.07 | 10 |
| 4 | 0.0463 | 4.63 | 7 |
| 5+ | 0.0626 | 6.26 | 5 |
The fitted Geometric distribution has \(\hat p = 0.4255\). The fit is reasonable, though the observed right tail (\(x = 3, 4\)) is a little heavier than the Geometric predicts.
For the marketing-call data of Experiment 7 (\(\hat p = 0.4255,\ \hat q = 0.5745\)), obtain the Geometric probabilities using the recurrence relation.
To fit a Geometric distribution using the recurrence relation.
Applying it:
Blank working table (fill in each probability from the recurrence):
| x | P(x) = q · P(x−1) |
|---|---|
| 0 | |
| 1 | |
| 2 | |
| 3 |
The recurrence gives the same probabilities as the direct method (0.4255, 0.2445, 0.1404, 0.0807, …), confirming the Geometric fit with \(\hat p = 0.4255\).
Lots of size \(N_L = 50\) each contain \(M = 10\) defective items. From each lot a sample of \(n = 5\) items is drawn without replacement and the number of defectives \(X\) is observed, over 500 lots:
| x | 0 | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|---|
| f | 155 | 225 | 95 | 20 | 5 | 0 |
Fit a Hypergeometric distribution.
To fit a Hypergeometric distribution to sampling-without-replacement data whose parameters are known from the sampling scheme.
Applying it:
Blank working table (compute p(x) and the expected frequencies):
| x | p(x) | Expected f = 500 p(x) | Observed f |
|---|---|---|---|
| 0 | 155 | ||
| 1 | 225 | ||
| 2 | 95 | ||
| 3 | 20 | ||
| 4 | 5 | ||
| 5 | 0 | ||
| Total | 1.000 | 500 | 500 |
| x | p(x) | Expected f | Observed f |
|---|---|---|---|
| 0 | 0.3105 | 155.3 | 155 |
| 1 | 0.4313 | 215.7 | 225 |
| 2 | 0.2098 | 104.9 | 95 |
| 3 | 0.0442 | 22.1 | 20 |
| 4 | 0.0040 | 2.0 | 5 |
| 5 | 0.0001 | 0.05 | 0 |
| Total | 1.000 | 500.0 | 500 |
Recurrence (optional): \(P(x+1) = \dfrac{(M - x)(n - x)}{(x + 1)(N_L - M - n + x + 1)}\,P(x)\).
The fitted Hypergeometric distribution (\(N_L = 50,\ M = 10,\ n = 5\)) gives expected frequencies in strong agreement with the observed ones, so the Hypergeometric is the correct model for this without-replacement sampling.