An average tells us where the centre of a distribution is. Two distributions can have the
same average but very different spreads. Dispersion is the extent to which
data values are scattered around the central value.
Bowley: "Dispersion is the measure of variation of items."
Spiegel: "The degree to which numerical data tend to spread about an average value is
called the variation or dispersion of the data."
Why study Dispersion?
To test reliability of an average — small dispersion ⇒ representative average.
To compare two or more series.
To control variability in business, pharmacy, quality control.
Forms the basis of further analysis (correlation, regression, ANOVA).
Illustration: Series A (50, 50, 50) and Series B (10, 50, 90) both have mean 50, but
A has zero dispersion while B is highly dispersed.
Fig 4.0 — The average alone cannot tell these two series apart; only a measure of
dispersion can. This is exactly why the mean needs a companion spread statistic.
Types of Measures
Absolute measures — expressed in original units (Range, QD, MD, SD).
Relative measures — unit-free coefficients used to compare different series
(Coefficient of Range, Coefficient of QD, Coefficient of MD, Coefficient of Variation).
Fig 4.0b — Classification of the measures of dispersion.Absolute measures carry the original units of the data; each has a matching relative measure — a unit-free coefficient — used to compare series measured in different units or with very different means. Every box is worked through in the sections that follow.
2. Range
DEFINITION
Range is the difference between the largest and smallest values in the data.
FORMULAE
\[
\text{Range} \;=\; L - S
\]
\[
\text{Coefficient of Range} \;=\; \dfrac{L - S}{L + S}
\]
Where \(L\) = largest value, \(S\) = smallest value.
EXAMPLE 1
Data: 12, 18, 25, 30, 45, 50.
Range = 50 − 12 = 38.
Coefficient of Range = (50 − 12)/(50 + 12) = 38/62 = 0.613.
EXAMPLE 2 (Grouped)
Class 0–10, 10–20, 20–30, 30–40, 40–50; frequencies 5, 8, 15, 7, 5.
Use boundaries: \(L = 50,\; S = 0\). Range = 50. Coefficient = 50/50 = 1.
Merits: simple, easy to understand. Demerits: based on only two
extreme values; affected severely by outliers; not based on all observations.
3. Quartile Deviation (QD) / Semi-Inter-Quartile Range
DEFINITION
Quartile Deviation is half the difference between the third and first quartiles.
It measures the spread of the middle 50 % of data.
Merits: not affected by extreme values; useful for open-ended classes.
Demerits: ignores extreme 50 % of data; not based on all observations.
4. Mean Deviation (MD)
DEFINITION
Mean Deviation is the arithmetic mean of the absolute deviations of observations
from any central value (mean, median or mode). MD about median is the least.
Merits: uses all observations, easy to understand.
Demerits: uses absolute values, hence not suitable for further algebraic treatment.
5. Standard Deviation (SD) and Variance
DEFINITION
Standard Deviation is the positive square root of the arithmetic mean of the squared deviations
of observations from their mean. The square of the SD is called Variance. Symbol \(\sigma\).
Because \(\sigma^2 = \overline{x^2} - \bar{x}^2\) is a difference of two non-negative-weighted means,
it is always \(\ge 0\) — the variance can never be negative, a useful check on any calculation.
It also proves \(\overline{x^2} \ge \bar{x}^2\) (a special case of the AM inequality from Unit 3).
Quality control — control charts use \(\bar x \pm 3\sigma\) limits to flag defects.
Investment / Finance — SD of returns measures risk; portfolios are compared by CV.
Sales forecasting — SD of forecast errors quantifies forecasting accuracy.
Market research — variability in customer ratings.
In Pharmacy / Bio-medical
Drug potency — manufacturing tolerance specified as mean ± a multiple of SD.
Clinical trials — sample size depends on SD of the response.
Bio-equivalence tests rely on variance of pharmacokinetic parameters.
Reference ranges for laboratory tests use mean ± 2 SD (covers ≈ 95 % of healthy population).
Empirical Rule (Normal data)
About 68 % of data lies in \(\bar x \pm \sigma\).
About 95 % lies in \(\bar x \pm 2\sigma\).
About 99.7 % lies in \(\bar x \pm 3\sigma\).
Fig 4.1 — The empirical (68–95–99.7) rule for a normal distribution: successively wider
bands of \(\pm 1\sigma\), \(\pm 2\sigma\), \(\pm 3\sigma\) capture 68 %, 95 % and 99.7 % of the data.
This is what makes \(\bar x \pm 2\sigma\) a natural reference range (≈ 95 %) in the applications above.
7. Quick Comparison of All Measures
Measure
Uses all data?
Affected by outliers?
Algebraic treatment?
Range
No
Yes (severely)
No
Quartile Deviation
No
No
Limited
Mean Deviation
Yes
Less than SD
Limited (absolute values)
Standard Deviation
Yes
Yes
Yes (best)
8. Lorenz Curve and Gini Coefficient
The Lorenz curve plots the cumulative share of a variable (e.g. income) against
the cumulative share of the population, both from lowest to highest. The line of equal
distribution is the 45° diagonal; the farther the Lorenz curve bows below it, the greater the
inequality.
The Gini coefficient measures this gap: it is twice the area between the diagonal
and the Lorenz curve, ranging from 0 (perfect equality) to 1 (perfect inequality). For \(n\) values
\(x_1,\dots,x_n\) with mean \(\bar x\),
\[ G = \dfrac{\displaystyle\sum_{i=1}^{n}\sum_{j=1}^{n}|x_i - x_j|}{2\,n^{2}\,\bar x}. \]
EXAMPLE
Incomes of five persons: 10, 20, 30, 40, 50 (mean = 30). The mean absolute difference gives
\(G = 0.267\), a moderate level of inequality. If all incomes were equal, \(G = 0\).
Key Take-aways from Unit 4
Dispersion measures spread; Range is crudest, SD is best.
Empirical rule: 68 %–95 %–99.7 % rule for normal data.
Extra Practical Problems
PRACTICE
Additional worked problems with step-by-step procedures to support self-study, matching this unit's topics.
STEP-BY-STEP PROCEDURE
Range \(= \) max value \(-\) min value.
Quartile Deviation: find \(Q_1\) and \(Q_3\) (use \(\frac{i(n+1)}{4}\)th item for ungrouped, or
\(Q_i = l + \frac{\frac{iN}{4}-C}{f}h\) for grouped); then \(QD = \dfrac{Q_3-Q_1}{2}\).
Mean Deviation: choose an average \(A\) (mean, median or mode); compute
\(MD = \dfrac{\sum |X_i - A|}{n}\) (or \(\dfrac{\sum f_i|X_i-A|}{N}\) for frequency data). MD about the median is least.
Standard Deviation: \(\sigma = \sqrt{\dfrac{\sum (X_i-\bar X)^2}{n}}\) (ungrouped) or
\(\sqrt{\dfrac{\sum f_i(X_i-\bar X)^2}{N}}\) (grouped). Variance \(= \sigma^2\).
Relative measures: Coefficient of Variation \(CV = \dfrac{\sigma}{\bar X}\times 100\) — lower CV ⇒ more consistent.
Combined SD of two groups:
\(\sigma^2 = \dfrac{1}{n_1+n_2}\big[n_1(\sigma_1^2+d_1^2) + n_2(\sigma_2^2+d_2^2)\big]\), where
\(d_1 = \bar X_1 - \bar X\), \(d_2 = \bar X_2 - \bar X\) and \(\bar X\) is the combined mean.
Problem 1 — All Measures of Dispersion (ungrouped)
Mean Deviation \(= \dfrac{5152}{542} = 9.51\) yrs.
Standard Deviation \(= \sqrt{\dfrac{76458.49}{542}} = \sqrt{141.07} = 11.88\) yrs.
Problem 3 — Coefficient of Variation (consistency of two series)
DATA
Goals scored per match by teams A and B. For A: mean \(= 1.05\), SD \(= 1.31\) →
\(CV_A = \dfrac{1.31}{1.05}\times 100 = 124.76\). For B: mean \(= 1.2\), SD \(= 1.30\) →
\(CV_B = \dfrac{1.30}{1.2}\times 100 = 108.33\).
Since \(CV_B < CV_A\), team B is more consistent.
Problem 4 — Comparison of Two Firms (combined mean & variance)
DATA
Firm A: \(n_A=586\), mean ₹52.50, variance 100. Firm B: \(n_B=648\), mean ₹47.50, variance 121.
Total wages: A \(= 586\times52.50 = 30765\); B \(= 648\times47.50 = 30780\). Firm B pays more.
\(CV_A = \dfrac{10}{52.50}\times100 = 19.05\); \(CV_B = \dfrac{11}{47.50}\times100 = 23.16\) → B more variable.
Combined mean \(\bar x = \dfrac{586(52.50)+648(47.50)}{1234} = 49.87\).
Sample 1: \(n_1=100\), mean 15, SD 3. Whole group: \(n=250\), mean 15.6, SD \(\sqrt{13.44}\).
Find the SD of sample 2.
\(n_2 = 150\); from the combined mean, \(\bar x_2 = 16\); \(d_1 = -0.6, d_2 = 0.4\). Substituting into
the combined-variance formula and solving gives \(\sigma_2 = \) 4.
Problem 6 — Corrected Mean & SD
DATA
200 candidates: mean 40, SD 15. Scores 43 and 35 were misread as 34 and 53. Correct them.
Corrected \(\sum X = 8000 - (34+53) + (43+35) = 7991\) → corrected mean \(= 7991/200 = 39.95\).