Skip to the content

Topics Covered

Concept of Dispersion Range Quartile Deviation Mean Deviation Standard Deviation Variance Coefficient of Variation Applications
On this page
  1. 1. Concept of Dispersion
  2. 2. Range
  3. 3. Quartile Deviation (QD) / Semi-Inter-Quartile Range
  4. 4. Mean Deviation (MD)
  5. 5. Standard Deviation (SD) and Variance
  6. 6. Applications of Variance and SD
  7. 7. Quick Comparison of All Measures
  8. 8. Lorenz Curve and Gini Coefficient
  9. Key Take-aways from Unit 4

1. Concept of Dispersion

DEFINITION

An average tells us where the centre of a distribution is. Two distributions can have the same average but very different spreads. Dispersion is the extent to which data values are scattered around the central value.

Bowley: "Dispersion is the measure of variation of items."

Spiegel: "The degree to which numerical data tend to spread about an average value is called the variation or dispersion of the data."

Why study Dispersion?

  1. To test reliability of an average — small dispersion ⇒ representative average.
  2. To compare two or more series.
  3. To control variability in business, pharmacy, quality control.
  4. Forms the basis of further analysis (correlation, regression, ANOVA).
Illustration: Series A (50, 50, 50) and Series B (10, 50, 90) both have mean 50, but A has zero dispersion while B is highly dispersed.
mean = 50 (both) A SD = 0 50, 50, 50 — all values sit on the mean B 10 50 90 SD ≈ 32.7
Fig 4.0 — The average alone cannot tell these two series apart; only a measure of dispersion can. This is exactly why the mean needs a companion spread statistic.

Types of Measures

Measures of Dispersion Absolute (in original units) Range Quartile Deviation Mean Deviation Standard Deviation & Variance Relative (unit-free coefficients) Coefficient of Range Coefficient of Q.D. Coefficient of M.D. Coefficient of Variation
Fig 4.0b — Classification of the measures of dispersion. Absolute measures carry the original units of the data; each has a matching relative measure — a unit-free coefficient — used to compare series measured in different units or with very different means. Every box is worked through in the sections that follow.

2. Range

DEFINITION

Range is the difference between the largest and smallest values in the data.

FORMULAE \[ \text{Range} \;=\; L - S \] \[ \text{Coefficient of Range} \;=\; \dfrac{L - S}{L + S} \]

Where \(L\) = largest value, \(S\) = smallest value.

EXAMPLE 1

Data: 12, 18, 25, 30, 45, 50.
Range = 50 − 12 = 38.
Coefficient of Range = (50 − 12)/(50 + 12) = 38/62 = 0.613.

EXAMPLE 2 (Grouped)

Class 0–10, 10–20, 20–30, 30–40, 40–50; frequencies 5, 8, 15, 7, 5.
Use boundaries: \(L = 50,\; S = 0\). Range = 50. Coefficient = 50/50 = 1.

Merits: simple, easy to understand. Demerits: based on only two extreme values; affected severely by outliers; not based on all observations.

3. Quartile Deviation (QD) / Semi-Inter-Quartile Range

DEFINITION

Quartile Deviation is half the difference between the third and first quartiles. It measures the spread of the middle 50 % of data.

FORMULAE \[ \text{QD} \;=\; \dfrac{Q_3 - Q_1}{2} \] \[ \text{Coefficient of QD} \;=\; \dfrac{Q_3 - Q_1}{Q_3 + Q_1} \]

For grouped data, the quartile formula is:

\[ Q_k \;=\; L + \dfrac{\dfrac{kN}{4} - C}{f}\,h, \quad k = 1, 2, 3. \]
EXAMPLE 1 (Ungrouped)

Data: 5, 7, 9, 10, 12, 15, 18, 20, 25 (n = 9).
\(Q_1\) at position \((n+1)/4 = 2.5\) ⇒ between 2nd (7) and 3rd (9) ⇒ \(Q_1 = 8\).
\(Q_3\) at \(3(n+1)/4 = 7.5\) ⇒ between 7th (18) and 8th (20) ⇒ \(Q_3 = 19\).
QD = (19 − 8)/2 = 5.5; Coeff. of QD = (19 − 8)/(19 + 8) = 11/27 = 0.407.

EXAMPLE 2 (Grouped)
ClassfCF
0–1055
10–20813
20–301225
30–401035
40–50540

\(N = 40\); \(N/4 = 10\) ⇒ \(Q_1\) class is 10–20; \(L=10,C=5,f=8,h=10\). \(Q_1 = 10 + (10-5)/8 \times 10 = 16.25\).

\(3N/4 = 30\) ⇒ \(Q_3\) class is 30–40; \(L=30,C=25,f=10,h=10\). \(Q_3 = 30 + (30-25)/10 \times 10 = 35\).

QD = (35 − 16.25)/2 = 9.375; Coeff. = 18.75/51.25 = 0.366.

Merits: not affected by extreme values; useful for open-ended classes. Demerits: ignores extreme 50 % of data; not based on all observations.

4. Mean Deviation (MD)

DEFINITION

Mean Deviation is the arithmetic mean of the absolute deviations of observations from any central value (mean, median or mode). MD about median is the least.

UNGROUPED DATA \[ \text{MD}_{\bar x} = \dfrac{\sum |x_i - \bar{x}|}{n}, \quad \text{MD}_{M} = \dfrac{\sum |x_i - M|}{n} \]
FREQUENCY DATA \[ \text{MD}_A = \dfrac{\sum f_i |x_i - A|}{N} \]

where \(A\) = the chosen average (mean, median or mode).

COEFFICIENT \[ \text{Coeff. of MD} \;=\; \dfrac{\text{MD}}{A} \]
EXAMPLE 1 (about Mean)

Data: 4, 6, 8, 10, 12. Mean = 8.
Deviations |x − 8|: 4, 2, 0, 2, 4. Sum = 12.
MD = 12/5 = 2.4; Coefficient = 2.4/8 = 0.3.

EXAMPLE 2 (about Median, frequency)
xf|x − M|f|x − M|
1031030
155525
20 (Median)700
254520
3011010
Total2085

MDM = 85 / 20 = 4.25.

Merits: uses all observations, easy to understand. Demerits: uses absolute values, hence not suitable for further algebraic treatment.

5. Standard Deviation (SD) and Variance

DEFINITION

Standard Deviation is the positive square root of the arithmetic mean of the squared deviations of observations from their mean. The square of the SD is called Variance. Symbol \(\sigma\).

5.1 Formulae

UNGROUPED \[ \sigma \;=\; \sqrt{\dfrac{\sum (x_i - \bar{x})^2}{n}} \quad\;\; \text{Variance } \sigma^2 \;=\; \dfrac{\sum (x_i - \bar{x})^2}{n} \]
FREQUENCY DATA \[ \sigma \;=\; \sqrt{\dfrac{\sum f_i (x_i - \bar{x})^2}{N}} \]
SHORTCUT (computational) \[ \sigma^2 \;=\; \dfrac{\sum f_i x_i^2}{N} - \left(\dfrac{\sum f_i x_i}{N}\right)^{\!2} \;=\; \dfrac{\sum f_i x_i^2}{N} - \bar{x}^2 \]
DERIVATION — where the shortcut comes from

Expand the squared deviation and use \(\sum f_i x_i = N\bar{x}\):

\[ \sigma^2 = \frac{1}{N}\sum f_i (x_i - \bar{x})^2 = \frac{1}{N}\sum f_i\big(x_i^2 - 2\bar{x}x_i + \bar{x}^2\big) = \frac{\sum f_i x_i^2}{N} - 2\bar{x}\!\underbrace{\frac{\sum f_i x_i}{N}}_{=\,\bar{x}} + \bar{x}^2 = \frac{\sum f_i x_i^2}{N} - \bar{x}^2. \]

Because \(\sigma^2 = \overline{x^2} - \bar{x}^2\) is a difference of two non-negative-weighted means, it is always \(\ge 0\) — the variance can never be negative, a useful check on any calculation. It also proves \(\overline{x^2} \ge \bar{x}^2\) (a special case of the AM inequality from Unit 3).

STEP-DEVIATION

With \(u_i = (x_i - A)/h\):

\[ \sigma \;=\; h \sqrt{\dfrac{\sum f_i u_i^2}{N} - \left(\dfrac{\sum f_i u_i}{N}\right)^{\!2}} \]
COEFFICIENT OF SD & CV \[ \text{Coefficient of SD} = \dfrac{\sigma}{\bar{x}}, \qquad \text{Coefficient of Variation (CV)} = \dfrac{\sigma}{\bar{x}} \times 100 \% \]
EXAMPLE 1 (Ungrouped)

Data: 4, 8, 6, 10, 12. \(\bar x = 40/5 = 8\).
Deviations: −4, 0, −2, 2, 4. Squared: 16, 0, 4, 4, 16. Sum = 40.
Variance = 40/5 = 8; SD = \(\sqrt{8} \approx 2.828\). CV = 2.828/8 × 100 = 35.36 %.

EXAMPLE 2 (Grouped)
Classfxfxfx²
0–105525125
10–208151201 800
20–3015253759 375
30–407352458 575
40–5054522510 125
Total4099030 000

\(\bar x = 990/40 = 24.75\). Variance = \(30000/40 - 24.75^2 = 750 - 612.5625 = 137.4375\).

SD = \(\sqrt{137.4375} \approx \mathbf{11.72}\). CV = 11.72/24.75 × 100 ≈ 47.4 %.

5.2 Properties of Standard Deviation

  1. SD is independent of change of origin but not of scale: if \(y = a + bx\), then \(\sigma_y = |b| \sigma_x\).
  2. SD is the least RMS deviation; \(\sigma^2 \le \dfrac{\sum (x - A)^2}{n}\) for any \(A\); equality at \(A = \bar x\).
  3. SD of \(n\) consecutive natural numbers \(1,2,\ldots,n\) is \(\sigma = \sqrt{\dfrac{n^2 - 1}{12}}\).
  4. Combined SD of two groups: \[ \sigma_{12}^2 = \dfrac{n_1(\sigma_1^2 + d_1^2) + n_2(\sigma_2^2 + d_2^2)}{n_1 + n_2} \] where \(d_i = \bar x_i - \bar x_{12}\).
  5. SD ≥ MD ≥ QD generally (for a normal distribution: QD : MD : SD ≈ 10 : 12 : 15).

5.3 Coefficient of Variation (CV)

USE

CV is a unit-free relative measure used to compare variability of two or more series. The series with smaller CV is more consistent / less variable.

\[ \text{CV} = \dfrac{\sigma}{\bar x}\times 100\ \%. \]
EXAMPLE 1 (Comparing Consistency)

Batsman A: mean = 50, SD = 12. Batsman B: mean = 40, SD = 8.
CVA = 12/50 × 100 = 24 %; CVB = 8/40 × 100 = 20 %.
B is more consistent (lower CV) although A has higher average.

EXAMPLE 2 (Combined SD)

Group 1: \(n_1 = 50,\; \bar x_1 = 60,\; \sigma_1 = 5\). Group 2: \(n_2 = 100,\; \bar x_2 = 70,\; \sigma_2 = 8\).
Combined mean \(\bar x_{12} = (50\cdot 60 + 100\cdot 70)/150 = 10000/150 = 66.67\).
\(d_1 = 60-66.67 = -6.67;\; d_2 = 70-66.67 = 3.33\).
\(\sigma_{12}^2 = \dfrac{50(25 + 44.49) + 100(64 + 11.09)}{150} = \dfrac{3474.5 + 7509}{150} = 73.22\).
\(\sigma_{12} \approx \mathbf{8.56}\).

6. Applications of Variance and SD

In Business

In Pharmacy / Bio-medical

Empirical Rule (Normal data)

μ−3σμ−2σμ−σ μ μ+σμ+2σμ+3σ 68% 95% 99.7%
Fig 4.1 — The empirical (68–95–99.7) rule for a normal distribution: successively wider bands of \(\pm 1\sigma\), \(\pm 2\sigma\), \(\pm 3\sigma\) capture 68 %, 95 % and 99.7 % of the data. This is what makes \(\bar x \pm 2\sigma\) a natural reference range (≈ 95 %) in the applications above.

7. Quick Comparison of All Measures

MeasureUses all data?Affected by outliers?Algebraic treatment?
RangeNoYes (severely)No
Quartile DeviationNoNoLimited
Mean DeviationYesLess than SDLimited (absolute values)
Standard DeviationYesYesYes (best)

8. Lorenz Curve and Gini Coefficient

The Lorenz curve plots the cumulative share of a variable (e.g. income) against the cumulative share of the population, both from lowest to highest. The line of equal distribution is the 45° diagonal; the farther the Lorenz curve bows below it, the greater the inequality.

The Gini coefficient measures this gap: it is twice the area between the diagonal and the Lorenz curve, ranging from 0 (perfect equality) to 1 (perfect inequality). For \(n\) values \(x_1,\dots,x_n\) with mean \(\bar x\),

\[ G = \dfrac{\displaystyle\sum_{i=1}^{n}\sum_{j=1}^{n}|x_i - x_j|}{2\,n^{2}\,\bar x}. \]
EXAMPLE

Incomes of five persons: 10, 20, 30, 40, 50 (mean = 30). The mean absolute difference gives \(G = 0.267\), a moderate level of inequality. If all incomes were equal, \(G = 0\).

Key Take-aways from Unit 4

Extra Practical Problems

PRACTICE

Additional worked problems with step-by-step procedures to support self-study, matching this unit's topics.

STEP-BY-STEP PROCEDURE
  1. Range \(= \) max value \(-\) min value.
  2. Quartile Deviation: find \(Q_1\) and \(Q_3\) (use \(\frac{i(n+1)}{4}\)th item for ungrouped, or \(Q_i = l + \frac{\frac{iN}{4}-C}{f}h\) for grouped); then \(QD = \dfrac{Q_3-Q_1}{2}\).
  3. Mean Deviation: choose an average \(A\) (mean, median or mode); compute \(MD = \dfrac{\sum |X_i - A|}{n}\) (or \(\dfrac{\sum f_i|X_i-A|}{N}\) for frequency data). MD about the median is least.
  4. Standard Deviation: \(\sigma = \sqrt{\dfrac{\sum (X_i-\bar X)^2}{n}}\) (ungrouped) or \(\sqrt{\dfrac{\sum f_i(X_i-\bar X)^2}{N}}\) (grouped). Variance \(= \sigma^2\).
  5. Relative measures: Coefficient of Variation \(CV = \dfrac{\sigma}{\bar X}\times 100\) — lower CV ⇒ more consistent.
  6. Combined SD of two groups: \(\sigma^2 = \dfrac{1}{n_1+n_2}\big[n_1(\sigma_1^2+d_1^2) + n_2(\sigma_2^2+d_2^2)\big]\), where \(d_1 = \bar X_1 - \bar X\), \(d_2 = \bar X_2 - \bar X\) and \(\bar X\) is the combined mean.

Problem 1 — All Measures of Dispersion (ungrouped)

DATA

10, 7, 5, 9, 9, 10, 7, 3, 12 → sorted 3, 5, 7, 7, 9, 9, 10, 10, 12; \(n=9\), mean \(=8\).

Problem 2 — All Measures of Dispersion (grouped)

DATA

Age distribution of 542 members: classes 20–30 … 80–90 with frequencies 3, 61, 132, 153, 140, 51, 2.

Problem 3 — Coefficient of Variation (consistency of two series)

DATA

Goals scored per match by teams A and B. For A: mean \(= 1.05\), SD \(= 1.31\) → \(CV_A = \dfrac{1.31}{1.05}\times 100 = 124.76\). For B: mean \(= 1.2\), SD \(= 1.30\) → \(CV_B = \dfrac{1.30}{1.2}\times 100 = 108.33\).

Since \(CV_B < CV_A\), team B is more consistent.

Problem 4 — Comparison of Two Firms (combined mean & variance)

DATA

Firm A: \(n_A=586\), mean ₹52.50, variance 100. Firm B: \(n_B=648\), mean ₹47.50, variance 121.

Problem 5 — SD of a Combined Sample

DATA

Sample 1: \(n_1=100\), mean 15, SD 3. Whole group: \(n=250\), mean 15.6, SD \(\sqrt{13.44}\). Find the SD of sample 2.

\(n_2 = 150\); from the combined mean, \(\bar x_2 = 16\); \(d_1 = -0.6, d_2 = 0.4\). Substituting into the combined-variance formula and solving gives \(\sigma_2 = \) 4.

Problem 6 — Corrected Mean & SD

DATA

200 candidates: mean 40, SD 15. Scores 43 and 35 were misread as 34 and 53. Correct them.

Corrected \(\sum X = 8000 - (34+53) + (43+35) = 7991\) → corrected mean \(= 7991/200 = 39.95\).

\(\sum X^2 = 200(225+1600) = 365000\); corrected \(\sum X^2 = 365000 - (34^2+53^2) + (43^2+35^2) = 364109\); corrected SD \(= \sqrt{364109/200 - 39.95^2} = \sqrt{224.54} = 14.98\).

REMEMBER

Unsolved Exercises

PRACTICE
  1. Variance of (i) 5,5,5,5,5 (ii) 4,5,6. (Ans: 0; 0.67)
  2. Mean 50, SD 10 for 10 figures. New mean/SD if every figure is (i) +4, (ii) ×2, (iii) ×2 then −4. (Ans: 54,10; 100,20; 96,20)
  3. MD and SD: classes 0–5…35–40, freq 2, 5, 7, 13, 21, 16, 8, 3. (Ans: MD 6.23, SD 8.00 using ÷N)
  4. Mean of 100 obs = 50, CV = 40%. Find SD. (Ans: 20)
  5. Mean 17, variance 33 for 10 figures; remove the inaccurate figure 26. New mean & variance of 9. (Ans: 16, 26.67)
  6. Samples size 50 & 100, means 54.1 & 50.3, SDs 8 & 7. Combined mean & SD of 150. (Ans: 51.57, 7.56)
  7. Firm A (500, ₹186, var 81) vs Firm B (600, ₹175, var 100): larger payroll? more variable? combined mean & variance? (Ans: B; B; ₹180, 121.36)