A measure of central tendency is a single value that represents the entire data by indicating the centre around which the values cluster. Also called an average.
The five averages prescribed in the syllabus: Arithmetic Mean (AM), Median, Mode, Geometric Mean (GM), Harmonic Mean (HM).
The arithmetic mean is the sum of all observations divided by the number of observations.
Take \(x_i\) = mid-value of class \(i\):
\[ \bar{x} \;=\; \dfrac{\sum f_i x_i}{N} \]Let \(A\) = assumed mean, \(d_i = x_i - A\):
\[ \bar{x} \;=\; A + \dfrac{\sum f_i d_i}{N} \]For equal class width \(h\), let \(u_i = (x_i - A)/h\):
\[ \bar{x} \;=\; A + \dfrac{\sum f_i u_i}{N}\, h \]Marks of 7 students: 45, 60, 55, 70, 65, 50, 80.
\(\sum x = 425;\; n = 7\). \(\bar{x} = 425/7 = 60.71\).
| Class | f | Mid x | u = (x − 45)/10 | fu |
|---|---|---|---|---|
| 20–30 | 5 | 25 | −2 | −10 |
| 30–40 | 8 | 35 | −1 | −8 |
| 40–50 | 15 | 45 (A) | 0 | 0 |
| 50–60 | 10 | 55 | 1 | 10 |
| 60–70 | 2 | 65 | 2 | 4 |
| Total | 40 | — | — | −4 |
\(\bar{x} = 45 + (-4/40)\cdot 10 = 45 - 1 = \mathbf{44}\).
Property 1 is a direct consequence of the definition \(\bar{x} = \tfrac{1}{n}\sum x_i\):
\[ \sum_{i=1}^{n}(x_i - \bar{x}) \;=\; \sum x_i - n\bar{x} \;=\; n\bar{x} - n\bar{x} \;=\; 0. \]
This is exactly why the mean is the natural "balance point" of the data: the positive and negative deviations cancel. Property 2 sharpens it — differentiating \(S(A)=\sum (x_i-A)^2\) and setting \(S'(A) = -2\sum(x_i-A) = 0\) gives \(A=\bar{x}\), so the sum of squared deviations is least about the mean (the basis of the variance in Unit 4 and least-squares regression later).
Section A has 30 students with mean 60; Section B has 20 students with mean 70.
\(\bar{x}_{12} = \dfrac{30(60)+20(70)}{30+20} = \dfrac{1800+1400}{50} = \dfrac{3200}{50} = 64\).
Data: 4, 8, 6, 10, 12. Mean = 40/5 = 8.
Deviations: −4, 0, −2, 2, 4. Sum = 0. ✓
| Merits | Demerits |
|---|---|
| Rigidly defined; based on all observations; easy to compute; suitable for algebraic treatment. | Highly affected by extreme values; cannot be obtained graphically; not suitable for qualitative data; cannot be computed if any value is missing or open-ended class. |
The median is the middle value of an arranged data set. It divides the distribution into two equal halves — half the values are less than the median and half are more.
Arrange data in ascending order. Then:
Form cumulative frequencies. Locate the value corresponding to \(\dfrac{N+1}{2}\)-th cumulative frequency.
Median class = the class containing the \(\dfrac{N}{2}\)-th observation.
Data: 12, 4, 9, 7, 15, 18, 11. (n = 7, odd)
Sorted: 4, 7, 9, 11, 12, 15, 18. Median = 4th term = 11.
Marks of 50 students:
| Class | f | CF |
|---|---|---|
| 0–10 | 5 | 5 |
| 10–20 | 8 | 13 |
| 20–30 | 12 | 25 ← median class |
| 30–40 | 15 | 40 |
| 40–50 | 10 | 50 |
\(N/2 = 25\). Median class = 20–30 (since CF first reaches 25 here). \(L = 20,\; C = 13,\; f = 12,\; h = 10\).
Median = \(20 + \dfrac{25 - 13}{12}\times 10 = 20 + 10 = \mathbf{30}\) marks.
Draw both ogives (less-than and more-than). The X-coordinate of their point of intersection = median. Alternatively, on a less-than ogive, locate \(N/2\) on the Y-axis, draw a horizontal line to meet the curve, then drop a perpendicular to the X-axis.
| Merits | Demerits |
|---|---|
| Not affected by extreme values; can be located graphically; suitable for qualitative ranks; can be found in open-ended classes. | Not based on all observations; not suitable for further algebraic treatment; needs data to be arranged. |
The mode is the value that occurs most frequently in a data set. A distribution may have one mode (unimodal), two modes (bimodal) or more (multimodal).
Data: 3, 5, 7, 5, 8, 5, 6, 9, 5. The value 5 appears 4 times — most often. Mode = 5.
Use the same table as Median Example 2. Modal class = 30–40 (highest f = 15).
\(L = 30,\; f_1 = 15,\; f_0 = 12,\; f_2 = 10,\; h = 10\).
Mode = \(30 + \dfrac{15-12}{2(15)-12-10}\times 10 = 30 + \dfrac{3}{8}\times 10 = 30 + 3.75 = \mathbf{33.75}\).
Draw the histogram of the data. In the modal rectangle, draw two diagonals from the top corners of the modal rectangle to the corresponding top corners of the adjacent rectangles. Drop a perpendicular from their intersection to the X-axis — that point is the mode.
| Merits | Demerits |
|---|---|
| Not affected by extreme values; easy to locate visually; useful for qualitative data (most popular brand). | Not rigidly defined; not based on all observations; may not exist or may be multiple; not suitable for algebraic treatment. |
Valid for a moderately skewed distribution.
If Mean = 50 and Median = 45, then Mode = 3(45) − 2(50) = 135 − 100 = 35.
If Mode = 30 and Median = 36, then 30 = 3(36) − 2 Mean ⇒ 2 Mean = 108 − 30 = 78 ⇒ Mean = 39.
The geometric mean of \(n\) positive observations is the \(n\)-th root of their product.
GM of 4, 8, 16. \(\text{GM} = \sqrt[3]{4\cdot 8\cdot 16} = \sqrt[3]{512} = 8\).
A company's sales rose 10 % then 20 % then 30 % in three successive years. The average annual growth factor is \(\sqrt[3]{1.10 \times 1.20 \times 1.30}\) = \(\sqrt[3]{1.716} \approx 1.197\). Average growth rate ≈ 19.7 %.
Use: for ratios, percentages, growth rates, compound interest. Limitation: cannot be computed if any value is zero or negative.
The harmonic mean is the reciprocal of the arithmetic mean of the reciprocals of the observations.
HM of 2, 4, 8.
\(\sum 1/x = 1/2 + 1/4 + 1/8 = 7/8\). HM = \(3 \div (7/8) = 24/7 \approx 3.43\).
A man covers half the distance at 40 km/h and half at 60 km/h. Average speed for equal distances at different speeds is the harmonic mean:
HM = \(\dfrac{2}{\frac{1}{40}+\frac{1}{60}} = \dfrac{2}{\frac{5}{120}} = 48\) km/h.
Use: for averaging rates and ratios (speeds, prices per unit). Limitation: cannot be used if any value is zero.
Equality holds only when all observations are equal.
For two values \(a, b\) this falls straight out of the definitions:
\[ \text{AM}\times\text{HM} \;=\; \frac{a+b}{2}\cdot\frac{2ab}{a+b} \;=\; ab \;=\; \big(\sqrt{ab}\,\big)^2 \;=\; \text{GM}^2. \]
So the geometric mean is itself the geometric mean of the other two averages, \(\text{GM} = \sqrt{\text{AM}\times\text{HM}}\) — which is how Example 2 recovers a missing average.
For 4 and 16: AM = 10, GM = \(\sqrt{64}=8\), HM = \(2/(1/4+1/16) = 2 \times 16/5 = 6.4\).
Check: \(10 \geq 8 \geq 6.4\) ✓ and \(\text{GM}^2 = 64 = 10 \times 6.4\) ✓.
If AM = 25 and HM = 16, then GM = \(\sqrt{25 \times 16} = \sqrt{400} = 20\).
| Situation | Best Average |
|---|---|
| Symmetric data, no outliers | Arithmetic Mean |
| Skewed data or outliers present | Median |
| Most fashionable / popular item | Mode |
| Growth rates, ratios, percentages | Geometric Mean |
| Average speeds, prices, rates | Harmonic Mean |
| Open-ended class intervals | Median or Mode |
Additional worked problems with step-by-step procedures to support self-study, matching this unit's topics.
10, 7, 11, 9, 9, 10, 7, 9, 12.
\(\bar X = \dfrac{\sum X_i}{n} = \dfrac{10+7+11+9+9+10+7+9+12}{9} = \dfrac{84}{9} = 9.33.\)
| \(X_i\) | 2 | 9 | 16 | 35 | 32 | 89 | 95 | 65 | 55 |
|---|---|---|---|---|---|---|---|---|---|
| \(f_i\) | 8 | 2 | 5 | 7 | 6 | 8 | 9 | 6 | 2 |
\(\bar X = \dfrac{\sum f_i X_i}{\sum f_i} = \dfrac{2618}{53} = 49.40.\)
Classes 0–10, 10–20, …, 80–90 with frequencies 8, 2, 5, 7, 6, 8, 9, 6, 2. Using mid-points \(X_i\): \(\sum f_i = 53\), \(\sum f_i X_i = 2355\).
\(\bar X = \dfrac{2355}{53} = 44.43.\)
Series 1: 5 numbers, mean 40. Series 2: 4 numbers, mean 50.
\(\bar X_{12} = \dfrac{n_1\bar X_1 + n_2\bar X_2}{n_1+n_2} = \dfrac{5(40)+4(50)}{9} = \dfrac{400}{9} = 44.44.\)
Odd: 5, 20, 15, 35, 18, 25, 40 → sorted 5, 15, 18, 20, 25, 35, 40; \(n=7\); median = \(\left(\frac{n+1}{2}\right)\)th = 4th term = 20.
Even: 8, 20, 50, 25, 15, 30 → sorted 8, 15, 20, 25, 30, 50; \(n=6\); median = mean of 3rd and 4th = \(\frac{20+25}{2} = \) 22.5.
\(X\): 1–9 with \(f\): 8, 10, 11, 16, 20, 25, 15, 9, 6 (\(N=120\)); cumulative 8, 18, 29, 45, 65, 90, 105, 114, 120.
\(N/2 = 60\); c.f. just greater than 60 is 65, whose \(X\)-value is 5. Hence median = 5.
Wages 20–30 … 80–90 with labourers 3, 5, 20, 10, 5, 7, 2; cumulative 3, 8, 28, 38, 43, 50, 52.
\(N/2 = 26\); c.f. just above is 28 → median class 40–50, with \(l=40, C=8, f=20, h=10\):
\(\text{Median} = l + \dfrac{N/2 - C}{f}\cdot h = 40 + \dfrac{26-8}{20}\cdot 10 = \) ₹49.
4, 2, 4, 3, 2, 2, 1, 2. The value 2 occurs four times — more than any other — so mode = 2.
Size \(X\): 1–12 with frequency \(f\): 3, 8, 15, 23, 35, 40, 32, 28, 20, 45, 14, 6.
The frequency 45 (at \(X=10\)) is out of step with the otherwise single-peak pattern, so the mode is not simply 10. The grouping method combines frequencies in 1s, 2s and 3s:
| Column | Max frequency | Value(s) of X |
|---|---|---|
| (i) singles | 45 | 10 |
| (ii) 2s | 75 | 5, 6 |
| (iii) 2s (offset) | 72 | 6, 7 |
| (iv) 3s | 98 | 4, 5, 6 |
| (v) 3s (offset) | 107 | 5, 6, 7 |
| (vi) 3s (offset) | 100 | 6, 7, 8 |
The value 6 appears the most times across the columns, so mode = 6 (10 was an irregular item).
Classes 0–10 … 70–80 with frequencies 5, 8, 7, 12, 28, 20, 10, 10. Maximum frequency 28 → modal class 40–50, \(l=40, f_1=28, f_0=12, f_2=20, h=10\):
\(\text{Mode} = l + \dfrac{f_1 - f_0}{2f_1 - f_0 - f_2}\cdot h = 40 + \dfrac{28-12}{56-12-20}\cdot 10 = 40 + 6.667 = \) 46.667.
Ungrouped: 3, 13, 11, 15, 5, 4, 2 (\(n=7\)). \(\log G = \frac1n\sum\log X_i = \frac{5.4106}{7} = 0.772944\); \(G = \text{antilog}(0.772944) = \) 5.928.
Grouped: classes 0–10, 10–20, 20–30, 30–40 with \(f\) = 1, 3, 4, 2 (\(N=10\)), mid-values 5, 15, 25, 35. \(\log G = \dfrac{1}{N}\sum f_i\log X_i = \dfrac{12.91}{10} = 1.29\); \(G = \text{antilog}(1.29) = \) 19.53.
Ungrouped: 10, 7, 11, 9, 9, 10, 7, 9, 12. \(H = \dfrac{n}{\sum (1/X_i)} = \) 9.06.
Grouped: ages 19–26 with students 5, 8, 7, 12, 28, 20, 10, 10. \(H = \dfrac{N}{\sum (f_i/X_i)} = \dfrac{100}{4.377} = \) 22.84.
Average speed: a cyclist goes at 10 mph and returns at 15 mph; the average speed is the harmonic mean \(H = \dfrac{2}{\frac1{10}+\frac1{15}} = \) 12 mph (not the simple mean 12.5).
3, 13, 11, 11, 5, 4, 2 → sorted 2, 3, 4, 5, 11, 11, 13; \(n=7\).
\(Q_1 = \left(\frac{1(n+1)}{4}\right)\)th = 2nd value = 3.
\(D_3 = \left(\frac{3(n+1)}{10}\right)\)th = 2.4th value \(= 3 + 0.4(4-3) = \) 3.4.
\(P_{20} = \left(\frac{20(n+1)}{100}\right)\)th = 1.6th value \(= 2 + 0.6(3-2) = \) 2.6.
Eight coins tossed 256 times; number of heads \(x\): 0–8 with \(f\): 1, 9, 26, 59, 72, 52, 29, 7, 1; cumulative 1, 10, 36, 95, 167, 219, 248, 255, 256.
Classes 0–15 … 135–150 with frequencies 1, 4, 17, 28, 25, 18, 13, 6, 5, 3 (\(N=120\)); cumulative 1, 5, 22, 50, 75, 93, 106, 112, 117, 120. Using \(Q_i = l + \dfrac{\frac{iN}{4}-C}{f}h\), \(D_i = l + \dfrac{\frac{iN}{10}-C}{f}h\), \(P_i = l + \dfrac{\frac{iN}{100}-C}{f}h\):