An attribute is a qualitative characteristic of an individual that cannot be measured quantitatively but can only be classified (e.g., honesty, gender, eye colour, smoking habit).
If a population is classified by attribute A, an individual either possesses A (denoted \(A\)) or does not possess A (denoted \(\alpha\), the "negation"). Similarly we use \(B, \beta\), \(C, \gamma\), etc.
Dividing a population into exactly two classes, those with the attribute and those without, is called dichotomy. The two classes are mutually exclusive (no one is in both) and exhaustive (everyone is in one). When an attribute is split into more than two classes, for example eye colour into blue, grey and brown, the classification is manifold. This unit works with dichotomies; manifold tables come back in section 7, with the coefficient of contingency.
Let \(A\) = drinking and \(B\) = singing, so \(\alpha\) = non-drinking and \(\beta\) = non-singing. Read each letter as a property the class must have:
Capital letters are called positive attributes and Greek letters negative attributes.
Class frequencies are classified by order = number of attributes specified.
With \(n\) attributes, the number of class frequencies of order \(r\) is \(\binom{n}{r}2^r\), and the total number of class frequencies of all orders is \(3^n\).
Why. A class of order \(r\) is fixed by two choices.
Multiplying gives \(\binom{n}{r}2^r\). Adding over \(r = 0, 1, \dots, n\) and using the binomial theorem \(\sum_r \binom{n}{r} a^r b^{\,n-r} = (a+b)^n\) with \(a = 2\), \(b = 1\):
\[ \sum_{r=0}^{n}\binom{n}{r}2^r = (2+1)^n = 3^n . \]Another way to see \(3^n\): each attribute is either present, absent, or not mentioned, which is three options per attribute.
| Order \(r\) | Count | Class frequencies |
|---|---|---|
| 0 | 1 | \(N\) |
| 1 | 6 | \((A), (\alpha)\), \((B), (\beta)\), \((C), (\gamma)\) |
| 2 | 12 | \((AB), (A\beta)\), \((\alpha B), (\alpha\beta)\), \((AC), (A\gamma)\), \((\alpha C), (\alpha\gamma)\), \((BC), (B\gamma)\), \((\beta C), (\beta\gamma)\) |
| 3 | 8 | \((ABC), (AB\gamma)\), \((A\beta C), (\alpha BC)\), \((A\beta\gamma), (\alpha B\gamma)\), \((\alpha\beta C), (\alpha\beta\gamma)\) |
| Total | 27 | \(= 3^3\) |
A class frequency with no Greek letter, such as \((A)\), \((AB)\) or \((ABC)\), is a positive class frequency; \(N\) counts as one too. Of order \(r\) there are only \(\binom{n}{r}\) of them (choose the attributes; each is present), so there are \(\sum_r \binom{n}{r} = 2^n\) in all. For three attributes these are the 8 frequencies
\[ N,\; (A),\; (B),\; (C),\; (AB),\; (AC),\; (BC),\; (ABC). \]They matter because every other class frequency can be computed from them (section 3), so a table of \(n\) attributes is fully known from its \(2^n\) positive class frequencies.
The frequencies of the highest-order classes (where every attribute is specified as either present or absent) are called ultimate class frequencies. For 2 attributes there are \(2^2 = 4\) ultimate frequencies; for 3 attributes there are \(2^3 = 8\).
Lower-order frequencies can be expressed as sums of ultimate frequencies. E.g.
Every class splits into two smaller classes by bringing in one more attribute: the members who have it and the members who do not. So any class frequency is the sum of two frequencies of the next order:
\[ N = (A) + (\alpha), \qquad (A) = (AB) + (A\beta), \qquad (AB) = (ABC) + (AB\gamma). \]Repeating the split until every attribute is specified writes any class frequency as a sum of ultimate frequencies. For three attributes, for example,
\[ (C) = (AC) + (\alpha C) = (ABC) + (A\beta C) + (\alpha BC) + (\alpha\beta C), \]and \(N\) is the sum of all 8 ultimate frequencies. For this reason the ultimate frequencies are called the fundamental set: knowing them, we know every class frequency.
The positive class frequencies are also a complete set: every frequency with Greek letters can be written in terms of them. A symbolic shorthand does this quickly.
Now replace each Greek letter, multiply out, and read each term back as a frequency.
The last line comes from expanding \((1-A)(1-B)(1-C)\): the terms of one letter enter with a minus sign, the terms of two letters with a plus sign, and the term of three letters with a minus sign. It is the inclusion–exclusion principle written for frequencies.
For any \(n\) attributes \(A_1, A_2, \dots, A_n\),
\[ (A_1 A_2 \cdots A_n) \;\ge\; (A_1) + (A_2) + \cdots + (A_n) - (n-1)N . \]Proof (by induction on \(n\)).
Use. In percentages (\(N = 100\)) with three attributes it reads \((ABC) \ge (A) + (B) + (C) - 200\). If 65%, 90% and 60% of candidates pass three papers, at least \(65 + 90 + 60 - 200 = 15\)% pass all three.
A set of class frequencies is consistent if it does not violate the basic axiom that every frequency must be non-negative.
The four ultimate frequencies must all be \(\ge 0\). Since
the consistency conditions reduce to:
For one attribute the only condition is \(0 \le (A) \le N\). For three attributes the 8 ultimate frequencies must all be \(\ge 0\). Write each one through the positive class frequencies (the dichotomy algebra of section 3) and rearrange it as a bound on \((ABC)\):
| Ultimate frequency (must be \(\ge 0\)) | Condition on \((ABC)\) | |
|---|---|---|
| \((ABC)\) | \((ABC) \ge 0\) | (i) |
| \((AB\gamma) = (AB) - (ABC)\) | \((ABC) \le (AB)\) | (ii) |
| \((A\beta C) = (AC) - (ABC)\) | \((ABC) \le (AC)\) | (iii) |
| \((\alpha BC) = (BC) - (ABC)\) | \((ABC) \le (BC)\) | (iv) |
| \((A\beta\gamma) = (A) - (AB) - (AC) + (ABC)\) | \((ABC) \ge (AB) + (AC) - (A)\) | (v) |
| \((\alpha B\gamma) = (B) - (AB) - (BC) + (ABC)\) | \((ABC) \ge (AB) + (BC) - (B)\) | (vi) |
| \((\alpha\beta C) = (C) - (AC) - (BC) + (ABC)\) | \((ABC) \ge (AC) + (BC) - (C)\) | (vii) |
| \((\alpha\beta\gamma)\) | \((ABC) \le (AB) + (AC) + (BC)\) \(- (A) - (B) - (C) + N\) | (viii) |
The data are consistent exactly when all eight hold. Conditions (i), (v), (vi) and (vii) are lower bounds on \((ABC)\); (ii), (iii), (iv) and (viii) are upper bounds.
Often only the frequencies up to order two are given. They are consistent when some value of \((ABC)\) satisfies (i)–(viii), that is, when every lower bound is at most every upper bound. Taking the 16 pairs one at a time, twelve of them give back the two-attribute conditions of 4.1 for the pairs \((A,B)\), \((A,C)\) and \((B,C)\). The other four are new:
From (i) and (viii):
\[ (AB) + (AC) + (BC) \ge (A) + (B) + (C) - N . \]From (iv) and (v):
\[ (AB) + (AC) - (BC) \le (A) . \]From (iii) and (vi):
\[ (AB) + (BC) - (AC) \le (B) . \]From (ii) and (vii):
\[ (AC) + (BC) - (AB) \le (C) . \]For example, (iv) and (v) together say \((AB) + (AC) - (A) \le (ABC) \le (BC)\), which is possible only if \((AB) + (AC) - (A) \le (BC)\), the second condition. So the second-order data of three attributes are consistent if and only if each pair passes the two-attribute test and these four inequalities hold. (This was also confirmed by checking every table with \(N = 6\) by computer.)
Given: \(N = 1000,\; (A) = 600,\; (B) = 500,\; (AB) = 200\). Check consistency.
\((A\beta) = 600 - 200 = 400 \ge 0\) ✓.
\((\alpha B) = 500 - 200 = 300 \ge 0\) ✓.
\((\alpha\beta) = 1000 - 600 - 500 + 200 = 100 \ge 0\) ✓.
Data are consistent.
\(N = 100,\; (A) = 70,\; (B) = 60,\; (AB) = 20\). Check.
\((\alpha\beta) = 100 - 70 - 60 + 20 = -10 < 0\) ✗ — inconsistent.
Two attributes \(A\) and \(B\) are independent if the proportion of \(A\) in the population is the same as the proportion of \(A\) within \(B\):
\[ \dfrac{(AB)}{(B)} \;=\; \dfrac{(A)}{N}. \]Equivalently:
\[ (AB) \;=\; \dfrac{(A)(B)}{N}. \]Among 1000 people: 200 have attribute \(A\), 300 have \(B\), and 60 have both.
Expected \((AB)\) under independence = \(200 \cdot 300 / 1000 = 60\). Observed = 60.
\(A\) and \(B\) are independent.
Same population, but \((AB) = 80\). Expected = 60, so \(80 > 60\) ⇒ positive association.
The definition says that \(B\) makes no difference to how common \(A\) is. That idea can be written in three ways, and all three are the same condition.
For two attributes the following are equivalent:
Proof that 1 gives 2. If two fractions are equal, say \(a/b = c/d = k\), then \(a = kb\) and \(c = kd\), so \(a + c = k(b + d)\) and \((a+c)/(b+d) = k\) as well. Apply this with \(a = (AB)\), \(b = (B)\), \(c = (A\beta)\), \(d = (\beta)\):
\[ \frac{(AB)}{(B)} = \frac{(AB) + (A\beta)}{(B) + (\beta)} = \frac{(A)}{N}, \]and multiplying by \((B)\) gives 2.
Proof that 2 gives 1. Using \((A\beta) = (A) - (AB)\) and \((\beta) = N - (B)\), then substituting 2:
\[ \frac{(A\beta)}{(\beta)} = \frac{(A) - (A)(B)/N}{N - (B)} = \frac{(A)\,[N - (B)]/N}{N - (B)} = \frac{(A)}{N} = \frac{(AB)}{(B)} . \]Proof that 2 and 3 are the same. Expand the cross-product difference using \((A\beta) = (A) - (AB)\), \((\alpha B) = (B) - (AB)\) and \((\alpha\beta) = N - (A) - (B) + (AB)\):
\[ (AB)(\alpha\beta) = N(AB) - (A)(AB) - (B)(AB) + (AB)^2, \] \[ (A\beta)(\alpha B) = (A)(B) - (A)(AB) - (B)(AB) + (AB)^2 . \]Subtracting, everything cancels except two terms:
So the cross products are equal exactly when \(\delta = 0\), that is, when 2 holds. \(\blacksquare\)
\(\delta\) is how far the observed \((AB)\) is above the frequency expected under independence. The other three cells are off by the same amount. For instance, using \((\beta) = N - (B)\),
\[ (A\beta) - \frac{(A)(\beta)}{N} = (A) - (AB) - (A) + \frac{(A)(B)}{N} = -\delta , \]and in the same way \((\alpha B) - (\alpha)(B)/N = -\delta\) and \((\alpha\beta) - (\alpha)(\beta)/N = +\delta\). Two consequences follow.
The tests of independence turn into tests of association by replacing “=” with “>” or “<”. \(A\) and \(B\) are positively associated when
\[ \frac{(AB)}{(B)} > \frac{(A\beta)}{(\beta)}, \quad \text{or equivalently} \quad (AB) > \frac{(A)(B)}{N}, \quad \text{i.e. } \delta > 0, \]and negatively associated when the inequalities are reversed. The three forms agree because
\[ \frac{(AB)}{(B)} - \frac{(A\beta)}{(\beta)} = \frac{(AB)(\beta) - (A\beta)(B)}{(B)(\beta)} = \frac{N\delta}{(B)(\beta)} , \]where the numerator was simplified exactly as in the key identity, and \((B)(\beta) > 0\).
Yule's coefficient of association:
\[ Q \;=\; \dfrac{(AB)(\alpha\beta) - (A\beta)(\alpha B)}{(AB)(\alpha\beta) + (A\beta)(\alpha B)}. \]By the key identity of section 5, the numerator \(a - b\) equals \(N\delta\). So \(Q\) has the sign of \(\delta\): positive for positive association, negative for negative association, and zero exactly when the attributes are independent.
Because a single empty cell already gives \(Q = \pm 1\), \(Q\) reports “complete” association in tables where the two attributes are far from going together in every case. For example, \((AB) = 10\), \((A\beta) = 0\), \((\alpha B) = 90\), \((\alpha\beta) = 900\) gives \(Q = 1\), although only 10 of the 100 B's are A's.
Yule's coefficient of colligation (many books write it \(Y\)):
\[ \omega \;=\; \dfrac{\sqrt{(AB)(\alpha\beta)} - \sqrt{(A\beta)(\alpha B)}}{\sqrt{(AB)(\alpha\beta)} + \sqrt{(A\beta)(\alpha B)}}. \]Dividing the numerator and denominator by \(\sqrt{(AB)(\alpha\beta)}\) gives the equivalent form
\[ \omega = \frac{1 - \sqrt{K}}{1 + \sqrt{K}}, \qquad K = \frac{(A\beta)(\alpha B)}{(AB)(\alpha\beta)} . \]Write \(p = \sqrt{(AB)(\alpha\beta)}\) and \(q = \sqrt{(A\beta)(\alpha B)}\). Both are \(\ge 0\) and \(\omega = (p - q)/(p + q)\), which has the same form as \(Q = (a-b)/(a+b)\). The argument used for \(Q\) therefore gives \(-1 \le \omega \le 1\) directly, with
Also \(|\omega| \le |Q|\): since \(\omega^2 \le 1\), the denominator \(1 + \omega^2\) is at most 2, so \(|Q| = 2|\omega|/(1+\omega^2) \ge |\omega|\). Fig 5.2 shows the curve.
\(Q\) and \(\omega\) need a \(2\times2\) table. For a manifold classification into an \(r \times s\) table, association is measured through the \(\chi^2\) statistic, \(\chi^2 = \sum (O - E)^2/E\), where each expected frequency \(E\) is (row total \(\times\) column total)\(/N\), as under independence.
Mean square contingency:
\[ \phi^2 = \frac{\chi^2}{N} . \]Karl Pearson's coefficient of contingency:
\[ C = \sqrt{\frac{\chi^2}{\chi^2 + N}} = \sqrt{\frac{\phi^2}{1 + \phi^2}} . \]Tschuprow's coefficient:
\[ T = \sqrt{\frac{\phi^2}{\sqrt{(r-1)(s-1)}}} . \]Take the table of Worked Problem 7: \((AB) = 35\), \((A\beta) = 25\), \((\alpha B) = 15\), \((\alpha\beta) = 25\), with \((A) = 60\), \((\alpha) = 40\), \((B) = (\beta) = 50\), \(N = 100\).
Expected frequencies. \(60 \times 50/100 = 30\) for \((AB)\) and \((A\beta)\); \(40 \times 50/100 = 20\) for \((\alpha B)\) and \((\alpha\beta)\). Each cell is off by \(\delta = 5\).
\[ \chi^2 = \frac{5^2}{30} + \frac{5^2}{30} + \frac{5^2}{20} + \frac{5^2}{20} = 0.8333 + 0.8333 + 1.25 + 1.25 = 4.1667 . \]Then \(\phi^2 = 4.1667/100 = 0.041667\), \(\phi = T = 0.204\), and
\[ C = \sqrt{\frac{4.1667}{104.1667}} = \sqrt{0.04} = 0.2 . \]All three say the association is weak, as \(Q = 0.4\) did.
Among 200 people: \((AB) = 80,\; (A\beta) = 30,\; (\alpha B) = 40,\; (\alpha\beta) = 50\).
\(Q = (80 \cdot 50 - 30 \cdot 40)/(80 \cdot 50 + 30 \cdot 40)\) \(= (4000 - 1200)/(4000 + 1200)\) \(= 2800/5200 = 0.538\).
A moderate positive association: \(Q\) is a little over one half.
\(\omega = (\sqrt{4000} - \sqrt{1200})/(\sqrt{4000} + \sqrt{1200})\) \(= (63.25 - 34.64)/(63.25 + 34.64)\) \(= 28.61/97.89 = 0.292\).
Verify: \(2\omega/(1+\omega^2) = 0.584/(1.0853) ≈ 0.538\) ✓.
Among 1000 individuals: \((AB) = 60, (A\beta) = 140, (\alpha B) = 240, (\alpha\beta) = 560\).
Check: \((A) = 60+140 = 200,\; (B) = 60+240 = 300\); expected \((AB) = 200 \cdot 300 / 1000 = 60\) — matches observed.
\(Q = (60 \cdot 560 - 140 \cdot 240)/(60 \cdot 560 + 140 \cdot 240)\) \(= (33600 - 33600)/(33600 + 33600) = 0\). Attributes are independent.
Eight problems in the textbook's order, grouped by topic, then its eight exercises with the answers checked. Every figure was recomputed exactly. Where the textbook prints a different figure or conclusion, the correct one is used and the difference is noted.
Given. \(N = 23713\), \((A) = 1618\), \((B) = 2015\), \((C) = 770\), \((AB) = 587\), \((AC) = 335\), \((BC) = 428\), \((ABC) = 156\). Find the remaining class frequencies.
Plan. Three attributes have \(3^3 = 27\) class frequencies; 8 are given, so 19 remain. Work upwards in order: first order, then second, then third, each time subtracting from a frequency already known (section 3).
Step 1: first order. \((\alpha) = N - (A)\), and likewise for \(\beta\), \(\gamma\):
\((\alpha) = 23713 - 1618 = 22095\), \((\beta) = 23713 - 2015 = 21698\), \((\gamma) = 23713 - 770 = 22943\).
Step 2: second order. Each pair of attributes gives a \(2\times2\) table, completed from its margins.
| Pair | Frequency | Working | Value |
|---|---|---|---|
| A, B | \((A\beta)\) | \((A) - (AB) = 1618 - 587\) | 1031 |
| \((\alpha B)\) | \((B) - (AB) = 2015 - 587\) | 1428 | |
| \((\alpha\beta)\) | \((\alpha) - (\alpha B) = 22095 - 1428\) | 20667 | |
| A, C | \((A\gamma)\) | \((A) - (AC) = 1618 - 335\) | 1283 |
| \((\alpha C)\) | \((C) - (AC) = 770 - 335\) | 435 | |
| \((\alpha\gamma)\) | \((\alpha) - (\alpha C) = 22095 - 435\) | 21660 | |
| B, C | \((B\gamma)\) | \((B) - (BC) = 2015 - 428\) | 1587 |
| \((\beta C)\) | \((C) - (BC) = 770 - 428\) | 342 | |
| \((\beta\gamma)\) | \((\beta) - (\beta C) = 21698 - 342\) | 21356 |
Step 3: third order (the ultimate frequencies).
| Frequency | Working | Value |
|---|---|---|
| \((AB\gamma)\) | \((AB) - (ABC) = 587 - 156\) | 431 |
| \((A\beta C)\) | \((AC) - (ABC) = 335 - 156\) | 179 |
| \((\alpha BC)\) | \((BC) - (ABC) = 428 - 156\) | 272 |
| \((A\beta\gamma)\) | \((A\beta) - (A\beta C) = 1031 - 179\) | 852 |
| \((\alpha B\gamma)\) | \((\alpha B) - (\alpha BC) = 1428 - 272\) | 1156 |
| \((\alpha\beta C)\) | \((\beta C) - (A\beta C) = 342 - 179\) | 163 |
| \((\alpha\beta\gamma)\) | \((\alpha\beta) - (\alpha\beta C) = 20667 - 163\) | 20504 |
Check. The 8 ultimate frequencies must add to \(N\): \(156 + 431 + 179 + 272 + 852 + 1156 + 163 + 20504 = 23713\). ✓ As a second check, the formula of section 3 gives \((\alpha\beta\gamma) = 23713 - 1618 - 2015 - 770 + 587 + 335 + 428 - 156 = 20504\).
Note. The textbook prints \((A\beta) = 1037\) but then uses 1031, which is correct (\(1618 - 587 = 1031\)). It also writes \((\alpha\beta C) = (\beta C) - (\alpha\beta C)\); the subtracted term should be \((A\beta C)\).
Given. \(N = 500\), \((A) = 400\), \((B) = 380\), \((AB) = 270\). Are the data consistent?
Method. Compute the four ultimate frequencies; all must be \(\ge 0\).
\[ (A\beta) = 400 - 270 = 130, \qquad (\alpha B) = 380 - 270 = 110, \] \[ (\alpha\beta) = N - (A) - (B) + (AB) = 500 - 400 - 380 + 270 = -10 . \]Conclusion. \((\alpha\beta) = -10 < 0\), which is impossible for a count, so the data are inconsistent. Equivalently, the condition \((AB) \ge (A) + (B) - N = 280\) fails, since \(270 < 280\): if 400 of 500 have A and 380 have B, at least 280 must have both.
Given. \((A)/N = x\), \((B)/N = 2x\), \((C)/N = 3x\) and \((AB)/N = (AC)/N = (BC)/N = y\). Show that for consistent data neither \(x\) nor \(y\) can exceed \(\tfrac14\).
The bound is reached. At \(x = \tfrac14\), (1) and (2) force \(y = \tfrac14\). Taking also \((ABC)/N = \tfrac14\), the ultimate frequencies (as fractions of \(N\)) are \(\tfrac14\) for \((ABC)\) and \((\alpha B\gamma)\), \(\tfrac12\) for \((\alpha\beta C)\), and 0 for the other five: all \(\ge 0\), so \(x = y = \tfrac14\) is consistent.
Given. \((A) = (B) = (C) = \tfrac12 N\); 80% of the A's are B's and 75% of the A's are C's. Find the limits to the percentage of B's that are C's.
Step 1: translate. \((AB) = 0.8(A) = 0.4N\) and \((AC) = 0.75(A) = 0.375N\). The quantity asked for is \(100(BC)/(B) = 100 \times 2(BC)/N\) per cent.
Step 2: the four three-attribute conditions of section 4 (the two-attribute conditions only say it lies between 0% and 100%).
Answer. Between 55% and 95% of the B's are C's. Both ends are attainable: a search over all values of \((BC)\) and \((ABC)\) finds consistent tables exactly on this interval.
Note. The textbook writes “\((AC)\) = 75% of \((C)\)”; the problem says 75% of the A's. The value is the same here only because \((A) = (C)\).
Given. \(N = 1000\), \((A) = 100\), \((B) = 300\), and A and B are independent. Find \((AB)\) and \((\alpha\beta)\).
Step 1. By the independence criterion, \((AB) = \dfrac{(A)(B)}{N} = \dfrac{100 \times 300}{1000} = 30\).
Step 2. If A and B are independent, so are \(\alpha\) and \(\beta\) (section 5). With \((\alpha) = 900\) and \((\beta) = 700\),
\[ (\alpha\beta) = \frac{(\alpha)(\beta)}{N} = \frac{900 \times 700}{1000} = 630 . \]Check. \((\alpha\beta) = N - (A) - (B) + (AB) = 1000 - 100 - 300 + 30 = 630\). ✓
Note. The textbook's middle step reads \(100 \times 300/1000\); it should be \(900 \times 700/1000\). Its answer, 630, is right.
Given. \(N = 1000\), \((A) = 450\), \((B) = 650\), \((AB) = 310\). Are A and B associated?
Step 1. Under independence we would expect \((A)(B)/N = 450 \times 650/1000 = 292.5\).
Step 2. Observed \((AB) = 310 > 292.5\), so \(\delta = 17.5 > 0\): A and B are positively associated.
How strongly? The table is \((AB) = 310\), \((A\beta) = 140\), \((\alpha B) = 340\), \((\alpha\beta) = 210\), so
\[ Q = \frac{310 \times 210 - 140 \times 340}{310 \times 210 + 140 \times 340} = \frac{65100 - 47600}{65100 + 47600} = \frac{17500}{112700} = 0.155 , \]a weak positive association. The numerator is \(N\delta = 1000 \times 17.5\), as the key identity says.
Given. Of 100 people, 60 take a morning walk (A), 50 are physically fit (B), and 35 do both. Find Yule's coefficient of association and interpret it.
Step 1: complete the table.
| \(B\) (fit) | \(\beta\) (not fit) | Total | |
|---|---|---|---|
| \(A\) (walk) | 35 | 25 | 60 |
| \(\alpha\) (no walk) | 15 | 25 | 40 |
| Total | 50 | 50 | 100 |
Step 2: the cross products. \((AB)(\alpha\beta) = 35 \times 25 = 875\) and \((A\beta)(\alpha B) = 25 \times 15 = 375\).
\[ Q = \frac{875 - 375}{875 + 375} = \frac{500}{1250} = 0.4 . \]Step 3: colligation, as a check. \(K = 375/875 = 0.4286\), \(\sqrt K = 0.6547\), so
\[ \omega = \frac{1 - 0.6547}{1 + 0.6547} = \frac{0.3453}{1.6547} = 0.209, \qquad \frac{2\omega}{1+\omega^2} = \frac{0.4174}{1.0436} = 0.4 = Q . \;\checkmark \]Interpretation. \(Q = 0.4\) is a fairly low degree of positive association: walkers are fit more often (35 of 60, 58%) than non-walkers (15 of 40, 37.5%), but far from always. The expected \((AB)\) under independence is \(60 \times 50/100 = 30\), below the observed 35.
Given. Of 200 students, 150 are boys; 120 boys and 40 girls passed. Let A = boy and B = passed. Is there any association?
| \(A\) (boys) | \(\alpha\) (girls) | Total | |
|---|---|---|---|
| \(B\) (passed) | 120 | 40 | 160 |
| \(\beta\) (failed) | 30 | 10 | 40 |
| Total | 150 | 50 | 200 |
Conclusion. \(Q = 0\): sex and result are independent. Directly, 120 of 150 boys (80%) and 40 of 50 girls (80%) passed.
Note. The textbook's table gives the passed total as 100 (it is \(120 + 40 = 160\)), and its formula has \((A\beta)(\alpha\beta)\) in the denominator in place of \((AB)(\alpha\beta)\). Its answer, \(Q = 0\), is right.