An observed test score \(X\) is the sum of a true score \(T\) (the examinee's actual ability on whatever the test measures) plus a random measurement error \(E\):
\[ X \;=\; T + E, \]with two key assumptions:
where \(\sigma_X^2\) is the observed-score variance, \(\sigma_T^2\) the true-score variance, and \(\sigma_E^2\) the error variance.
The reliability coefficient \(r_{xx}\) is the proportion of observed-score variance that is "true" (consistent) variance:
\[ r_{xx} \;=\; \dfrac{\sigma_T^2}{\sigma_X^2} \;=\; 1 - \dfrac{\sigma_E^2}{\sigma_X^2}. \]Range: \(0 \le r_{xx} \le 1\). Reliability = 1 means perfect consistency (no measurement error); 0 means scores are pure noise.
The standard error of measurement (SEM) is the SD of the error term \(E\):
\[ \text{SEM} \;=\; \sigma_E \;=\; \sigma_X \sqrt{1 - r_{xx}}. \]Interpretation: if a person's true score is \(T\), their observed score on a repeated administration is approximately \(N(T, \sigma_E^2)\). A 95 % confidence interval around the observed score is \(X \pm 1.96\, \text{SEM}\).
A test has SD = 12 and reliability 0.84. SEM = \(12 \sqrt{1 - 0.84} = 12 \sqrt{0.16} = 12 \times 0.4 = 4.8\).
An individual scoring 75 has 95 % CI: 75 ± 1.96 × 4.8 = 75 ± 9.4 ⇒ (65.6, 84.4) — their true score is likely in this range.
Two tests have the same SEM but different reliabilities — the one with higher SD has higher reliability for the same SEM, because reliability is a ratio.
The index of reliability is the correlation between observed score and true score:
\[ \rho_{XT} \;=\; \sqrt{r_{xx}}. \]Range: 0 to 1. It is the upper bound on how well a test can correlate with anything else.
Because \(\sigma_T = \sigma_X \sqrt{r_{xx}}\) and \(\text{Cov}(X, T) = \text{Var}(T) = \sigma_T^2\):
\(\rho_{XT} = \text{Cov}(X, T)/(\sigma_X \sigma_T) = \sigma_T^2/(\sigma_X \sigma_T) = \sigma_T/\sigma_X = \sqrt{r_{xx}}\).
The validity coefficient (Section 9) of a test cannot exceed its index of reliability. A test cannot correlate higher with any external criterion than it does with its own true score.
If \(r_{xx} = 0.81\), then index of reliability = 0.9. The maximum possible correlation of this test with any other measure (e.g., academic GPA) is 0.9.
Two tests are parallel if they have:
Under parallel-test assumptions, the reliability of either test is the correlation between scores on the two tests:
This is the basis of nearly every empirical reliability estimation method.
When two equivalent forms (Form A and Form B) of a test are constructed, the correlation between scores on the two forms estimates reliability. Requires identical difficulty, content coverage, and item statistics.
Administer the same test twice to the same group, with a time gap (1–4 weeks typical), then correlate the two sets of scores:
\[ r_{tt} \;=\; \text{Corr}(X_1, X_2). \]This is the coefficient of stability.
Pros: simple, intuitive.
Cons:
Administer two parallel forms either simultaneously (coefficient of equivalence) or with a gap (coefficient of stability & equivalence). Correlation of forms gives reliability.
Administer the test once; split into two halves (e.g., odd vs even items); correlate the half-scores. Since each half is shorter than the full test, apply the Spearman-Brown correction:
where \(r_{hh}\) is the correlation between the two halves. More generally, lengthening a test by a factor of \(k\) gives:
\[ r_{kk} \;=\; \dfrac{k\, r_{xx}}{1 + (k - 1) r_{xx}}. \]where \(\sigma_d^2\) is the variance of the differences between the two halves, and \(\sigma_X^2\) is the variance of the total scores.
Rulon's method does not require computing the correlation between the halves — only the variance of differences and the total variance. Convenient when correlation calculation is awkward.
Half-test correlation \(r_{hh} = 0.7\). Full-test reliability = \(2(0.7)/(1 + 0.7) = 1.4/1.7 = 0.82\).
For 50 students taking a 40-item test split into odd/even halves: SD of differences = 4.0; SD of total = 12. \(r_{xx} = 1 - 16/144 = 1 - 0.111 = 0.889\) — high reliability.
For tests with dichotomous items (right/wrong, 1/0), Kuder and Richardson (1937) derived formulas that don't require splitting the test.
where \(k\) is the number of items, \(p_i\) is the proportion answering item \(i\) correctly, \(q_i = 1 - p_i\), and \(\sigma_X^2\) is the variance of the total test scores.
Assumes all items have the same difficulty. Let \(\bar p\) be the average proportion correct:
KR-21 ≤ KR-20 (lower bound when item difficulties vary).
where \(\sigma_i^2\) is the variance of item \(i\). For dichotomous items, \(\sigma_i^2 = p_i q_i\) so α = KR-20.
A 10-item test (k = 10) with item p-values 0.5, 0.6, 0.7, 0.4, 0.5, 0.8, 0.3, 0.6, 0.5, 0.7. \(\sum p q = 0.25 + 0.24 + 0.21 + 0.24 + 0.25 + 0.16 + 0.21 + 0.24 + 0.25 + 0.21 = 2.26\). Total-score variance \(\sigma_X^2 = 6.0\).
\(r = (10/9)(1 - 2.26/6.0) = 1.111 \times 0.623 = 0.693\).
For the same test, average \(p = (0.5+0.6+\ldots+0.7)/10 = 0.56\); \(\bar q = 0.44\); \(\bar p \bar q = 0.2464\); \(k \bar p \bar q = 2.464\).
\(r_{KR\text{-}21} = (10/9)(1 - 2.464/6.0) = 1.111 \times 0.589 = 0.654\). Slightly lower than KR-20 (as expected).
Reliability rises with test length (more items = better averaging out of errors). The Spearman-Brown formula (Section 6.3) quantifies the relationship.
A 20-item test has reliability 0.6. How many items needed for reliability 0.9?
From \(r_{kk} = k r_{xx} / (1 + (k-1) r_{xx})\): solve for \(k\) given \(r_{xx} = 0.6\) and \(r_{kk} = 0.9\):
\(k = r_{kk}(1 - r_{xx})/[r_{xx}(1 - r_{kk})] = 0.9 \times 0.4/(0.6 \times 0.1) = 0.36/0.06 = 6\). So multiply length by 6 → need 120 items.
Validity is the extent to which a test actually measures what it claims to measure. A test of "mathematical ability" should correlate with grades in math classes — not with running speed.
Numerically: \(r_{xy}\) = correlation between test score \(X\) and criterion \(Y\). Validity is typically 0.3–0.6 in practice.
Administer the test to a sample, also obtain criterion scores, compute Pearson correlation.
Lengthening a test increases reliability and (typically) increases validity, but with diminishing returns. Eventually adding more items provides minimal benefit.
The correction for attenuation:
\[ r_{tx} \;=\; \dfrac{r_{xy}}{\sqrt{r_{xx} r_{yy}}}. \]This is the correlation if both test and criterion had perfect reliability — an upper bound on the true relationship.
An aptitude test correlates 0.5 with first-year college GPA. Test reliability = 0.9; GPA reliability = 0.8. Corrected validity = 0.5/\(\sqrt{0.9 \times 0.8}\) = 0.5/0.849 = 0.59.
The validity coefficient cannot exceed \(\sqrt{r_{xx}}\). If \(r_{xx} = 0.64\), maximum validity = 0.8.
| Aspect | Reliability | Validity |
|---|---|---|
| Definition | Consistency of measurement | Accuracy of measurement (measures the right thing) |
| Numerical range | 0 to 1 | −1 to 1 |
| What's required | Two administrations / split | An external criterion |
| Symbol | \(r_{xx}\) | \(r_{xy}\) |
| Necessary condition for validity? | Yes (reliability is upper bound on validity) | No (a test can be reliable without being valid) |
Analogy: reliability is a precise scale (always reads the same weight); validity is an accurate scale (reads your true weight). A consistently mis-calibrated scale is reliable but not valid.
The Intelligence Quotient (IQ) is a numerical measure of intelligence relative to a normed reference group.
The original Stern (1912) formula:
Mental Age was the age at which an average child solves the same items the examinee does. This original formula is rarely used today (it doesn't work for adults).
Most modern tests (Wechsler, Stanford-Binet) use a deviation IQ — a standard score with mean 100 and SD 15 (Wechsler) or 16 (Stanford-Binet) in the norming population:
| IQ Range | Classification | Approx. % of population |
|---|---|---|
| ≥ 130 | Very superior | 2.2 % |
| 120–129 | Superior | 6.7 % |
| 110–119 | High average | 16.1 % |
| 90–109 | Average | 50.0 % |
| 80–89 | Low average | 16.1 % |
| 70–79 | Borderline | 6.7 % |
| < 70 | Intellectual disability | 2.2 % |
A child of CA = 8 yr performs at Mental Age 10 on a Binet test. IQ = 10/8 × 100 = 125.
An adult scores 1.5 SDs above the mean. IQ = 100 + 15 × 1.5 = 122.5. This places them at approximately the 93rd percentile.