Skip to the content

Topics Covered

Reliability Concept Error Variance Index of Reliability Parallel Tests Test-Retest Rulon Method Kuder-Richardson Validity IQ
On this page
  1. 1. Classical Test Theory — The True-Score Model
  2. 2. Reliability of Test Scores
  3. 3. Error Variance and Standard Error of Measurement
  4. 4. Index of Reliability
  5. 5. Parallel Tests
  6. 6. Methods of Determining Test Reliability
  7. 7. Reliability and Test Length
  8. 8. Validity of Test Scores
  9. 9. Comparison — Reliability vs Validity
  10. 10. Intelligence Quotient (IQ)
  11. Key Take-aways

1. Classical Test Theory — The True-Score Model

FOUNDATIONAL ASSUMPTION

An observed test score \(X\) is the sum of a true score \(T\) (the examinee's actual ability on whatever the test measures) plus a random measurement error \(E\):

\[ X \;=\; T + E, \]

with two key assumptions:

Variance Decomposition

\[ \sigma_X^2 \;=\; \sigma_T^2 + \sigma_E^2, \]

where \(\sigma_X^2\) is the observed-score variance, \(\sigma_T^2\) the true-score variance, and \(\sigma_E^2\) the error variance.

2. Reliability of Test Scores

DEFINITION

The reliability coefficient \(r_{xx}\) is the proportion of observed-score variance that is "true" (consistent) variance:

\[ r_{xx} \;=\; \dfrac{\sigma_T^2}{\sigma_X^2} \;=\; 1 - \dfrac{\sigma_E^2}{\sigma_X^2}. \]

Range: \(0 \le r_{xx} \le 1\). Reliability = 1 means perfect consistency (no measurement error); 0 means scores are pure noise.

Rules of Thumb

3. Error Variance and Standard Error of Measurement

DEFINITION

The standard error of measurement (SEM) is the SD of the error term \(E\):

\[ \text{SEM} \;=\; \sigma_E \;=\; \sigma_X \sqrt{1 - r_{xx}}. \]

Interpretation: if a person's true score is \(T\), their observed score on a repeated administration is approximately \(N(T, \sigma_E^2)\). A 95 % confidence interval around the observed score is \(X \pm 1.96\, \text{SEM}\).

Why SEM Matters

EXAMPLE 1

A test has SD = 12 and reliability 0.84. SEM = \(12 \sqrt{1 - 0.84} = 12 \sqrt{0.16} = 12 \times 0.4 = 4.8\).

An individual scoring 75 has 95 % CI: 75 ± 1.96 × 4.8 = 75 ± 9.4 ⇒ (65.6, 84.4) — their true score is likely in this range.

EXAMPLE 2

Two tests have the same SEM but different reliabilities — the one with higher SD has higher reliability for the same SEM, because reliability is a ratio.

4. Index of Reliability

DEFINITION

The index of reliability is the correlation between observed score and true score:

\[ \rho_{XT} \;=\; \sqrt{r_{xx}}. \]

Range: 0 to 1. It is the upper bound on how well a test can correlate with anything else.

Why \(\sqrt{r_{xx}}\)?

Because \(\sigma_T = \sigma_X \sqrt{r_{xx}}\) and \(\text{Cov}(X, T) = \text{Var}(T) = \sigma_T^2\):

\(\rho_{XT} = \text{Cov}(X, T)/(\sigma_X \sigma_T) = \sigma_T^2/(\sigma_X \sigma_T) = \sigma_T/\sigma_X = \sqrt{r_{xx}}\).

Validity Ceiling

The validity coefficient (Section 9) of a test cannot exceed its index of reliability. A test cannot correlate higher with any external criterion than it does with its own true score.

EXAMPLE

If \(r_{xx} = 0.81\), then index of reliability = 0.9. The maximum possible correlation of this test with any other measure (e.g., academic GPA) is 0.9.

5. Parallel Tests

DEFINITION

Two tests are parallel if they have:

  1. The same true score for every examinee (\(T_1 = T_2 = T\)).
  2. The same error variance (\(\sigma_{E_1}^2 = \sigma_{E_2}^2\)).
  3. Uncorrelated errors (\(\text{Cov}(E_1, E_2) = 0\)).

Under parallel-test assumptions, the reliability of either test is the correlation between scores on the two tests:

\[ r_{xx} \;=\; \text{Corr}(X_1, X_2) \;=\; \rho_{X_1 X_2}. \]

This is the basis of nearly every empirical reliability estimation method.

Alternate-Form Reliability

When two equivalent forms (Form A and Form B) of a test are constructed, the correlation between scores on the two forms estimates reliability. Requires identical difficulty, content coverage, and item statistics.

6. Methods of Determining Test Reliability

6.1 Test–Retest Method

Administer the same test twice to the same group, with a time gap (1–4 weeks typical), then correlate the two sets of scores:

\[ r_{tt} \;=\; \text{Corr}(X_1, X_2). \]

This is the coefficient of stability.

Pros: simple, intuitive.
Cons:

6.2 Equivalent-Forms Method

Administer two parallel forms either simultaneously (coefficient of equivalence) or with a gap (coefficient of stability & equivalence). Correlation of forms gives reliability.

6.3 Split-Half Method

Administer the test once; split into two halves (e.g., odd vs even items); correlate the half-scores. Since each half is shorter than the full test, apply the Spearman-Brown correction:

\[ r_{xx} \;=\; \dfrac{2 r_{hh}}{1 + r_{hh}}, \]

where \(r_{hh}\) is the correlation between the two halves. More generally, lengthening a test by a factor of \(k\) gives:

\[ r_{kk} \;=\; \dfrac{k\, r_{xx}}{1 + (k - 1) r_{xx}}. \]

6.4 Rulon's Method (Split-Half Variant)

RULON FORMULA \[ r_{xx} \;=\; 1 - \dfrac{\sigma_d^2}{\sigma_X^2}, \]

where \(\sigma_d^2\) is the variance of the differences between the two halves, and \(\sigma_X^2\) is the variance of the total scores.

Rulon's method does not require computing the correlation between the halves — only the variance of differences and the total variance. Convenient when correlation calculation is awkward.

EXAMPLE 1 — Spearman-Brown

Half-test correlation \(r_{hh} = 0.7\). Full-test reliability = \(2(0.7)/(1 + 0.7) = 1.4/1.7 = 0.82\).

EXAMPLE 2 — Rulon

For 50 students taking a 40-item test split into odd/even halves: SD of differences = 4.0; SD of total = 12. \(r_{xx} = 1 - 16/144 = 1 - 0.111 = 0.889\) — high reliability.

6.5 Kuder-Richardson Formulas (Method of Rational Equivalence)

For tests with dichotomous items (right/wrong, 1/0), Kuder and Richardson (1937) derived formulas that don't require splitting the test.

KR-20 (Kuder-Richardson 20)

\[ r_{KR\text{-}20} \;=\; \dfrac{k}{k - 1}\left(1 - \dfrac{\sum_{i=1}^{k} p_i q_i}{\sigma_X^2}\right), \]

where \(k\) is the number of items, \(p_i\) is the proportion answering item \(i\) correctly, \(q_i = 1 - p_i\), and \(\sigma_X^2\) is the variance of the total test scores.

KR-21 (simpler, less accurate)

Assumes all items have the same difficulty. Let \(\bar p\) be the average proportion correct:

\[ r_{KR\text{-}21} \;=\; \dfrac{k}{k-1}\left(1 - \dfrac{k\, \bar p\, \bar q}{\sigma_X^2}\right). \]

KR-21 ≤ KR-20 (lower bound when item difficulties vary).

Cronbach's α (generalization for any score type)

\[ \alpha \;=\; \dfrac{k}{k - 1}\left(1 - \dfrac{\sum_{i=1}^{k} \sigma_i^2}{\sigma_X^2}\right), \]

where \(\sigma_i^2\) is the variance of item \(i\). For dichotomous items, \(\sigma_i^2 = p_i q_i\) so α = KR-20.

EXAMPLE 1 — KR-20

A 10-item test (k = 10) with item p-values 0.5, 0.6, 0.7, 0.4, 0.5, 0.8, 0.3, 0.6, 0.5, 0.7. \(\sum p q = 0.25 + 0.24 + 0.21 + 0.24 + 0.25 + 0.16 + 0.21 + 0.24 + 0.25 + 0.21 = 2.26\). Total-score variance \(\sigma_X^2 = 6.0\).

\(r = (10/9)(1 - 2.26/6.0) = 1.111 \times 0.623 = 0.693\).

EXAMPLE 2 — KR-21

For the same test, average \(p = (0.5+0.6+\ldots+0.7)/10 = 0.56\); \(\bar q = 0.44\); \(\bar p \bar q = 0.2464\); \(k \bar p \bar q = 2.464\).

\(r_{KR\text{-}21} = (10/9)(1 - 2.464/6.0) = 1.111 \times 0.589 = 0.654\). Slightly lower than KR-20 (as expected).

7. Reliability and Test Length

Reliability rises with test length (more items = better averaging out of errors). The Spearman-Brown formula (Section 6.3) quantifies the relationship.

EXAMPLE

A 20-item test has reliability 0.6. How many items needed for reliability 0.9?

From \(r_{kk} = k r_{xx} / (1 + (k-1) r_{xx})\): solve for \(k\) given \(r_{xx} = 0.6\) and \(r_{kk} = 0.9\):

\(k = r_{kk}(1 - r_{xx})/[r_{xx}(1 - r_{kk})] = 0.9 \times 0.4/(0.6 \times 0.1) = 0.36/0.06 = 6\). So multiply length by 6 → need 120 items.

8. Validity of Test Scores

DEFINITION

Validity is the extent to which a test actually measures what it claims to measure. A test of "mathematical ability" should correlate with grades in math classes — not with running speed.

Types of Validity

  1. Content validity: do the items cover the entire content domain? (Expert review.)
  2. Criterion-related validity: correlation with an external criterion.
    • Predictive: test correlates with later performance.
    • Concurrent: test correlates with current performance.
  3. Construct validity: does the test measure the theoretical construct? (Convergent & discriminant validity.)
  4. Face validity: does the test look like it measures what it should?

Validity Coefficient

Numerically: \(r_{xy}\) = correlation between test score \(X\) and criterion \(Y\). Validity is typically 0.3–0.6 in practice.

Calculating Validity

Administer the test to a sample, also obtain criterion scores, compute Pearson correlation.

Validity and Test Length

Lengthening a test increases reliability and (typically) increases validity, but with diminishing returns. Eventually adding more items provides minimal benefit.

Relation to Reliability

The correction for attenuation:

\[ r_{tx} \;=\; \dfrac{r_{xy}}{\sqrt{r_{xx} r_{yy}}}. \]

This is the correlation if both test and criterion had perfect reliability — an upper bound on the true relationship.

EXAMPLE 1

An aptitude test correlates 0.5 with first-year college GPA. Test reliability = 0.9; GPA reliability = 0.8. Corrected validity = 0.5/\(\sqrt{0.9 \times 0.8}\) = 0.5/0.849 = 0.59.

EXAMPLE 2 — Validity ceiling

The validity coefficient cannot exceed \(\sqrt{r_{xx}}\). If \(r_{xx} = 0.64\), maximum validity = 0.8.

9. Comparison — Reliability vs Validity

AspectReliabilityValidity
DefinitionConsistency of measurementAccuracy of measurement (measures the right thing)
Numerical range0 to 1−1 to 1
What's requiredTwo administrations / splitAn external criterion
Symbol\(r_{xx}\)\(r_{xy}\)
Necessary condition for validity?Yes (reliability is upper bound on validity)No (a test can be reliable without being valid)

Analogy: reliability is a precise scale (always reads the same weight); validity is an accurate scale (reads your true weight). A consistently mis-calibrated scale is reliable but not valid.

10. Intelligence Quotient (IQ)

DEFINITION

The Intelligence Quotient (IQ) is a numerical measure of intelligence relative to a normed reference group.

Classical (Stern–Binet) IQ

The original Stern (1912) formula:

\[ \text{IQ} \;=\; \dfrac{\text{Mental Age (MA)}}{\text{Chronological Age (CA)}} \times 100. \]

Mental Age was the age at which an average child solves the same items the examinee does. This original formula is rarely used today (it doesn't work for adults).

Modern (Deviation) IQ

Most modern tests (Wechsler, Stanford-Binet) use a deviation IQ — a standard score with mean 100 and SD 15 (Wechsler) or 16 (Stanford-Binet) in the norming population:

\[ \text{IQ} \;=\; 100 + 15 \cdot Z. \]

IQ Classification (Wechsler)

IQ RangeClassificationApprox. % of population
≥ 130Very superior2.2 %
120–129Superior6.7 %
110–119High average16.1 %
90–109Average50.0 %
80–89Low average16.1 %
70–79Borderline6.7 %
< 70Intellectual disability2.2 %
EXAMPLE 1 — Classical IQ

A child of CA = 8 yr performs at Mental Age 10 on a Binet test. IQ = 10/8 × 100 = 125.

EXAMPLE 2 — Deviation IQ

An adult scores 1.5 SDs above the mean. IQ = 100 + 15 × 1.5 = 122.5. This places them at approximately the 93rd percentile.

Important Limitations

Key Take-aways