Psychological & educational statistics applies statistical methods to mental measurement — intelligence, aptitude, attitude, achievement etc. Raw test scores by themselves are often meaningless; they must be scaled to allow comparison across tests, across populations, or across time. This unit develops the standard scaling techniques.
Suppose a test item is answered correctly by a proportion \(p\) of a large sample of examinees. The difficulty of the item can be expressed in standard-normal units, so all items live on a common scale.
If ability is assumed normally distributed, the probability of passing an item is related to the area under the normal curve above the item's difficulty point. So z-values give a linear, additive scale on the latent ability dimension.
An item is passed by 84 % of students. \(p = 0.84\), \(q = 0.16\). The z-value such that \(P(Z > z) = 0.84\) is \(z = -1.0\).
So the item's difficulty is \(\sigma_i = -1.0\) — an easy item.
Only 16 % pass an item. \(p = 0.16\). \(z\) such that \(P(Z > z) = 0.16\) is \(z = 1.0\). Difficulty \(\sigma_i = +1.0\) — a hard item.
By combining many items, an entire test can be calibrated to a common difficulty scale, enabling item-banking and adaptive testing.
The Z score expresses a raw score as the number of standard deviations above or below the mean of the reference distribution:
\[ Z \;=\; \dfrac{X - \mu}{\sigma}. \]\(\mu\) and \(\sigma\) are the mean and SD of the group being normed. The Z-score distribution has mean 0 and SD 1.
A student scores 80 in Math (class mean 70, SD 8) and 75 in English (mean 65, SD 5). Which is the relatively better performance?
\(Z_{\text{Math}} = (80-70)/8 = 1.25\); \(Z_{\text{English}} = (75-65)/5 = 2.0\). English is the stronger performance.
For a sample mean 60 and SD 10, a raw score of 75 has \(Z = 1.5\) — 1.5 SDs above the mean, equivalent to the 93rd percentile under normality.
A standard score rescales the Z-score onto a more readable scale with a specified mean and SD:
where \(\mu^{*}\) and \(\sigma^{*}\) are the chosen target mean and SD.
| Scale | Target mean | Target SD | Formula |
|---|---|---|---|
| Z-score | 0 | 1 | \(Z\) |
| T-score | 50 | 10 | \(50 + 10 Z\) |
| SAT-score | 500 | 100 | \(500 + 100 Z\) |
| IQ | 100 | 15 | \(100 + 15 Z\) |
| Stanine | 5 | 2 | \(5 + 2 Z\), clipped 1–9 |
All linear standard scores preserve relative spacing — they have the same correlations with other variables as the underlying Z.
A normalized score first transforms raw scores to percentile ranks, then converts those percentiles back to z-values under a normal distribution. This changes the shape of the distribution to normal, regardless of the original shape.
A student is at the 84th percentile in a non-normal distribution. Normalized Z = \(\Phi^{-1}(0.84) = 1.0\). Normalized T-score = 50 + 10 × 1.0 = 60.
Raw scores are highly right-skewed. Normalizing brings them onto a symmetric bell curve, enabling parametric tests (t, ANOVA) that assume normality.
A T-score is a standard score with mean 50 and SD 10. It can be either linear (from raw scores via Z) or normalized (from percentile via inverse normal).
\[ T \;=\; 50 + 10 Z. \]Widely used in psychological testing (MMPI personality inventory, NEO-PI, etc.).
A raw score \(X = 78\) on a test with mean 70 and SD 8 gives \(Z = 1.0\); T = 50 + 10 × 1.0 = 60.
The percentile score (or percentile rank) of a raw score \(X\) is the percentage of examinees whose raw scores are less than or equal to \(X\). It expresses the raw score's position in the distribution.
Alternative formula (interpolated): \(P_X = \dfrac{C + 0.5 f}{N} \times 100\), where \(C\) is the cumulative frequency below \(X\)'s class and \(f\) is the frequency of \(X\)'s class.
The 90th percentile is the score at or below which 90 % of the group falls. Computed by interpolating in a cumulative-frequency distribution.
A student scores 82 on a test taken by 1000 students; 850 score 82 or below. Percentile rank = 85 — i.e., the student is at the 85th percentile (better than 85 % of peers).
If raw scores follow \(N(60, 10)\), what raw score corresponds to the 90th percentile?
\(z = \Phi^{-1}(0.90) = 1.282\). Raw score = 60 + 1.282 × 10 = 72.82. So scoring ~73 puts you at the 90th percentile.
Sometimes we have only ranks (1st, 2nd, …, Nth) — for example, ordered preferences or class rank — and wish to convert them to a normal-scaled score.
Caution about direction: if rank 1 means the highest performer, treat it as the highest score; the corresponding percentile is close to 100, so z is positive and large.
For \(N = 100\) students ranked 1 (best) to 100 (worst):
When raters assign one of several ordered categories (e.g., Poor / Average / Good / Very Good / Excellent), the categories need to be converted to numeric scores for further statistical analysis.
This procedure assumes the latent attribute being rated is continuous and normally distributed; the discrete rating categories are an approximation.
200 raters classify a product on a 5-point scale:
| Category | Frequency | Cum % | Midpoint % | z |
|---|---|---|---|---|
| Poor | 10 | 5 | 2.5 | −1.96 |
| Average | 40 | 25 | 15 | −1.04 |
| Good | 80 | 65 | 45 | −0.13 |
| Very Good | 50 | 90 | 77.5 | +0.76 |
| Excellent | 20 | 100 | 95 | +1.65 |
The z values are the normal-curve scaled scores of each category. Now raters who "Liked the product" can be assigned z = +0.76; those rating "Excellent" get z = +1.65. These scores are now amenable to mean, variance, regression analyses.
| Scale | Mean | SD | Use |
|---|---|---|---|
| Z | 0 | 1 | Universal; can be negative |
| T | 50 | 10 | Personality / clinical tests |
| SAT/GRE | 500 | 100 | College entrance |
| IQ | 100 | 15 (Wechsler) | Intelligence |
| Stanine | 5 | 2 | School achievement (1–9 integers) |
| Percentile | — | — | Position (uniform on 1–99) |