Skip to the content

Topics Covered

Sigma Scaling (Difficulty) Z Scaling Standard Scores Normalized Scores T-Scores Percentile Scores Scaling of Rankings Scaling of Ratings
On this page
  1. 1. Introduction
  2. 2. Sigma Scaling of Individual Test Items (Item Difficulty)
  3. 3. Scaling of Scores — Z Scaling
  4. 4. Standard Scores (Linear Transformations)
  5. 5. Normalized Scores
  6. 6. T-Scores
  7. 7. Percentile Scores
  8. 8. Scaling of Rankings in Terms of the Normal Curve
  9. 9. Scaling of Ratings in Terms of the Normal Curve
  10. Summary — Common Scaling Schemes
  11. Key Take-aways

1. Introduction

DEFINITION

Psychological & educational statistics applies statistical methods to mental measurement — intelligence, aptitude, attitude, achievement etc. Raw test scores by themselves are often meaningless; they must be scaled to allow comparison across tests, across populations, or across time. This unit develops the standard scaling techniques.

Why Scale?

2. Sigma Scaling of Individual Test Items (Item Difficulty)

IDEA

Suppose a test item is answered correctly by a proportion \(p\) of a large sample of examinees. The difficulty of the item can be expressed in standard-normal units, so all items live on a common scale.

Procedure

  1. For each test item compute the proportion failing: \(q = 1 - p\).
  2. Find the standard-normal abscissa (z-value) such that \(P(Z > z) = p\) — i.e., the upper \(p\)-percentile of \(N(0, 1)\).
  3. Call this z-value the item's sigma difficulty \(\sigma_i\). Easier items (large \(p\)) have negative sigma; difficult items have positive sigma.
  4. Optionally transform to a fixed mean and SD (often mean 5, SD 2) for reporting.

Why Z-values?

If ability is assumed normally distributed, the probability of passing an item is related to the area under the normal curve above the item's difficulty point. So z-values give a linear, additive scale on the latent ability dimension.

EXAMPLE 1

An item is passed by 84 % of students. \(p = 0.84\), \(q = 0.16\). The z-value such that \(P(Z > z) = 0.84\) is \(z = -1.0\).

So the item's difficulty is \(\sigma_i = -1.0\) — an easy item.

EXAMPLE 2

Only 16 % pass an item. \(p = 0.16\). \(z\) such that \(P(Z > z) = 0.16\) is \(z = 1.0\). Difficulty \(\sigma_i = +1.0\) — a hard item.

By combining many items, an entire test can be calibrated to a common difficulty scale, enabling item-banking and adaptive testing.

3. Scaling of Scores — Z Scaling

DEFINITION

The Z score expresses a raw score as the number of standard deviations above or below the mean of the reference distribution:

\[ Z \;=\; \dfrac{X - \mu}{\sigma}. \]

\(\mu\) and \(\sigma\) are the mean and SD of the group being normed. The Z-score distribution has mean 0 and SD 1.

Properties of Z Scaling

EXAMPLE 1

A student scores 80 in Math (class mean 70, SD 8) and 75 in English (mean 65, SD 5). Which is the relatively better performance?

\(Z_{\text{Math}} = (80-70)/8 = 1.25\); \(Z_{\text{English}} = (75-65)/5 = 2.0\). English is the stronger performance.

EXAMPLE 2

For a sample mean 60 and SD 10, a raw score of 75 has \(Z = 1.5\) — 1.5 SDs above the mean, equivalent to the 93rd percentile under normality.

4. Standard Scores (Linear Transformations)

A standard score rescales the Z-score onto a more readable scale with a specified mean and SD:

\[ \text{Standard score} \;=\; \mu^{*} + \sigma^{*} \cdot Z, \]

where \(\mu^{*}\) and \(\sigma^{*}\) are the chosen target mean and SD.

Common Standard Scales

ScaleTarget meanTarget SDFormula
Z-score01\(Z\)
T-score5010\(50 + 10 Z\)
SAT-score500100\(500 + 100 Z\)
IQ10015\(100 + 15 Z\)
Stanine52\(5 + 2 Z\), clipped 1–9

All linear standard scores preserve relative spacing — they have the same correlations with other variables as the underlying Z.

5. Normalized Scores

DEFINITION

A normalized score first transforms raw scores to percentile ranks, then converts those percentiles back to z-values under a normal distribution. This changes the shape of the distribution to normal, regardless of the original shape.

Procedure

  1. Rank examinees from worst to best.
  2. Compute the percentile rank \(P_i\) of each (proportion of group at or below that score).
  3. Find the corresponding normal-curve z-value: \(z_i = \Phi^{-1}(P_i / 100)\).
  4. (Optional) transform to a chosen scale (T = 50 + 10 z, IQ = 100 + 15 z, etc.).

Linear Z vs Normalized Z

EXAMPLE 1

A student is at the 84th percentile in a non-normal distribution. Normalized Z = \(\Phi^{-1}(0.84) = 1.0\). Normalized T-score = 50 + 10 × 1.0 = 60.

EXAMPLE 2

Raw scores are highly right-skewed. Normalizing brings them onto a symmetric bell curve, enabling parametric tests (t, ANOVA) that assume normality.

6. T-Scores

DEFINITION

A T-score is a standard score with mean 50 and SD 10. It can be either linear (from raw scores via Z) or normalized (from percentile via inverse normal).

\[ T \;=\; 50 + 10 Z. \]

Why T-Scores?

Widely used in psychological testing (MMPI personality inventory, NEO-PI, etc.).

EXAMPLE

A raw score \(X = 78\) on a test with mean 70 and SD 8 gives \(Z = 1.0\); T = 50 + 10 × 1.0 = 60.

7. Percentile Scores

DEFINITION

The percentile score (or percentile rank) of a raw score \(X\) is the percentage of examinees whose raw scores are less than or equal to \(X\). It expresses the raw score's position in the distribution.

Formula

\[ P_X \;=\; \dfrac{\text{Number of scores} \le X}{N} \times 100. \]

Alternative formula (interpolated): \(P_X = \dfrac{C + 0.5 f}{N} \times 100\), where \(C\) is the cumulative frequency below \(X\)'s class and \(f\) is the frequency of \(X\)'s class.

Inverse — Score at a Given Percentile

The 90th percentile is the score at or below which 90 % of the group falls. Computed by interpolating in a cumulative-frequency distribution.

Interpretation

EXAMPLE 1

A student scores 82 on a test taken by 1000 students; 850 score 82 or below. Percentile rank = 85 — i.e., the student is at the 85th percentile (better than 85 % of peers).

EXAMPLE 2

If raw scores follow \(N(60, 10)\), what raw score corresponds to the 90th percentile?

\(z = \Phi^{-1}(0.90) = 1.282\). Raw score = 60 + 1.282 × 10 = 72.82. So scoring ~73 puts you at the 90th percentile.

8. Scaling of Rankings in Terms of the Normal Curve

Sometimes we have only ranks (1st, 2nd, …, Nth) — for example, ordered preferences or class rank — and wish to convert them to a normal-scaled score.

Procedure

  1. For each rank \(r\), compute its percentile equivalent. The formula depends on which end rank 1 is: \[ P_r \;=\; \dfrac{r - 0.5}{N} \times 100 \quad (\text{rank 1 = lowest}), \qquad P_r \;=\; \dfrac{N - r + 0.5}{N} \times 100 \quad (\text{rank 1 = best}). \] Each of the \(N\) ranks stands for a slice \(1/N\) of the group; the \(0.5\) places the rank at the middle of its slice, so that no rank is put at the 0th or 100th percentile.
  2. Convert to z-value: \(z_r = \Phi^{-1}(P_r/100)\).
  3. (Optional) Linear-transform to T-score, IQ etc.

Caution about direction: if rank 1 means the highest performer, treat it as the highest score; the corresponding percentile is close to 100, so z is positive and large.

EXAMPLE

For \(N = 100\) students ranked 1 (best) to 100 (worst):

9. Scaling of Ratings in Terms of the Normal Curve

When raters assign one of several ordered categories (e.g., Poor / Average / Good / Very Good / Excellent), the categories need to be converted to numeric scores for further statistical analysis.

Procedure

  1. Tabulate frequencies in each rating category.
  2. Compute cumulative percentages.
  3. Find the midpoint percentile for each category (sum of cumulative below + half of current category's frequency).
  4. Convert to z-value via \(\Phi^{-1}\).
  5. (Optional) Re-scale to T / 5-point Likert / any scale of choice.

This procedure assumes the latent attribute being rated is continuous and normally distributed; the discrete rating categories are an approximation.

EXAMPLE

200 raters classify a product on a 5-point scale:

CategoryFrequencyCum %Midpoint %z
Poor1052.5−1.96
Average402515−1.04
Good806545−0.13
Very Good509077.5+0.76
Excellent2010095+1.65

The z values are the normal-curve scaled scores of each category. Now raters who "Liked the product" can be assigned z = +0.76; those rating "Excellent" get z = +1.65. These scores are now amenable to mean, variance, regression analyses.

Summary — Common Scaling Schemes

ScaleMeanSDUse
Z01Universal; can be negative
T5010Personality / clinical tests
SAT/GRE500100College entrance
IQ10015 (Wechsler)Intelligence
Stanine52School achievement (1–9 integers)
Percentile——Position (uniform on 1–99)

Key Take-aways