Narrow the table
What are you asking?
Is the data roughly normal?
What are you comparing?
What kind of measurement?
Are the two sets of numbers linked?
Showing all 13 tests.
| You want to know | The test | It assumes | If that fails, use | Taught in |
|---|---|---|---|---|
| Do these two things move together? Do students who attend more classes score higher? | Pearson’s r | Both numeric, the relationship roughly a straight line, and no wild outliers — one stray point can create or destroy a correlation on its own. | Spearman’s rank, which correlates the ranks instead. It needs only that the relationship goes one way, and it takes ordinal data. | Correlation · Spearman’s rank |
| Can I predict one from the other? What mark would 12 hours of study give? | Regression | The same as correlation, plus that you have decided which variable explains which. That decision comes from you, not from the data. | — correlation still measures the strength without committing to a direction, which is the honest answer when you cannot justify one. | Regression |
| Is this group’s mean different from a value I already know? Is the mean mark different from the pass mark of 40? | One-sample t-test | Numeric data; the values roughly normal, or a sample large enough for the central limit theorem to carry it. | Sign test, which asks only whether values fall above or below the median and assumes almost nothing. | Small Sample Tests · Non-parametric Tests |
| Same question, but my sample is large. A mean or a proportion, from 200 responses. | z-test | A large sample, so the sampling distribution is near normal whatever the population looks like. For a proportion, enough successes and failures. | — at this size the t and z tests agree closely anyway. | Large Sample Tests |
| Do these two separate groups differ? Do students taught two ways score differently? | Independent two-sample t-test | Numeric; each group roughly normal; similar spread in both. Unequal spread is handled by the Welch form rather than by giving up. | Mann–Whitney U, which compares ranks instead of means and does not need normality. | Small Sample Tests · Non-parametric Tests |
| Did the same subjects change? Marks before and after coaching, same students. | Paired t-test | That the differences are roughly normal — not the original values. This is the assumption most often checked on the wrong column. | Wilcoxon signed-rank, which ranks the differences. | Small Sample Tests · Non-parametric Tests |
| Do three or more groups differ? Yields under four fertiliser treatments. | One-way ANOVA | Numeric; roughly normal within each group; similar spread across groups; independent observations. | Kruskal–Wallis, the rank-based equivalent. | Analysis of Variance · Non-parametric Tests |
| Do two things each affect the outcome? Does the teaching method matter, and does it matter which class? | Two-way ANOVA | The same as one-way ANOVA, applied to each factor: numeric, roughly normal within groups, similar spread. | — with one factor, one-way ANOVA above; this site does not cover a non-parametric two-factor test. | Analysis of Variance |
| Is one group more variable than the other? Is machine A’s output more consistent than machine B’s? | F-test for two variances | Both samples drawn from normal populations. This test is unusually sensitive to that — non-normality misleads it more than it misleads a t-test. | — see the note on equal-variance tests below. | Small Sample Tests |
| Do the counts match what I expected? Are births spread evenly across the days of the week? | Chi-square goodness of fit | Actual counts, never percentages; observations independent; most expected counts at least 5. | — combine sparse categories until the expected counts are large enough. | Small Sample Tests |
| Are these two categorical variables related? Does drink preference depend on gender? | Chi-square test of independence | A contingency table of counts; independent observations; most expected counts at least 5. | — as above, or collapse categories. | Small Sample Tests · Inference in Python |
| The same question, but the counts are small. A 2×2 table where one cell holds three people. | Fisher’s exact test | Nothing about normality. It computes the probability directly rather than approximating it, which is exactly why it survives small counts. | — this is the answer when chi-square’s expected-count rule fails. Reach for it whenever an expected count drops below about five. | Categorical Outcomes — Clinical Trials |
| Is this sequence random, or is there a pattern? Do defects arrive at random along a production run? | Run test | A sequence in the order it occurred. Sorting the data first destroys the very thing the test looks at. | — | Non-parametric Tests |
Checking the assumptions
Every parametric row above rests on two of them, and they are the step most often skipped:
- Is the data roughly normal?
- Shapiro–Wilk is the usual formal check, alongside a histogram and a Q–Q plot — and looking at the picture matters as much as the p-value, because with a large sample the test flags departures too small to affect anything. On this site it is covered in the R material: Inferential Statistics & Hypothesis Testing in R.
- Do the groups have similar spread?
- This site does not yet teach Levene’s, Bartlett’s or the Brown–Forsythe test in full. These are the standard formal checks. Levene’s is covered only as output to read: the SPSS practical shows how its result decides which t row to use (Statistical Analysis Using SPSS, Practical §4), but not how the test works. Brown–Forsythe appears nowhere at all; and the word “Bartlett” does appear here, but attached to Bartlett’s formula and the Bartlett window in time series — a different thing that is easy to mistake for the test. In the meantime, compare the group standard deviations directly — a ratio beyond about two to one is the point at which the equal-variance assumption stops being safe — and for two groups use the Welch form of the t-test, which does not need the assumption.
- What if my groups are paired and there are more than two of them?
- This site does not cover Friedman’s test, repeated-measures ANOVA, or
McNemar’s test. They are the standard answers to three cells of the usual decision
tree, and each is named here only in passing — Friedman in a syllabus and one MCQ,
McNemar once in a data-mining unit, and “repeated measures” only in the context
of working out a sample size, which is a different question from testing.
What not to do: Kruskal–Wallis is not a substitute for Friedman, and one-way ANOVA is not a substitute for the repeated-measures form. Both assume the groups are independent, and pairing is exactly the thing they would throw away — you would be answering a question you were not asking, with more error than you need. For two paired groups the Wilcoxon signed-rank and paired t-tests above are the right tools; beyond two, this site cannot yet take you.
What a significant result does not tell you
- Not that the effect is large. With a big enough sample almost any difference reaches significance. Report the size of the difference and a confidence interval, not only the p-value.
- Not the probability that the null hypothesis is true. A p-value is the probability of data this extreme if the null were true — the conditional runs the other way.
- Not that the result will hold next time. Significance is not replication.
- And a non-significant result is not proof of no effect. It may only mean the sample was too small to see one.
Topics Covered
Beyond choosing a test
This page covers the comparison tests. For relationships and prediction see correlation and regression; for designed experiments see Design and Analysis of Experiments. Everything on the site is listed in the A–Z index.