Skip to the content

Topics Covered

Population, sample, parameter, statistic Sampling distributions Estimation Hypothesis testing The four tests Type I and Type II errors 📝 Practice problems Exam questions from this unit Mistakes that cost marks
On this page
  1. 5.1 Population, sample, parameter, statistic
  2. 5.2 Sampling distributions
  3. 5.3 Estimation
  4. 5.4 Hypothesis testing — the procedure
  5. 5.5 The four tests
  6. 5.6 Type I and Type II errors
  7. 📝 Practice problems
  8. Exam questions from this unit
  9. Mistakes that cost marks

Syllabus topics: Population and sample, parameters and statistics. Sampling distributions. Point and interval estimation (confidence intervals). Tests of significance — z-test, t-test, chi-square test and F-test. p-values and errors (Type I and II). Power of a statistical test.


5.1 Population, sample, parameter, statistic

THE BIG IDEA

You almost never measure everyone. You measure a few and reason about the rest. Statistical inference is the discipline of doing that reasoning honestly, with a stated degree of uncertainty.

IN DEPTH

A company claims its packets contain 500 g of rice. You cannot open every packet in the country. So you weigh 40 of them, find an average of 496 g, and ask: is the company lying, or is 496 just the sort of thing that happens when you weigh only 40 packets? That question — and the machinery for answering it — is this entire unit.

The four terms

Term Meaning Notation
Population Every member of the group of interest size N
Sample The subset actually measured size n
Parameter A number describing the population μ, σ, P
Statistic A number computed from the sample x̄, s, p̂

Greek letters for parameters, Roman letters for statistics. The parameter is fixed but unknown; the statistic is known but varies from sample to sample. Inference uses the second to estimate the first.

5.2 Sampling distributions

If you take many samples of the same size and compute the mean of each, those means form their own distribution — the sampling distribution of the mean.

From the Central Limit Theorem (Unit 3):

FORMULA

Mean of x̄ = μ Standard error = σ/√n Shape → normal as n grows

The standard error is not the standard deviation.

Standard deviation (σ or s) Standard error (σ/√n)
Describes Spread of individual values Spread of sample means
Depends on n? No Yes — shrinks as √n
Used for Describing data Inference

Confusing them is one of the most common errors in this unit. The standard error is always the smaller of the two (for n > 1), and it is what appears in every confidence interval and test statistic below.

5.3 Estimation

Point estimation

A single number as the best guess for a parameter: x̄ estimates μ, s estimates σ, p̂ estimates P.

Properties of a good estimator:

Property Meaning
Unbiased Its expected value equals the parameter: E(x̄) = μ
Consistent It converges to the parameter as n grows
Efficient It has the smallest variance among unbiased estimators
Sufficient It uses all the relevant information in the sample

(This is why the sample variance divides by n−1: doing so makes it unbiased. See Unit 1.)

The flaw of a point estimate is that it is almost certainly not exactly right, and it says nothing about how far off it might be.

Interval estimation — confidence intervals

NOTE

Estimate ± (critical value × standard error)

When σ is known (or n is large):

NOTE

x̄ ± z(α/2) · σ/√n

When σ is unknown (the usual case):

NOTE

x̄ ± t(α/2, n−1) · s/√n

Critical values worth memorising:

Confidence α z(α/2)
90% 0.10 1.645
95% 0.05 1.96
99% 0.01 2.576

WORKED EXAMPLE

n = 25, x̄ = 68, s = 5. Build a 95% confidence interval.

KEY INSIGHT

What "95% confident" actually means

If we repeated this sampling procedure many times and built an interval each time, about 95% of those intervals would contain the true population mean.

It does not mean "there is a 95% probability that μ lies in this particular interval". The true mean is a fixed number, not a random one — it is either in this interval or it is not. What is random is the interval, not μ.

Stating this correctly is worth marks; stating it the wrong way loses them.

Two behaviours to note and explain:

5.4 Hypothesis testing — the procedure

Follow these six steps every single time. Marks are given for each.

  1. State H₀ and H₁ — in symbols and in words
  2. Choose α — usually 0.05
  3. Compute the test statistic
  4. Find the p-value, or compare with the critical value
  5. Decide — p < α means reject H₀
  6. Conclude in the words of the original problem

The two hypotheses

Null H₀ Alternative H₁
Says No effect, no difference There is an effect
Contains Always = ≠, < or >
Status Assumed true until evidence says otherwise What you are trying to show

The courtroom analogy: H₀ is "innocent". You do not prove innocence; you either find enough evidence to reject it or you do not. "Fail to reject H₀" is not "accept H₀" — a jury returns "not guilty", never "innocent". Writing "accept H₀" is a standard mark deduction.

One-tailed vs two-tailed

Two-tailed One-tailed
H₁ μ ≠ μ₀ μ > μ₀ or μ < μ₀
Question "Is it different?" "Is it bigger?" / "Is it smaller?"
α split α/2 in each tail all α in one tail
z at α = 0.05 ±1.96 1.645

Decide the direction before seeing the data. Choosing a one-tailed test after looking at which way the result went is a form of cheating.

The p-value

NOTE

The p-value is the probability of observing a result at least as extreme as the one you got, assuming H₀ is true.

A small p-value means your data would be surprising if H₀ were true — so H₀ is doubtful.

p-value Reading
p < 0.01 Very strong evidence against H₀
p < 0.05 Strong evidence against H₀
0.05 < p < 0.10 Weak evidence
p > 0.10 Little or no evidence against H₀

What a p-value is not: it is not the probability that H₀ is true, and not the probability that your result happened by chance. Those misinterpretations are examined precisely because they are so common.

5.5 The four tests

The decision tree

  1. Comparing one mean to a known value, σ known or n > 30 → z-test
  2. Comparing one mean to a known value, σ unknown and n small → one-sample t-test
  3. Comparing two group means → two-sample t-test
  4. Same subjects measured twice → paired t-test
  5. Two categorical variables → chi-square test of independence
  6. Comparing two variances → F-test
  7. Comparing three or more means → ANOVA (uses F; beyond this syllabus)

z-test

Use when: σ is known, or n > 30.

FORMULA

z = (x̄ − μ₀) / (σ/√n)

WORKED EXAMPLE

Packets should average 70 g with a known σ of 3 g. A sample of 40 averages 71.2 g. Test at α = 0.05.

t-test

Use when: σ is unknown (the normal situation).

One-sample:

FORMULA

t = (x̄ − μ₀) / (s/√n), df = n − 1

Two-sample, equal variances (pooled):

FORMULA

t = (x̄₁ − x̄₂) / √(s²ₚ(1/n₁ + 1/n₂)), df = n₁ + n₂ − 2

where s²ₚ = [(n₁−1)s₁² + (n₂−1)s₂²] / (n₁ + n₂ − 2)

Paired:

FORMULA

t = d̄ / (s_d/√n), df = n − 1, where d is the difference within each pair

WORKED EXAMPLE

two-sample

Two teaching methods:

Why t and not z? The t distribution has heavier tails, accounting for the extra uncertainty in estimating σ from the sample. As n grows, t approaches z — by df = 30 they are nearly identical, which is where the "n > 30" rule of thumb comes from.

Chi-square test

Two uses. Both compare observed counts with expected counts.

FORMULA

χ² = Σ (O − E)² / E

Test of independence: are two categorical variables related?

Goodness of fit: does the data follow a claimed distribution?

WORKED EXAMPLE

Is region related to purchase type?

Region Premium (O) Standard (O) Total
North 30 70 100
South 45 55 100
East 25 75 100
Total 100 200 300

Assumption: every expected frequency should be at least 5. State that you checked it — here the smallest expected value is 33.33, so it holds.

Chi-square is always right-tailed: a large χ² means observed and expected differ a lot, which is the evidence against H₀.

F-test

Use when: comparing two variances — often to check the equal-variance assumption of the pooled t-test.

FORMULA

F = s₁² / s₂², with the larger variance on top df = (n₁ − 1, n₂ − 1)

WORKED EXAMPLE

, using the two groups above:

Putting the larger variance on top makes F ≥ 1 and lets you use the standard right-tail tables.

(All of these worked figures are computed in 05_inference_hypothesis_tests.py.)

5.6 Type I and Type II errors

H₀ is true H₀ is false
Reject H₀ Type I error (probability α) Correct — power (1 − β)
Fail to reject H₀ Correct (1 − α) Type II error (probability β)

In the courtroom analogy:

α is chosen by you. Setting α = 0.05 says you accept a 5% chance of rejecting a true H₀.

β follows from the design — the sample size, the effect size and α together determine it.

The trade-off

Lower α → fewer Type I errors, but more Type II errors. Tighten the standard of proof and more guilty people go free.

The only way to reduce both at once is to increase n.

Which error matters more depends on the context:

Power

FORMULA

Power = 1 − β = P(rejecting H₀ when it is genuinely false)

The probability of detecting a real effect. Conventionally, aim for 0.80.

Power increases when:

Factor Effect on power
Larger sample size n ↑
Larger true effect size ↑
Smaller population variance ↑
Larger α (e.g. 0.10 instead of 0.01) ↑ (but more Type I errors)
One-tailed instead of two-tailed ↑ (only if the direction is right)

A power analysis before collecting data tells you what sample size you need. Running an underpowered study wastes everyone's time: it will probably fail to detect a real effect, and you will not know whether the effect was absent or merely invisible.


📝 Practice problems

PROBLEM 1

A sample of 36 students has mean height 168 cm with sample standard deviation 6 cm. Construct a 95% confidence interval for the population mean height.

Solution.

Interpretation: if we repeated this sampling many times, about 95% of the intervals so constructed would contain the true mean height.

PROBLEM 2

A machine is supposed to fill bottles with 500 ml. A sample of 25 bottles gives a mean of 495 ml with a sample standard deviation of 8 ml. Test at α = 0.05 whether the machine is under-filling.

Solution.

PROBLEM 3

A survey of 200 people asks whether they prefer tea or coffee, split by gender. Test at α = 0.05 whether preference is independent of gender.

Tea Coffee Total
Male 40 60 100
Female 60 40 100
Total 100 100 200

Solution.

Tea (O, E) Coffee (O, E)
Male 40, 50 60, 50
Female 60, 50 40, 50

All expected frequencies are 50 ≥ 5 ✓


Exam questions from this unit

Two marks

  1. Distinguish a parameter from a statistic.
  2. What is the standard error, and how does it differ from the standard deviation?
  3. Define a Type I and a Type II error.
  4. What is the power of a test?
  5. State the correct interpretation of a 95% confidence interval.
  6. When do you use a t-test rather than a z-test?

Five marks

  1. Explain the steps of hypothesis testing.
  2. Construct a confidence interval for given sample data and interpret it.
  3. Explain Type I and Type II errors with the error table and the trade-off.
  4. Explain the chi-square test of independence with a worked example.
  5. Explain point and interval estimation, and the properties of a good estimator.

Ten marks

  1. Explain hypothesis testing in full — hypotheses, significance level, test statistics, p-values, errors and power — with a worked example.

  2. Explain the four tests of significance (z, t, chi-square, F) with their conditions, formulas and examples.

Mistakes that cost marks

COMMON ERRORS