Skip to the content

Topics Covered

Fundamentals of Sample Size Dichotomous Response Variables Continuous Response Variables Repeated Measures Variables Power and Significance Level Practical Considerations
On this page
  1. 1. Fundamentals of Sample Size Determination
  2. 2. Sample Size for Two Independent Samples of Dichotomous Response Variables
  3. 3. Sample Size for Two Independent Samples of Continuous Response Variables
  4. 4. Sample Size for Repeated Measures Variables

1. Fundamentals of Sample Size Determination

KEY CONCEPT

Sample size determination is the process of calculating the minimum number of subjects needed in a clinical trial to reliably detect a clinically meaningful treatment effect if one truly exists. An under-powered trial (too few subjects) may fail to detect a real effect (Type II error), while an over-powered trial (too many subjects) unnecessarily exposes more participants to experimental treatments and wastes resources.

1.1 Key Parameters

Every sample size calculation depends on four parameters:

THE FOUR PILLARS
  1. Significance level (\(\alpha\)): The probability of a Type I error (rejecting \(H_0\) when it is true). Conventionally set at \(\alpha = 0.05\) (two-sided) or \(\alpha = 0.025\) (one-sided).
  2. Power (\(1 - \beta\)): The probability of correctly rejecting \(H_0\) when the alternative hypothesis is true. Conventionally set at 80% or 90%. Power = \(1 - \beta\), where \(\beta\) is the Type II error rate.
  3. Effect size (\(\delta\)): The minimum clinically meaningful difference between the treatment groups that the trial is designed to detect. This is determined by clinical judgement, not statistics. For example, a 5 mmHg reduction in blood pressure, or a 15% improvement in response rate.
  4. Variability (\(\sigma^2\)): The variance of the outcome measure, estimated from previous studies, pilot data, or published literature.

Given any three of these parameters, the fourth can be determined. In practice, \(\alpha\) and power are fixed, \(\delta\) and \(\sigma\) are estimated from prior knowledge, and \(n\) is calculated.

1.2 General Formula Structure

GENERAL SAMPLE SIZE FORMULA

For a two-group comparison:

\[ n = \frac{(z_{\alpha/2} + z_{\beta})^2 \times f(\sigma, \delta)}{(\text{effect size})^2} \]

where \(z_{\alpha/2}\) and \(z_{\beta}\) are the standard normal quantiles, and \(f(\sigma, \delta)\) depends on the type of outcome variable (dichotomous, continuous, or repeated measures).

Common z-values: \(z_{0.025} = 1.96\) (two-sided \(\alpha = 0.05\)), \(z_{0.20} = 0.842\) (80% power), \(z_{0.10} = 1.282\) (90% power).

EXAMPLE 1 — Impact of Parameters on Sample Size

Consider a trial to compare two treatments where the effect size \(\delta = 5\) and \(\sigma = 15\). The sample size per group (continuous outcome) is approximately:

\[ n \approx \frac{2\sigma^2(z_{\alpha/2} + z_{\beta})^2}{\delta^2} \]

For \(\alpha = 0.05\) (two-sided) and 80% power: \(n = \frac{2 \times 225 \times (1.96 + 0.842)^2}{25} = \frac{450 \times 7.85}{25} = \frac{3532.5}{25} = 141.3 \approx 142\) per group.

If we increase power to 90%: \(n = \frac{450 \times (1.96 + 1.282)^2}{25} = \frac{450 \times 10.51}{25} = 189.2 \approx 190\) per group — a 34% increase in sample size for a 10% increase in power.

If the effect size is only \(\delta = 3\) (smaller, harder to detect): \(n = \frac{450 \times 7.85}{9} = 392.5 \approx 393\) per group — sample size nearly triples. This illustrates that the required sample size is extremely sensitive to the effect size (it enters the formula as \(\delta^2\) in the denominator).

EXAMPLE 2 — Clinical vs. Statistical Significance in Sample Size

A researcher is planning a trial of a new antihypertensive. The standard drug reduces systolic BP by a mean of 10 mmHg. The new drug is expected to reduce it by 12 mmHg — a difference of \(\delta = 2\) mmHg. With \(\sigma = 12\) mmHg and 80% power:

\[ n = \frac{2 \times 144 \times 7.85}{4} = \frac{2260.8}{4} = 565.2 \approx 566 \text{ per group} \]

A total of 1,132 patients are needed to detect a 2 mmHg difference. Is this clinically worthwhile? Most clinicians would say no — a 2 mmHg difference is trivial in practice. However, with a large enough sample, even this tiny difference would be "statistically significant." This highlights the importance of choosing a clinically meaningful effect size, not merely a statistically detectable one. If the researcher instead chooses \(\delta = 5\) mmHg as the minimum clinically important difference, the sample size drops to \(n = \frac{2 \times 144 \times 7.85}{25} = 90.4 \approx 91\) per group (182 total) — a much more feasible and meaningful trial.

2. Sample Size for Two Independent Samples of Dichotomous Response Variables

KEY CONCEPT

A dichotomous response is a binary outcome such as success/failure, response/no response, alive/dead, or event/no event. The treatment effect is measured as the difference in proportions (\(p_1 - p_2\)) or the odds ratio. Sample size formulas for dichotomous outcomes are among the most commonly used in clinical trial design.

2.1 Formula for Comparing Two Proportions

SAMPLE SIZE FOR DIFFERENCE IN PROPORTIONS

For a two-sided test comparing proportions \(p_1\) and \(p_2\) with equal allocation (\(n_1 = n_2 = n\)):

\[ n = \frac{(z_{\alpha/2} + z_{\beta})^2 \left[p_1(1-p_1) + p_2(1-p_2)\right]}{(p_1 - p_2)^2} \]

Alternatively, using the pooled proportion \(\bar{p} = \frac{p_1 + p_2}{2}\):

\[ n = \frac{(z_{\alpha/2} + z_{\beta})^2 \times 2\bar{p}(1-\bar{p})}{(p_1 - p_2)^2} \]

Total sample size: \(N = 2n\).

2.2 Using the Odds Ratio

SAMPLE SIZE VIA ODDS RATIO

If the treatment effect is specified as an odds ratio \(\psi = \frac{p_1/(1-p_1)}{p_2/(1-p_2)}\), then:

\[ p_1 = \frac{\psi p_2}{1 + p_2(\psi - 1)} \]

and the sample size formula above can be used with this calculated \(p_1\).

EXAMPLE 1 — Sample Size for a Cancer Response Trial

A Phase III trial compares a new chemotherapy regimen (Treatment 1) with the standard regimen (Treatment 2) for advanced lung cancer. The primary endpoint is the tumour response rate (complete or partial response). Based on prior studies, the expected response rates are:

The clinically meaningful difference is \(\delta = p_1 - p_2 = 0.15\). For \(\alpha = 0.05\) (two-sided) and 90% power:

\[ n = \frac{(1.96 + 1.282)^2 \times [0.40 \times 0.60 + 0.25 \times 0.75]}{(0.15)^2} \]

\[ = \frac{10.51 \times (0.24 + 0.1875)}{0.0225} = \frac{10.51 \times 0.4275}{0.0225} = \frac{4.493}{0.0225} = 199.7 \approx 200 \text{ per group} \]

Total sample size: \(N = 400\). Accounting for 10% dropout, enrol \(200/0.90 = 223\) per group, total \(N = 446\).

EXAMPLE 2 — Sample Size for a Vaccine Efficacy Trial

A vaccine trial aims to demonstrate that the vaccine reduces the infection rate from \(p_2 = 0.04\) (4% in the placebo group) to \(p_1 = 0.01\) (1% in the vaccine group). The expected vaccine efficacy is \(VE = 1 - 0.01/0.04 = 75\%\). For \(\alpha = 0.05\) (two-sided) and 80% power:

\[ n = \frac{(1.96 + 0.842)^2 \times [0.01 \times 0.99 + 0.04 \times 0.96]}{(0.03)^2} \]

\[ = \frac{7.85 \times (0.0099 + 0.0384)}{0.0009} = \frac{7.85 \times 0.0483}{0.0009} = \frac{0.3792}{0.0009} = 421.3 \approx 422 \text{ per group} \]

Total: 844 participants. However, because the event rate is low (4%), the number of events is small. With 422 per group, we expect only about 4 events in the vaccine group (422 × 0.01) and 17 in the placebo group (422 × 0.04). For rare events, alternative formulas (e.g., based on the Poisson distribution or the log-rank test for time-to-event data) may be more appropriate and yield different sample sizes.

3. Sample Size for Two Independent Samples of Continuous Response Variables

KEY CONCEPT

A continuous response is a quantitative outcome measured on an interval or ratio scale, such as blood pressure, cholesterol level, tumour size, or a quality-of-life score. The treatment effect is typically measured as the difference in means (\(\mu_1 - \mu_2\)). The sample size depends on the effect size relative to the standard deviation (Cohen's d = \(\delta/\sigma\)).

3.1 Formula for Comparing Two Means

SAMPLE SIZE FOR DIFFERENCE IN MEANS

For a two-sided two-sample t-test with equal variances and equal allocation:

\[ n = \frac{2\sigma^2(z_{\alpha/2} + z_{\beta})^2}{(\mu_1 - \mu_2)^2} \]

where \(\sigma^2\) is the common variance (assumed equal in both groups), \(\mu_1 - \mu_2 = \delta\) is the clinically meaningful difference, and \(n\) is the sample size per group.

Total sample size: \(N = 2n\).

3.2 Effect Size (Cohen's d)

STANDARDISED EFFECT SIZE

\[ d = \frac{\mu_1 - \mu_2}{\sigma} = \frac{\delta}{\sigma} \]

Conventions: \(d = 0.2\) (small), \(d = 0.5\) (medium), \(d = 0.8\) (large).

The formula can be rewritten as:

\[ n = \frac{2(z_{\alpha/2} + z_{\beta})^2}{d^2} \]

For \(\alpha = 0.05\) (two-sided) and 80% power: \(n \approx \frac{15.7}{d^2}\).

3.3 Accounting for Dropouts

ADJUSTED SAMPLE SIZE

If the expected dropout rate is \(r\) (expressed as a proportion), inflate the sample size:

\[ n_{\text{adjusted}} = \frac{n}{1 - r} \]

For example, with a 15% expected dropout rate: \(n_{\text{adjusted}} = n / 0.85\).

EXAMPLE 1 — Sample Size for a Cholesterol-Lowering Drug Trial

A pharmaceutical company plans a Phase III trial of a new statin. The primary endpoint is the change in LDL cholesterol from baseline to Week 12. Based on prior data:

For \(\alpha = 0.05\) (two-sided) and 90% power:

\[ n = \frac{2 \times 625 \times (1.96 + 1.282)^2}{100} = \frac{1250 \times 10.51}{100} = \frac{13137.5}{100} = 131.4 \approx 132 \text{ per group} \]

Total: \(N = 264\). With an expected 10% dropout: \(n_{\text{adj}} = 132/0.90 = 147\) per group, \(N = 294\).

Effect size: \(d = 10/25 = 0.4\) (between small and medium). This is a moderate effect that requires a reasonably sized trial. If the company only expects a 5 mg/dL advantage (\(d = 0.2\), small effect), the required sample size would be \(n = \frac{2 \times 625 \times 10.51}{25} = 525.5 \approx 526\) per group (1,052 total) — four times larger.

EXAMPLE 2 — Sample Size for a Blood Pressure Trial

A trial compares a new antihypertensive with placebo. The primary endpoint is the change in seated systolic blood pressure from baseline to 8 weeks. Parameters:

For \(\alpha = 0.05\) and 80% power:

\[ n = \frac{2 \times 324 \times (1.96 + 0.842)^2}{100} = \frac{648 \times 7.85}{100} = \frac{5086.8}{100} = 50.9 \approx 51 \text{ per group} \]

With 80% power, only 51 per group (102 total) is needed because the effect size is moderate and the variability is manageable. However, this assumes perfect adherence and no dropouts. With 20% dropout: \(n_{\text{adj}} = 51/0.80 = 64\) per group (128 total). The trial should also consider using a non-inferiority design if the objective is to show the new drug is not worse than an active comparator, which typically requires a larger sample size.

4. Sample Size for Repeated Measures Variables

KEY CONCEPT

Repeated measures designs involve measuring the same outcome variable at multiple time points on each subject (e.g., blood pressure measured at Weeks 0, 4, 8, and 12). Compared to a single endpoint analysis, repeated measures designs are more efficient because they utilise all available data and account for within-subject correlation. However, sample size calculation is more complex because it must account for the correlation structure.

4.1 The Role of Within-Subject Correlation

Measurements on the same subject are correlated — a patient's blood pressure at Week 4 is related to their blood pressure at Week 8. This within-subject correlation (\(\rho\)) affects the effective sample size. Higher correlation means less additional information from repeated measurements.

EFFECTIVE SAMPLE SIZE IN REPEATED MEASURES

If each subject has \(k\) measurements with a common within-subject correlation \(\rho\), the variance of the subject-level mean is:

\[ \text{Var}(\bar{Y}_i) = \frac{\sigma^2}{k}\left[1 + (k-1)\rho\right] \]

The factor \([1 + (k-1)\rho]\) is called the design effect. It inflates the variance relative to the case of independent measurements.

When \(\rho = 0\) (independent measurements): design effect = 1 (no inflation).

When \(\rho = 1\) (perfectly correlated): design effect = \(k\) (repeated measurements provide no additional information beyond the first).

4.2 Sample Size Formula for Repeated Measures ANOVA

SAMPLE SIZE FOR REPEATED MEASURES

For a two-group comparison of the time-averaged difference with \(k\) repeated measures:

\[ n = \frac{2\sigma^2(z_{\alpha/2} + z_{\beta})^2 \left[1 + (k-1)\rho\right]}{k \times \delta^2} \]

where \(\sigma^2\) is the variance at a single time point, \(\rho\) is the within-subject correlation, and \(\delta\) is the clinically meaningful difference in the time-averaged mean.

Compared to a single-measurement design:

\[ \frac{n_{\text{repeated}}}{n_{\text{single}}} = \frac{1 + (k-1)\rho}{k} \]

Repeated measures reduce the required sample size when \(\rho < 1\). The reduction is greatest when \(\rho\) is low (measurements provide largely independent information) and \(k\) is large (many measurements).

4.3 Sample Size for the Rate of Change (Slope)

SAMPLE SIZE FOR SLOPE COMPARISON

If the outcome of interest is the rate of change (slope) over time, and measurements are equally spaced at times \(t_1, t_2, \ldots, t_k\), the variance of the estimated slope per subject is:

\[ \text{Var}(\hat{\beta}_i) = \frac{\sigma_e^2}{\sum_{j=1}^{k}(t_j - \bar{t})^2} \]

where \(\sigma_e^2 = \sigma^2(1 - \rho)\) is the residual variance after accounting for the within-subject correlation. The sample size per group for comparing slopes is:

\[ n = \frac{2\sigma_e^2(z_{\alpha/2} + z_{\beta})^2}{\delta_{\text{slope}}^2 \times \sum(t_j - \bar{t})^2} \]

EXAMPLE 1 — Repeated Measures Diabetes Trial

A diabetes trial measures HbA1c at baseline, Week 12, and Week 24 (k = 3 time points). The treatment effect of interest is the time-averaged difference in HbA1c change. Parameters:

For \(\alpha = 0.05\) and 80% power:

\[ n = \frac{2 \times 1.44 \times 7.85 \times [1 + 2 \times 0.6]}{3 \times 0.25} = \frac{2 \times 1.44 \times 7.85 \times 2.2}{0.75} = \frac{49.72}{0.75} = 66.3 \approx 67 \text{ per group} \]

For comparison, a single-measurement design would require:

\[ n_{\text{single}} = \frac{2 \times 1.44 \times 7.85}{0.25} = \frac{22.61}{0.25} = 90.4 \approx 91 \text{ per group} \]

The repeated measures design reduces the required sample size by approximately 26% (from 91 to 67 per group) by exploiting the within-subject correlation to gain efficiency. The total number of measurements is higher (67 × 3 = 201 vs. 91 × 1 = 91), but the number of subjects is lower, which is usually the binding constraint.

EXAMPLE 2 — Slope Comparison in an Alzheimer's Trial

An Alzheimer's disease trial measures cognitive function (ADAS-Cog score) at baseline and every 6 months for 2 years (k = 5 measurements at t = 0, 6, 12, 18, 24 months). The treatment effect of interest is the difference in the rate of cognitive decline (slope of ADAS-Cog over time).

Parameters: \(\sigma = 6.0\), \(\rho = 0.7\), expected slope difference: \(\delta_{\text{slope}} = 1.5\) points per year.

Residual variance: \(\sigma_e^2 = 36 \times (1 - 0.7) = 10.8\)

\(\sum(t_j - \bar{t})^2 = (0-12)^2 + (6-12)^2 + (12-12)^2 + (18-12)^2 + (24-12)^2 = 144 + 36 + 0 + 36 + 144 = 360\)

For \(\alpha = 0.05\) and 80% power (converting \(\delta_{\text{slope}}\) to per-month: 1.5/12 = 0.125 per month):

\[ n = \frac{2 \times 10.8 \times 7.85}{0.125^2 \times 360} = \frac{169.56}{5.625} = 30.1 \approx 31 \text{ per group} \]

With 80% power, only 31 subjects per group (62 total) are needed to detect a 1.5-point/year difference in the rate of decline. This is far fewer than would be needed for a simple endpoint comparison because the slope analysis uses all 5 measurements efficiently. However, the assumed within-subject correlation of \(\rho = 0.7\) must be verified from prior studies — if the true correlation is lower, more subjects will be needed.