Skip to the content

Topics Covered

Parallel Designs Cross-over Designs Cross-sectional vs. Longitudinal Objectives and Endpoints Phase I Trial Design Phase II Trial Design Phase III with Sequential Stopping Bioequivalence Trials
On this page
  1. 1. Parallel vs. Cross-over Designs
  2. 2. Cross-sectional vs. Longitudinal Designs
  3. 3. Objectives and Endpoints of Clinical Trials
  4. 4. Design of Phase I Trials
  5. 5. Design of Single-stage and Multi-stage Phase II Trials
  6. 6. Design and Monitoring of Phase III Trials with Sequential Stopping
  7. 7. Design of Bioequivalence Trials

1. Parallel vs. Cross-over Designs

KEY CONCEPT

Parallel and cross-over are the two fundamental design architectures for comparing treatments in clinical trials. In a parallel design, each subject receives exactly one treatment and groups are compared between subjects. In a cross-over design, each subject receives multiple treatments in different periods, serving as their own control. The choice between them profoundly affects sample size, analysis methods, and the interpretation of results.

1.1 Parallel Design

In a parallel (or between-subject) design, eligible subjects are randomised into two or more groups. Each group receives a different treatment throughout the entire study. The primary comparison is between subjects — outcomes from Group A are compared with outcomes from Group B.

PARALLEL DESIGN STRUCTURE
SubjectPeriod 1Period 2
Group A (n₁)Treatment ATreatment A
Group B (n₂)Treatment BTreatment B

Analysis model:

\[ Y_{ij} = \mu + \tau_i + \varepsilon_{ij}, \quad j = 1, \ldots, n_i, \quad i = A, B \]

where \(\tau_i\) is the treatment effect and \(\varepsilon_{ij} \sim N(0, \sigma^2)\).

Test statistic:

\[ t = \frac{\bar{Y}_A - \bar{Y}_B}{s_p \sqrt{\frac{1}{n_1} + \frac{1}{n_2}}} \]

with pooled variance \(s_p^2 = \frac{(n_1-1)s_1^2 + (n_2-1)s_2^2}{n_1 + n_2 - 2}\).

Advantages of parallel design:

Disadvantages:

EXAMPLE 1 — Parallel Design in a Hypertension Trial

A pharmaceutical company wishes to compare a new antihypertensive drug (Drug A) with the current standard (Drug B). One hundred and twenty patients with essential hypertension are randomised 1:1 into two groups of 60. Group A receives Drug A for 8 weeks; Group B receives Drug B for 8 weeks. The primary endpoint is the change in systolic blood pressure (SBP) from baseline to week 8.

After 8 weeks, the results are:

GroupnMean ΔSBP (mmHg)SD
Drug A60−14.28.5
Drug B60−10.59.1

Pooled variance: \(s_p^2 = \frac{59 \times 72.25 + 59 \times 82.81}{118} = 77.53\), so \(s_p = 8.81\).

Test statistic:

\[ t = \frac{-14.2 - (-10.5)}{8.81 \times \sqrt{\frac{1}{60} + \frac{1}{60}}} = \frac{-3.7}{8.81 \times 0.1826} = \frac{-3.7}{1.609} = -2.30 \]

With \(df = 118\), the two-sided p-value ≈ 0.023, which is statistically significant at \(\alpha = 0.05\). The new drug produces a significantly greater reduction in SBP compared to the standard. This parallel design is appropriate because blood pressure medication is taken chronically and a washout period would raise ethical concerns about untreated hypertension.

EXAMPLE 2 — Parallel Design in an Oncology Trial

A Phase III trial compares a new immunotherapy (Treatment A) with standard chemotherapy (Treatment B) in patients with advanced non-small cell lung cancer (NSCLC). Three hundred patients are randomised 1:1. The primary endpoint is overall survival (OS) — the time from randomisation to death from any cause.

Because the outcome is a time-to-event, the analysis uses a log-rank test and Cox proportional hazards model:

\[ h(t \mid X) = h_0(t) \exp(\beta \cdot \text{Treatment}) \]

Results after 24 months of follow-up: median OS = 16.3 months (Treatment A) vs. 12.8 months (Treatment B); hazard ratio = 0.74 (95% CI: 0.58–0.95), log-rank p = 0.018.

A parallel design is the only feasible option here because immunotherapy may permanently alter the immune system, making it impossible for a patient to "cross over" and receive the other treatment without carry-over effects. Moreover, the disease is life-threatening, so withholding effective treatment during a washout period is unethical. The parallel design with a survival endpoint provides an unbiased comparison of the two treatment strategies.

1.2 Cross-over Design

In a cross-over (or within-subject) design, each subject receives two or more treatments in different periods, separated by a washout period. The simplest and most common form is the 2×2 cross-over (also called AB/BA design): subjects are randomised into two sequences — AB (Treatment A in Period 1, Treatment B in Period 2) and BA (the reverse).

2×2 CROSS-OVER DESIGN STRUCTURE
SequencePeriod 1WashoutPeriod 2
Group 1 (AB)Treatment A—Treatment B
Group 2 (BA)Treatment B—Treatment A

Model:

\[ Y_{ijk} = \mu + \pi_j + \tau_{d(i,j)} + s_{ik} + \varepsilon_{ijk} \]

where:

Carry-over effect (\(\lambda\)): The residual effect of the treatment from Period 1 that persists into Period 2. If \(\lambda_A \neq \lambda_B\), the cross-over analysis is biased and only Period 1 data should be used (effectively reverting to a parallel design).

Test for treatment effect (assuming no carry-over):

Using the within-subject difference \(d_k = Y_{i1k} - Y_{i2k}\):

\[ \hat{\tau}_A - \hat{\tau}_B = \frac{\bar{d}_1 - \bar{d}_2}{2} \]

\[ t = \frac{\bar{d}_1 - \bar{d}_2}{\hat{\sigma}_d \sqrt{\frac{1}{n_1} + \frac{1}{n_2}}} \]

where \(\hat{\sigma}_d\) is the pooled SD of the within-subject differences.

Advantages of cross-over design:

Disadvantages:

EXAMPLE 1 — Cross-over Design for an Asthma Drug

A 2×2 cross-over trial compares a new inhaled bronchodilator (Treatment A) with a placebo (Treatment B) in patients with mild-to-moderate asthma. Twenty-four patients are randomised into two sequences (12 in AB, 12 in BA). Each treatment period is 2 weeks, with a 1-week washout between periods. The primary endpoint is the mean peak expiratory flow rate (PEFR, L/min) during each treatment period.

Results:

SequencenPeriod 1 MeanPeriod 2 MeanMean Diff (P1−P2)
AB12425 (A)395 (B)30
BA12388 (B)418 (A)−30

Treatment effect estimate:

\[ \hat{\tau}_A - \hat{\tau}_B = \frac{30 - (-30)}{2} = 30 \text{ L/min} \]

Suppose the pooled SD of differences \(\hat{\sigma}_d = 28\). Then:

\[ t = \frac{30 - (-30)}{28 \times \sqrt{\frac{1}{12} + \frac{1}{12}}} = \frac{60}{28 \times 0.4082} = \frac{60}{11.43} = 5.25 \]

With \(df = 22\), p < 0.001. The new bronchodilator significantly improves PEFR by 30 L/min compared to placebo. The cross-over design is appropriate here because asthma is a chronic stable condition, the drug has a short half-life (allowing a 1-week washout), and there is no carry-over concern. With only 24 subjects, the trial achieves high power because between-subject variability is removed from the comparison.

EXAMPLE 2 — Cross-over Design for an Analgesic

A 2×2 cross-over trial evaluates two analgesics — Drug A (new formulation) and Drug B (standard) — for post-surgical dental pain. Thirty patients who have undergone third molar extraction are randomised to Sequence AB or BA (15 each). Each patient takes the assigned drug for 6 hours post-surgery, then after a 3-day washout, takes the other drug following a second dental procedure on the contralateral side. The primary endpoint is the total pain relief score (TOTPAR) on a 0–100 visual analogue scale over the 6-hour period.

Results:

SequencenPeriod 1 TOTPARPeriod 2 TOTPARDiff (P1−P2)SD(diff)
AB1572 (A)61 (B)1114
BA1558 (B)68 (A)−1015

Treatment effect:

\[ \hat{\tau}_A - \hat{\tau}_B = \frac{11 - (-10)}{2} = 10.5 \text{ points} \]

Pooled SD of differences: \(\hat{\sigma}_d = \sqrt{\frac{14 \times 196 + 14 \times 225}{28}} = 14.53\)

\[ t = \frac{11 - (-10)}{14.53 \times \sqrt{\frac{1}{15} + \frac{1}{15}}} = \frac{21}{14.53 \times 0.3651} = \frac{21}{5.305} = 3.96 \]

With \(df = 28\), p < 0.001. Drug A provides significantly greater pain relief than Drug B. The cross-over design is ideal here because dental pain from third molar extraction is a well-characterised, short-duration model, and a 3-day washout is sufficient for either analgesic to be fully eliminated. A Grizzle test for carry-over yields \(p = 0.78\), confirming no significant carry-over effect, so the within-subject analysis is valid.

2. Cross-sectional vs. Longitudinal Designs

KEY CONCEPT

Cross-sectional designs observe or measure subjects at a single point in time, providing a "snapshot" of the population. Longitudinal designs follow the same subjects over multiple time points, enabling the study of change, trajectory, and temporal relationships. In clinical trials, longitudinal designs are predominant because they capture treatment effects over time, disease progression, and safety signals that emerge during follow-up.

2.1 Cross-sectional Design

In a cross-sectional study, each subject is measured once. The data represent the status of the population at one specific time. In the clinical trial context, a cross-sectional endpoint is one measured at a single pre-specified time point (e.g., blood pressure at week 8), even though the trial may run over a longer period.

CROSS-SECTIONAL ANALYSIS

Comparison at a single time point:

\[ Y_i = \beta_0 + \beta_1 \cdot \text{Treatment}_i + \varepsilon_i \]

Only one observation per subject. The estimate of treatment effect is:

\[ \hat{\beta}_1 = \bar{Y}_{\text{Treatment}} - \bar{Y}_{\text{Control}} \]

Advantages: Quick, inexpensive, no follow-up needed, no attrition.

Limitations: Cannot establish temporal ordering (causality), susceptible to cohort effects, cannot study individual change over time.

EXAMPLE 1 — Cross-sectional Survey of Diabetes Prevalence

A public health researcher conducts a cross-sectional survey of 2,000 adults aged 30–65 in a rural district to estimate the prevalence of type 2 diabetes and its association with body mass index (BMI). Each participant is assessed once: fasting blood glucose is measured, BMI is recorded, and a questionnaire captures demographic information.

Results: 320 out of 2,000 (16.0%) have diabetes (fasting glucose ≥ 126 mg/dL). The mean BMI is 28.3 kg/m² in diabetics vs. 25.1 kg/m² in non-diabetics. A chi-square test of association between BMI category (normal/overweight/obese) and diabetes status yields \(\chi^2 = 42.7\), \(df = 2\), \(p < 0.001\).

While the cross-sectional design clearly shows an association between higher BMI and diabetes, it cannot determine whether obesity precedes diabetes or vice versa. This snapshot is valuable for public health planning but insufficient for establishing a causal pathway.

EXAMPLE 2 — Cross-sectional Endpoint in a Vaccine Trial

In a Phase II vaccine trial, 150 healthy volunteers are randomised to receive either the experimental vaccine or a placebo. The primary endpoint is the seroconversion rate (proportion with antibody titre ≥ 1:40) measured at Day 28 post-vaccination. This is a cross-sectional endpoint because each subject is assessed at one time point.

Results: Seroconversion in 52/75 (69.3%) of vaccine recipients vs. 3/75 (4.0%) of placebo recipients. The difference in proportions is 0.653 (95% CI: 0.534–0.773), with a chi-square test yielding \(\chi^2 = 68.1\), \(p < 0.001\).

The cross-sectional measurement at Day 28 is sufficient for the primary immunogenicity endpoint. However, the duration of immunity cannot be assessed from this single time point — a longitudinal follow-up study would be needed to evaluate antibody persistence at 6 months, 1 year, and beyond.

2.2 Longitudinal Design

In a longitudinal study, the same subjects are measured repeatedly over time. This allows estimation of individual trajectories, the timing of events, and how treatment effects evolve. Longitudinal designs are fundamental to most Phase II and III clinical trials.

LONGITUDINAL ANALYSIS — MIXED EFFECTS MODEL

The general linear mixed model for longitudinal data:

\[ Y_{ij} = \beta_0 + \beta_1 \cdot \text{Time}_{ij} + \beta_2 \cdot \text{Treatment}_i + \beta_3 \cdot \text{Time}_{ij} \times \text{Treatment}_i + b_{0i} + b_{1i} \cdot \text{Time}_{ij} + \varepsilon_{ij} \]

where:

Covariance structures for the repeated measures:

EXAMPLE 1 — Longitudinal Design in a Rheumatoid Arthritis Trial

A 52-week randomised trial compares a new biologic (Treatment A) with methotrexate (Treatment B) in 200 patients with active rheumatoid arthritis. The primary endpoint is the DAS28 score (Disease Activity Score with 28-joint count), measured at baseline, week 4, week 12, week 24, and week 52. Lower DAS28 indicates less disease activity.

A linear mixed model is fitted:

\[ \text{DAS28}_{ij} = \beta_0 + \beta_1 \cdot \text{Week}_j + \beta_2 \cdot \text{Trt}_i + \beta_3 \cdot \text{Week}_j \times \text{Trt}_i + b_{0i} + \varepsilon_{ij} \]

Results (estimated coefficients):

ParameterEstimateSEp-value
Intercept (\(\beta_0\))5.420.12<0.001
Week (\(\beta_1\))−0.0120.0040.002
Treatment (\(\beta_2\))−0.180.170.29
Week × Treatment (\(\beta_3\))−0.0280.005<0.001

The significant Week × Treatment interaction (\(\beta_3 = -0.028\), p < 0.001) indicates that the rate of DAS28 improvement is significantly faster with Treatment A. By week 52, the estimated difference is \(0.028 \times 52 = 1.46\) DAS28 units — a clinically meaningful separation. The longitudinal design captures the evolving treatment effect that would be missed by a simple end-of-study comparison.

EXAMPLE 2 — Longitudinal Design in an HIV Treatment Trial

A clinical trial compares a new antiretroviral regimen (Regimen A) with the standard of care (Regimen B) in 180 treatment-naïve HIV-positive patients. The primary endpoint is the CD4+ T-cell count (cells/μL), measured at baseline, month 1, month 3, month 6, month 12, and month 24. Higher CD4 counts indicate better immune function.

Using a random intercept and slope model with AR(1) covariance:

\[ \text{CD4}_{ij} = \beta_0 + \beta_1 \cdot \text{Month}_j + \beta_2 \cdot \text{Trt}_i + \beta_3 \cdot \text{Month}_j \times \text{Trt}_i + b_{0i} + b_{1i} \cdot \text{Month}_j + \varepsilon_{ij} \]

Key results:

ParameterEstimate95% CIp-value
\(\beta_3\) (Month × Trt)3.21(1.84, 4.58)<0.001

The positive \(\beta_3 = 3.21\) means Regimen A increases CD4 count at a rate 3.21 cells/μL per month faster than Regimen B. By month 24, the estimated cumulative advantage is \(3.21 \times 24 = 77\) cells/μL — a substantial immunological benefit.

The longitudinal analysis is essential here because HIV treatment works gradually. A cross-sectional comparison at month 24 alone would miss early differences and the shape of the CD4 trajectory, and would be more vulnerable to missing data bias from dropouts. The mixed model uses all available data under the missing-at-random (MAR) assumption.

3. Objectives and Endpoints of Clinical Trials

KEY CONCEPT

The objective of a clinical trial is the specific question it is designed to answer (e.g., "Is Drug X superior to placebo in reducing HbA1c?"). The endpoint is the measurable outcome used to assess whether the objective has been met (e.g., "change in HbA1c from baseline to week 24"). A well-designed trial has a single primary objective with a corresponding primary endpoint, supported by secondary and exploratory objectives and endpoints.

3.1 Types of Objectives

CLASSIFICATION OF TRIAL OBJECTIVES
Objective TypeDefinitionStatistical FrameworkExample
SuperiorityDemonstrate that the experimental treatment is better than the control\(H_0: \tau_T - \tau_C \leq 0\) vs. \(H_1: \tau_T - \tau_C > 0\)New drug vs. placebo
Non-inferiorityDemonstrate that the experimental treatment is not unacceptably worse than the active control\(H_0: \tau_T - \tau_C \leq -\Delta\) vs. \(H_1: \tau_T - \tau_C > -\Delta\)New drug vs. standard (for convenience/safety)
EquivalenceDemonstrate that the treatments are similar within a pre-specified margin\(H_0: |\tau_T - \tau_C| \geq \Delta\) vs. \(H_1: |\tau_T - \tau_C| < \Delta\)Generic vs. brand-name drug
Dose-responseCharacterise the relationship between dose and effectTrend test, modellingMultiple dose levels vs. placebo

where \(\Delta\) is the non-inferiority/equivalence margin — the largest clinically acceptable difference that would still allow the conclusion that the treatments are "similar enough."

3.2 Types of Endpoints

CLASSIFICATION OF ENDPOINTS
Endpoint TypeDescriptionExample
Clinical endpointDirectly measures how a patient feels, functions, or survivesOverall survival, symptom relief
Surrogate endpointA biomarker intended to substitute for a clinical endpointCD4 count (for HIV), tumour shrinkage (for cancer)
Composite endpointCombination of multiple clinical eventsMACE (Major Adverse Cardiovascular Events)
DichotomousBinary outcome (success/failure)Response vs. no response
ContinuousMeasured on a continuous scaleBlood pressure, HbA1c, FEV₁
Time-to-eventTime until an event occursTime to disease progression, time to death
Repeated measuresMultiple assessments over timeWeekly pain scores over 12 weeks
REGULATORY GUIDANCE

The primary endpoint must be specified a priori in the protocol. FDA and EMA require that the primary endpoint be clinically meaningful or validated surrogate. The number of primary endpoints should be limited (preferably one) to control the Type I error rate. Multiple primary endpoints require multiplicity adjustment (e.g., Bonferroni, Hochberg, gatekeeping procedures).

EXAMPLE 1 — Superiority Objective with a Composite Endpoint

The EMPA-REG OUTCOME trial investigated whether empagliflozin (an SGLT2 inhibitor) reduces cardiovascular events in patients with type 2 diabetes and established cardiovascular disease. Primary objective: Demonstrate superiority of empagliflozin 10 mg vs. placebo in reducing the risk of Major Adverse Cardiovascular Events (MACE).

Primary endpoint: Composite of cardiovascular death, non-fatal myocardial infarction, or non-fatal stroke (3-point MACE), a time-to-event endpoint.

Statistical framework:

\[ H_0: \text{HR} \geq 1 \quad \text{vs.} \quad H_1: \text{HR} < 1 \]

Results: HR = 0.86 (95.02% CI: 0.74–0.99), p = 0.04 for superiority. The composite endpoint increased the event rate (8.2% vs. 5.9% for CV death alone), providing greater statistical power. However, the individual components were also examined: CV death was significantly reduced (HR = 0.62, p < 0.001), while MI and stroke were not — highlighting both the power advantage and interpretive challenge of composite endpoints.

EXAMPLE 2 — Non-inferiority Objective with a Continuous Endpoint

A new oral anticoagulant (Drug A) is compared with warfarin (Drug B) for stroke prevention in atrial fibrillation. Drug A has the advantage of not requiring INR monitoring, but must be shown to be not unacceptably worse than warfarin.

Primary objective: Demonstrate non-inferiority of Drug A vs. warfarin for stroke prevention.

Primary endpoint: Incidence of stroke or systemic embolism (dichotomous).

Non-inferiority margin: \(\Delta = 1.46\) (i.e., the upper bound of the 95% CI for the hazard ratio must be below 1.46).

Results: Event rate 1.27%/year (Drug A) vs. 1.40%/year (warfarin); HR = 0.91 (95% CI: 0.73–1.12). Since the upper bound of the CI (1.12) is below the non-inferiority margin (1.46), non-inferiority is established. Furthermore, the point estimate < 1 suggests possible superiority, which can be tested in a hierarchical testing procedure: first test non-inferiority, then if successful, test superiority.

This example illustrates the critical role of the non-inferiority margin. The margin \(\Delta\) is typically set as a fraction (usually 50%) of the estimated treatment effect of the active control over placebo from historical meta-analyses, preserving a clinically meaningful fraction of the active control's benefit.

4. Design of Phase I Trials

KEY CONCEPT

Phase I trials are the first stage of testing a new drug in humans. Their primary objective is to determine the maximum tolerated dose (MTD) — the highest dose that does not produce unacceptable toxicity — and to characterise the drug's safety profile, pharmacokinetics, and pharmacodynamics. Phase I trials typically involve 20–80 healthy volunteers (for non-cytotoxic drugs) or patients with advanced disease (for oncology drugs), and use dose escalation schemes to explore a range of doses safely.

4.1 Traditional 3+3 Design

The 3+3 design is the most widely used Phase I design in oncology. It escalates doses sequentially in cohorts of 3 patients, using the dose-limiting toxicity (DLT) rate to decide whether to escalate, de-escalate, or expand the cohort.

3+3 DOSE ESCALATION RULES

At each dose level, enrol 3 patients and observe DLTs during a specified window (typically Cycle 1, ~21 days):

DLTs out of 3Action
0/3Escalate to the next higher dose level
1/3Expand the cohort to 6 patients at the same dose level
≥2/3Dose exceeds MTD; de-escalate. The previous dose level is the MTD.

Expansion rule (after enrolling 3 more):

DLTs out of 6Action
≤1/6Escalate to the next higher dose level
≥2/6Dose exceeds MTD; de-escalate

The target DLT rate is typically 1/6 ≈ 17% (range 16–33% depending on the modification). The MTD is the highest dose at which ≤1/6 patients experience DLTs.

4.2 Continual Reassessment Method (CRM)

The CRM is a model-based Bayesian dose-finding design. It uses a prespecified dose-toxicity model, updates the estimate after each patient (or cohort), and assigns the next patient to the dose whose estimated probability of DLT is closest to the target \(\theta\).

CRM MODEL

Dose-toxicity model (power model):

\[ p(d_i) = d_i^{\exp(\alpha)}, \quad i = 1, 2, \ldots, k \]

where \(d_i\) are "skeleton" probabilities (pre-specified prior guesses of DLT rates at each dose level) and \(\alpha\) is the unknown parameter to be estimated.

After observing data from n patients: The posterior distribution of \(\alpha\) is updated using Bayes' theorem:

\[ f(\alpha \mid \text{data}) \propto L(\text{data} \mid \alpha) \cdot \pi(\alpha) \]

where \(L\) is the likelihood and \(\pi(\alpha)\) is the prior (typically \(N(0, \sigma^2_\alpha)\)).

Next dose assignment:

\[ d^* = \arg\min_{d_i} |E[p(d_i) \mid \text{data}] - \theta| \]

where \(\theta\) is the target DLT rate (e.g., \(\theta = 0.25\)).

Safety constraint: No dose escalation by more than one level at a time.

EXAMPLE 1 — 3+3 Design for a Chemotherapy Drug

A Phase I trial evaluates a new cytotoxic agent at 5 dose levels: 50, 100, 200, 400, and 600 mg/m². The target DLT rate is ≤1/6 (≈17%).

Escalation process:

StepDose LevelPatientsDLTsAction
150 mg/m²30/3Escalate
2100 mg/m²30/3Escalate
3200 mg/m²31/3Expand to 6
4200 mg/m²+3 = 61/6 totalEscalate
5400 mg/m²32/3Exceeds MTD → De-escalate

Conclusion: Since 400 mg/m² produced DLTs in ≥2/3 patients, it exceeds the MTD. The previous dose level, 200 mg/m², is declared the MTD (with 1/6 DLT rate). This dose is recommended for Phase II testing. A total of 15 patients were enrolled.

EXAMPLE 2 — CRM Design for a Targeted Therapy

A Phase I trial uses the CRM to find the MTD of a new targeted therapy at 6 dose levels. The target DLT rate is \(\theta = 0.25\). Skeleton probabilities are \(d = (0.05, 0.12, 0.20, 0.30, 0.42, 0.55)\), and the prior for \(\alpha\) is \(N(0, 1.34^2)\).

Step-by-step CRM operation:

Patient 1: Assigned to dose level 3 (the level whose prior DLT probability is closest to \(\theta = 0.25\), since prior \(E[\alpha] = 0\) gives \(p(d_3) = 0.20\)). The patient does not experience a DLT.

Update: The posterior mean of \(\alpha\) shifts downward (the drug appears less toxic than the prior at level 3). The estimated DLT rates are:

Dose Level\(\hat{p}(d_i)\)
10.01
20.04
30.10
40.18
50.29
60.41

The dose closest to \(\theta = 0.25\) is now level 5 (\(\hat{p} = 0.29\)), but the safety constraint limits escalation to one level, so Patient 2 is assigned to dose level 4.

Patient 2: Assigned to dose level 4. Experiences a DLT.

Update: The posterior shifts upward. Estimated \(\hat{p}(d_4) = 0.28\), which is close to \(\theta = 0.25\). Patient 3 is assigned to dose level 4 again. This patient does not experience a DLT, so \(\hat{p}(d_4)\) adjusts back toward 0.25.

After 12 patients, the CRM converges: dose level 4 is the MTD with posterior mean DLT rate \(\hat{p} = 0.26\), closest to \(\theta = 0.25\). The CRM is more efficient than the 3+3 design because it uses all accumulated data to guide dose assignment and treats more patients near the MTD rather than at subtherapeutic doses.

5. Design of Single-stage and Multi-stage Phase II Trials

KEY CONCEPT

Phase II trials evaluate whether a new treatment has sufficient biological activity (efficacy signal) to warrant a definitive Phase III trial. They typically use tumour response rate (in oncology) or another binary endpoint. The key design question is: "How many responses are needed to conclude that the treatment is promising enough for further study?" Single-stage designs test once at the end; multi-stage designs allow early stopping for futility (and sometimes efficacy), saving patients and resources.

5.1 Single-stage Design (A'Hern)

The single-stage (or one-stage) design enrols all \(n\) patients, observes the number of responses \(r\), and compares \(r\) to a critical value \(r_0\). If \(r > r_0\), the treatment is declared promising.

SINGLE-STAGE DESIGN

Given:

Hypotheses:

\[ H_0: p \leq p_0 \quad \text{vs.} \quad H_1: p \geq p_1 \]

Decision rule: Reject \(H_0\) if the number of responses \(r > r_0\).

The values of \(n\) and \(r_0\) are chosen to satisfy:

\[ P(r > r_0 \mid p = p_0) \leq \alpha \quad \text{and} \quad P(r > r_0 \mid p = p_1) \geq 1 - \beta \]

using the exact binomial distribution:

\[ P(r > r_0 \mid p) = \sum_{k=r_0+1}^{n} \binom{n}{k} p^k (1-p)^{n-k} \]

5.2 Two-stage Design (Fleming / Simon Optimal / Simon Minimax)

Two-stage designs allow early termination if the treatment shows insufficient activity after the first stage, sparing patients from receiving an ineffective treatment.

SIMON TWO-STAGE DESIGN

Stage 1: Enrol \(n_1\) patients. If the number of responses \(r_1 \leq r_{1,c}\) (the futility boundary), stop the trial and declare the treatment unpromising. Otherwise, proceed to Stage 2.

Stage 2: Enrol an additional \(n_2\) patients. If the total number of responses \(r_1 + r_2 > r_{2,c}\), declare the treatment promising.

Error rates:

\[ P(\text{reject } H_0 \mid p = p_0) \leq \alpha \]

\[ P(\text{reject } H_0 \mid p = p_1) \geq 1 - \beta \]

Simon Optimal: Minimises the expected sample size under \(H_0\) (i.e., when the drug is inactive), saving the most patients on average when the treatment doesn't work.

Simon Minimax: Minimises the maximum total sample size \(n = n_1 + n_2\), useful when patient accrual is limited.

Expected sample size under \(H_0\):

\[ E[N \mid p_0] = n_1 + (1 - PET(p_0)) \times n_2 \]

where \(PET(p_0) = P(r_1 \leq r_{1,c} \mid p_0)\) is the probability of early termination under the null.

EXAMPLE 1 — Single-stage Phase II Design

A Phase II trial evaluates a new targeted agent in patients with relapsed non-Hodgkin lymphoma. The standard therapy has a response rate of approximately 20%. The investigators consider the new agent promising if the response rate is at least 40%.

Design parameters: \(p_0 = 0.20\), \(p_1 = 0.40\), \(\alpha = 0.05\), \(\beta = 0.20\) (80% power).

Using exact binomial calculations:

Try \(n = 35\), \(r_0 = 11\):

\[ P(r > 11 \mid p = 0.20) = \sum_{k=12}^{35} \binom{35}{k} (0.20)^k (0.80)^{35-k} = 0.034 \leq 0.05 \checkmark \]

\[ P(r > 11 \mid p = 0.40) = \sum_{k=12}^{35} \binom{35}{k} (0.40)^k (0.60)^{35-k} = 0.805 \geq 0.80 \checkmark \]

(Note: the smaller threshold \(r_0 = 10\) would give \(P(r>10\mid p=0.20) = 0.075 > 0.05\), failing the Type I error constraint — hence \(r_0 = 11\) is required.)

Design: Enrol 35 patients. If more than 11 responses are observed (i.e., ≥12), declare the treatment promising. If 11 or fewer responses are observed, the treatment is not pursued further.

Suppose the trial observes 14 responses out of 35 (40%). Since 14 > 11, the treatment is declared promising and a Phase III trial is warranted. The 95% confidence interval for the response rate is (24.5%, 57.3%), which includes \(p_1 = 0.40\) and excludes \(p_0 = 0.20\), consistent with the design conclusion.

EXAMPLE 2 — Simon Two-stage Optimal Design

A Phase II trial evaluates a new combination chemotherapy in patients with advanced pancreatic cancer. The historical response rate is \(p_0 = 0.10\). A response rate of \(p_1 = 0.30\) would be considered promising. Design parameters: \(\alpha = 0.05\), \(\beta = 0.20\).

Simon Optimal two-stage design:

Operating characteristics:

Trial conduct: In Stage 1, 3 out of 15 patients respond. Since 3 > 2, the trial continues to Stage 2. After 40 patients, 11 responses are observed. Since 11 ≥ 9, the treatment is declared promising. The two-stage design saved patients on average (expected 19.6 under \(H_0\) vs. 40 always in a single-stage design), while the early stopping rule protects patients from an ineffective treatment 81.6% of the time when the drug truly has only a 10% response rate.

6. Design and Monitoring of Phase III Trials with Sequential Stopping

KEY CONCEPT

Phase III trials are large, confirmatory studies comparing a new treatment with the standard of care. Because they enrol hundreds or thousands of patients over several years, it is ethically imperative to monitor the accumulating data for early evidence of benefit or harm. Sequential stopping (or group sequential) methods provide statistical boundaries that define when a trial may be stopped early while controlling the overall Type I error rate at the desired level (e.g., \(\alpha = 0.05\)).

6.1 Group Sequential Methods

In a group sequential design, the data are analysed at pre-specified interim analyses (after groups of patients have completed follow-up). At each analysis, the test statistic is compared to stopping boundaries. If the statistic crosses the boundary, the trial may be stopped early.

GROUP SEQUENTIAL FRAMEWORK

Notation:

Decision rules at analysis \(k\):

Overall Type I error:

\[ P(\text{reject } H_0 \text{ at any analysis} \mid H_0) = \alpha \]

The boundaries must be chosen to ensure the overall \(\alpha\) is maintained despite multiple looks at the data.

6.2 Common Boundary Methods

THREE CLASSICAL BOUNDARY METHODS

1. Pocock boundary: Uses the same critical value at each analysis (\(u_1 = u_2 = \cdots = u_K\)), but this value is larger than the fixed-sample \(z_{\alpha/2}\) to account for multiple testing. Simple but may be too aggressive in early stopping.

Analysis1234
Pocock \(u_k\) (K=4, \(\alpha=0.05\))2.3612.3612.3612.361

2. O'Brien-Fleming boundary: Uses very stringent boundaries at early analyses and relaxes toward the fixed-sample critical value at the final analysis. Very conservative early — requires overwhelming evidence to stop early. Most widely used in practice.

Analysis1234
O'Brien-Fleming \(u_k\) (K=4, \(\alpha=0.05\))4.0492.8632.3382.024

3. Haybittle-Peto boundary: Uses a very conservative critical value (e.g., \(z = 3.0\), p < 0.001) at interim analyses and the standard \(z = 1.96\) at the final analysis. Simple, but the overall \(\alpha\) is only approximately controlled.

Information fraction: \(t_k = n_k / n_K\), the proportion of total information accumulated at analysis \(k\). Analyses are typically scheduled at equal information fractions (e.g., \(t_k = 0.25, 0.50, 0.75, 1.00\) for K = 4).

6.3 Data Monitoring Committee (DMC)

A Data Monitoring Committee (also called Data and Safety Monitoring Board, DSMB) is an independent group of experts that reviews the interim analyses and makes recommendations about trial continuation, modification, or termination. The DMC operates under a charter that specifies the stopping guidelines, analysis plans, and confidentiality procedures.

EXAMPLE 1 — O'Brien-Fleming Boundary in a Cardiovascular Trial

A Phase III trial compares a new statin with a standard statin in 10,000 patients with hypercholesterolemia. The primary endpoint is the incidence of major cardiovascular events (MCE) over 5 years. Three interim analyses are planned at information fractions \(t = 0.33, 0.67, 1.00\) (i.e., \(K = 3\) analyses total), with an overall two-sided \(\alpha = 0.05\).

O'Brien-Fleming boundaries:

AnalysisInformation FractionEvents\(u_k\) (two-sided)Equivalent p-value
Interim 10.332003.471<0.001
Interim 20.674002.4540.014
Final1.006002.0040.045

Trial conduct: At Interim 1 (200 events), the observed z-statistic is 1.85. Since 1.85 < 3.471, the trial continues. At Interim 2 (400 events), the observed z-statistic is 2.72. Since 2.72 > 2.454, the boundary is crossed and the DMC recommends stopping the trial for efficacy. The conditional power at this point exceeds 99%, confirming that continuing the trial would almost certainly reach the same conclusion.

The O'Brien-Fleming design required overwhelming evidence (\(p < 0.001\)) at the first interim to stop, protecting against premature stopping on unreliable early data. By the second analysis with 400 events, the evidence was compelling enough that continuing would have been unethical. The overall \(\alpha\) expenditure across the three analyses is controlled at 0.05.

EXAMPLE 2 — Pocock Boundary with Futility Stopping in an Oncology Trial

A Phase III trial compares a new immunotherapy combination (Treatment A) with standard immunotherapy alone (Treatment B) in 600 patients with metastatic melanoma. The primary endpoint is progression-free survival (PFS). Two interim analyses are planned at \(t = 0.50\) and \(t = 1.00\) (\(K = 2\)), with overall two-sided \(\alpha = 0.05\).

Pocock boundaries with symmetric efficacy/futility stopping:

AnalysisEventsEfficacy \(u_k\)Futility \(l_k\)
Interim 11802.178−2.178
Final3602.178−2.178

Trial conduct: At the interim analysis (180 PFS events), the log-rank test yields \(Z = 1.50\). Since \(-2.178 < 1.50 < 2.178\), the trial continues — there is neither sufficient evidence of efficacy nor of futility.

At the final analysis (360 events), the log-rank test yields \(Z = 2.45\). Since \(2.45 > 2.178\), the efficacy boundary is crossed. The estimated hazard ratio is 0.78 (95% CI: 0.63–0.96), indicating a 22% reduction in the risk of progression with the combination.

The Pocock boundary's constant critical value (2.178 at both analyses) makes early stopping easier than O'Brien-Fleming if the effect is large and consistent. However, the trade-off is a higher critical value at the final analysis compared to the fixed-sample \(z = 1.96\), which slightly reduces power. In this case, the trial ran to completion, and the treatment benefit was confirmed at the final analysis.

7. Design of Bioequivalence Trials

KEY CONCEPT

Bioequivalence (BE) trials demonstrate that two pharmaceutical products (typically a generic and a brand-name drug, or two formulations of the same drug) have similar bioavailability — meaning they reach the systemic circulation at the same rate and to the same extent. Unlike superiority trials, BE trials aim to show that the treatments are not meaningfully different, using an equivalence testing framework. The standard criterion is that the 90% confidence interval for the ratio of geometric means (test/reference) must fall within the acceptance range of 80–125%.

7.1 Pharmacokinetic Parameters

KEY PK PARAMETERS FOR BIOEQUIVALENCE

\[ \text{AUC}_{0-\infty} = \text{AUC}_{0-t} + \frac{C_{\text{last}}}{\lambda_z} \]

where \(\lambda_z\) is the terminal elimination rate constant estimated from the log-linear portion of the concentration–time curve.

7.2 Standard Bioequivalence Design

2×2 CROSS-OVER BIOEQUIVALENCE DESIGN

The standard BE study uses a randomised, two-period, two-sequence cross-over design:

SequencePeriod 1WashoutPeriod 2
RT (Group 1)Reference (R)≥ 5 half-livesTest (T)
TR (Group 2)Test (T)≥ 5 half-livesReference (R)

Washout period: Must be at least 5 elimination half-lives to ensure <5% of the drug remains from Period 1.

Log-transformation: PK parameters (\(C_{\max}\), AUC) are analysed on the natural log scale because they are typically log-normally distributed:

\[ Y = \ln(\text{PK parameter}) \]

ANOVA model (on log scale):

\[ Y_{ijk} = \mu + \pi_j + \tau_{d(i,j)} + s_{ik} + \varepsilon_{ijk} \]

where \(\pi_j\) = period effect, \(\tau_{d(i,j)}\) = formulation effect (T or R), \(s_{ik}\) = subject random effect, \(\varepsilon_{ijk}\) = error.

Point estimate of the ratio (on original scale):

\[ \widehat{T/R} = \exp(\bar{Y}_T - \bar{Y}_R) \]

90% confidence interval:

\[ \exp\left[(\bar{Y}_T - \bar{Y}_R) \pm t_{\alpha/2, \, df} \times SE(\bar{Y}_T - \bar{Y}_R)\right] \]

Bioequivalence criterion:

\[ 0.80 \leq \text{lower 90% CI bound} \quad \text{and} \quad \text{upper 90% CI bound} \leq 1.25 \]

This is equivalent to testing two one-sided hypotheses (TOST procedure):

\[ H_{01}: \mu_T - \mu_R \leq \ln(0.80) \quad \text{vs.} \quad H_{11}: \mu_T - \mu_R > \ln(0.80) \]

\[ H_{02}: \mu_T - \mu_R \geq \ln(1.25) \quad \text{vs.} \quad H_{12}: \mu_T - \mu_R < \ln(1.25) \]

Both one-sided tests must be rejected at \(\alpha = 0.05\) (which together form the 90% CI).

REGULATORY NOTE

The 80–125% criterion is asymmetric on the original scale but symmetric on the log scale: \(\ln(0.80) = -0.2231\) and \(\ln(1.25) = 0.2231\). For narrow therapeutic index drugs (e.g., warfarin, digoxin), some regulators require a tighter range of 90–111%. For highly variable drugs (intra-subject CV > 30%), scaled average bioequivalence may be applied with widened limits.

7.3 Sample Size for Bioequivalence

SAMPLE SIZE FORMULA FOR BIOEQUIVALENCE

For a 2×2 cross-over BE study with the TOST procedure:

\[ n \geq \frac{(t_{\alpha, \, 2n-2} + t_{\beta/2, \, 2n-2})^2 \times \sigma_W^2} {(\ln(1.25) - |\ln(\theta)|)^2} \]

where:

This is solved iteratively because the t-distribution quantiles depend on \(n\).

EXAMPLE 1 — Bioequivalence Trial for a Generic Metformin Tablet

A pharmaceutical company develops a generic 500 mg metformin tablet and must demonstrate bioequivalence with the brand-name product (Glucophage®). A 2×2 cross-over study is designed with 24 healthy volunteers (12 in RT, 12 in TR). Each subject receives a single 500 mg dose in each period, with a 7-day washout (metformin half-life ≈ 6 hours, so 5 × 6 = 30 hours washout is sufficient; 7 days provides ample margin).

Blood samples are collected at 0, 0.5, 1, 1.5, 2, 3, 4, 6, 8, 12, and 24 hours post-dose in each period. The key PK parameters are calculated:

ParameterTest (Generic) MeanReference (Brand) Mean
\(C_{\max}\) (ng/mL)11681205
\(\text{AUC}_{0-t}\) (ng·h/mL)78008010
\(T_{\max}\) (h)2.52.3

Statistical analysis (on ln scale):

For \(C_{\max}\): Point estimate T/R = 1168/1205 = 0.969 (96.9%).

On log scale: \(\ln(0.969) = -0.0315\). The 90% CI on the log scale is \((-0.1453, 0.0823)\).

Exponentiating: 90% CI = \((\exp(-0.1453), \exp(0.0823)) = (0.865, 1.086)\) or (86.5%, 108.6%).

Since 80% ≤ 86.5% and 108.6% ≤ 125%, bioequivalence is established for \(C_{\max}\).

For \(\text{AUC}_{0-t}\): Point estimate T/R = 7800/8010 = 0.974 (97.4%).

90% CI on original scale: (87.2%, 108.9%). Since this falls entirely within 80–125%, bioequivalence is established for AUC.

Since both \(C_{\max}\) and AUC meet the bioequivalence criteria, the generic metformin tablet is declared bioequivalent to the brand-name product and may be approved for marketing.

EXAMPLE 2 — Bioequivalence Failure for a Modified-Release Formulation

A company develops a modified-release (MR) formulation of a drug and must show bioequivalence with the immediate-release (IR) reference. The MR formulation is designed to release the drug slowly over 12 hours. A 2×2 cross-over study enrolls 30 healthy volunteers.

Results:

ParameterTest (MR) MeanReference (IR) MeanT/R Ratio90% CI
\(C_{\max}\) (ng/mL)45820.549(44.8%, 67.2%)
\(\text{AUC}_{0-t}\) (ng·h/mL)5104901.041(95.3%, 113.8%)

Analysis:

Conclusion: Bioequivalence is not established. The MR formulation delivers the same total amount of drug (equivalent AUC) but at a significantly lower peak concentration. This is expected for a modified-release product — the drug is absorbed more slowly, producing a lower \(C_{\max}\) and later \(T_{\max}\). A standard BE comparison is inappropriate here; instead, the MR formulation should be evaluated on its own merits, perhaps using a different acceptance criterion or a clinical endpoint study to demonstrate that the lower \(C_{\max}\) does not compromise efficacy while potentially improving the safety profile (fewer peak-related side effects).

This example underscores that the 80–125% criterion is designed for products that are intended to be pharmaceutically equivalent. When the formulation is intentionally different (e.g., MR vs. IR), a standard BE trial may fail even though the new formulation is clinically acceptable for its intended purpose.