A surrogate endpoint (or surrogate biomarker) is a biomarker that is intended to substitute for a clinical endpoint. A surrogate is expected to predict clinical benefit (or harm, or lack of benefit) based on epidemiologic, therapeutic, pathophysiologic, or other scientific evidence. Surrogate endpoints can dramatically reduce the time and cost of clinical trials, but their use is valid only when they have been rigorously demonstrated to predict the true clinical outcome — a requirement that is often difficult to satisfy.
| Reason | Explanation | Example |
|---|---|---|
| Practical efficiency | Clinical endpoints may take years to observe (e.g., overall survival in early-stage cancer) | Tumour shrinkage (response rate) instead of survival |
| Rarity of events | The clinical event may be too rare to study feasibly | Fracture rate in osteoporosis → bone mineral density |
| Ethical considerations | It may be unethical to wait for clinical outcomes when surrogate data are available | CD4 count in HIV → delay in using effective therapy |
| Regulatory acceleration | Regulators may grant accelerated approval based on surrogate data | HbA1c in diabetes (FDA-accepted surrogate) |
| Screening and dose-finding | Phase I/II trials use surrogates to screen for activity and find the right dose | Pharmacokinetic parameters (AUC, Cmax) |
Not every biomarker qualifies as a valid surrogate. The classic criteria were articulated by Prentice (1989) and further developed by Freedman, Graubard, and Schatzkin (1992).
A biomarker \(S\) is a valid surrogate for the true clinical endpoint \(T\) if all four conditions hold:
Formally, condition 4 requires:
\[ P(T \mid Z, S) = P(T \mid S) \]
or equivalently, in a regression model for \(T\):
\[ E[T \mid Z, S] = E[T \mid S] \]
i.e., the coefficient of \(Z\) is zero after adjusting for \(S\). This is a very stringent requirement — the surrogate must fully mediate the treatment effect.
Freedman et al. (1992) proposed a less stringent, more practical measure: the proportion of the treatment effect explained by the surrogate:
\[ PTE = 1 - \frac{\beta_{Z \mid S}}{\beta_Z} \]
where:
Interpretation:
Freedman recommended that PTE should be at least 0.50–0.75 for a surrogate to be considered adequate. However, PTE is unstable with small samples and can be outside [0, 1] due to sampling variability.
1. Endpoint selection: Choose a surrogate with the strongest available evidence of validity (from prior randomised trials, meta-analyses, and biological understanding).
2. Sample size implications:
If the surrogate allows a smaller effect size \(\delta_S\) to be detected (compared to the clinical endpoint effect \(\delta_T\)), and if the surrogate has lower variability:
\[ n_{\text{surrogate}} = \frac{(z_{\alpha/2} + z_\beta)^2 \times 2\sigma_S^2}{\delta_S^2} \]
Sample size reduction factor:
\[ \frac{n_{\text{surrogate}}}{n_{\text{clinical}}} = \frac{\sigma_S^2 / \delta_S^2}{\sigma_T^2 / \delta_T^2} \]
If the surrogate has higher event rates and lower variance, this ratio can be < 0.5, meaning fewer than half the subjects are needed.
3. Timing of assessment: Surrogates are often assessed earlier than clinical endpoints, which changes the study duration and the risk of missing late-emerging effects.
4. Regulatory classification:
| Category | Definition | Example |
|---|---|---|
| Validated surrogate | Regulators accept it as a basis for full approval | Blood pressure (for stroke/heart disease), HbA1c (for diabetic complications) |
| Reasonably likely surrogate | Supports accelerated approval; confirmatory trial still required | Tumour response rate (for some cancers) |
| Candidate surrogate | Used in Phase I/II for screening; not accepted for approval | Most novel biomarkers |
A clinical trial evaluates a new antiretroviral regimen in treatment-naïve HIV patients. The traditional clinical endpoint is a composite of AIDS-defining events or death, which may take years to occur. Instead, the investigators use CD4+ T-cell count and HIV RNA viral load as surrogate endpoints.
Validation evidence:
Design: Three hundred patients are randomised 1:1 to the new regimen or the standard. The primary surrogate endpoint is the proportion with undetectable viral load (< 50 copies/mL) at week 48. A secondary surrogate endpoint is the change in CD4 count at week 48.
Results:
| Endpoint | New Regimen | Standard | Effect | p-value |
|---|---|---|---|---|
| Viral load < 50 copies/mL | 78/150 (52%) | 55/150 (36.7%) | OR = 1.84 | 0.009 |
| CD4 change (cells/μL) | +185 | +142 | Δ = 43 | 0.014 |
The new regimen produces significantly better surrogate outcomes. Based on the validated surrogate relationship, these results are expected to translate into fewer AIDS events and deaths in the long term. The FDA accepts these surrogate data for full approval in HIV, given the extensive validation. However, the trial protocol includes a 2-year follow-up for clinical events to provide additional confirmatory evidence.
A Phase III trial tests a new drug that shrinks tumours in advanced non-small cell lung cancer. The surrogate endpoint is objective response rate (ORR) — the proportion of patients with a complete or partial response by RECIST criteria. The true clinical endpoint is overall survival (OS).
Rationale for the surrogate: Tumour shrinkage is biologically plausible as a predictor of survival, and several cancer types (e.g., breast, lymphoma) show a strong ORR–OS correlation. However, in NSCLC, the evidence is mixed.
Trial results:
| Endpoint | New Drug | Standard | Effect | p-value |
|---|---|---|---|---|
| ORR (surrogate) | 35% (52/150) | 20% (30/150) | OR = 2.17 | 0.003 |
| Median OS (clinical) | 9.8 months | 10.2 months | HR = 1.05 | 0.71 |
The paradox: The new drug significantly improves the surrogate (tumour response rate) but provides no survival benefit. Possible explanations:
Lesson: ORR is a poor surrogate for OS in NSCLC. The Prentice criterion 4 is violated — the treatment effect on OS is not fully captured by ORR. This example demonstrates the danger of relying on an inadequately validated surrogate: a trial could conclude the drug is effective (based on the surrogate) when patients derive no meaningful clinical benefit. After several such examples, oncology trials now increasingly use progression-free survival (PFS) or OS as primary endpoints, with ORR as a secondary endpoint only.
Analysing surrogate endpoint data requires methods that go beyond simply testing whether the treatment affects the surrogate. The analyst must evaluate the strength of the surrogate–clinical endpoint association, the proportion of treatment effect explained, and the predictive value of the surrogate for the clinical outcome. Multi-trial (meta-analytic) approaches are now considered the gold standard for surrogate validation.
Stage 1: Fit a model for the clinical endpoint \(T\) on the treatment \(Z\) alone:
\[ T_i = \alpha_0 + \alpha_1 Z_i + \varepsilon_{1i} \]
\(\hat{\alpha}_1\) is the total treatment effect on the clinical endpoint.
Stage 2: Fit a model for the clinical endpoint \(T\) on both the surrogate \(S\) and the treatment \(Z\):
\[ T_i = \gamma_0 + \gamma_1 Z_i + \gamma_2 S_i + \varepsilon_{2i} \]
\(\hat{\gamma}_1\) is the direct treatment effect (after adjusting for the surrogate). \(\hat{\gamma}_2\) measures the surrogate–clinical endpoint association.
Proportion of treatment effect explained:
\[ PTE = 1 - \frac{\hat{\gamma}_1}{\hat{\alpha}_1} \]
Adjusted association: The coefficient \(\hat{\gamma}_2\) (and its significance) quantifies the association between \(S\) and \(T\) after adjusting for treatment.
The single-trial approach has limited power and generalisability. The modern approach, proposed by Buyse et al. (2000) and Gail et al. (2000), uses data from multiple randomised trials to evaluate the surrogate at the trial level.
Individual-level model (within trial \(i\)):
\[ S_{ij} = \mu_{Si} + \alpha_i Z_{ij} + \varepsilon_{Sij} \]
\[ T_{ij} = \mu_{Ti} + \beta_i Z_{ij} + \varepsilon_{Tij} \]
where \(\alpha_i\) = treatment effect on the surrogate in trial \(i\), and \(\beta_i\) = treatment effect on the true endpoint in trial \(i\).
Trial-level model (between trials):
\[ \begin{pmatrix} \mu_{Si} \\ \mu_{Ti} \end{pmatrix} \sim N \left( \begin{pmatrix} \mu_S \\ \mu_T \end{pmatrix}, \Sigma_\mu \right) \]
\[ \begin{pmatrix} \alpha_i \\ \beta_i \end{pmatrix} \sim N \left( \begin{pmatrix} \alpha \\ \beta \end{pmatrix}, \Sigma_\tau \right) \]
Trial-level surrogacy is measured by the coefficient of determination between the trial-specific treatment effects on the surrogate (\(\alpha_i\)) and on the true endpoint (\(\beta_i\)):
\[ R^2_{\text{trial}} = \frac{\sigma^2_{\alpha\beta}}{\sigma^2_{\alpha\alpha} \times \sigma^2_{\beta\beta}} \]
where \(\Sigma_\tau = \begin{pmatrix} \sigma^2_{\alpha\alpha} & \sigma^2_{\alpha\beta} \\ \sigma^2_{\alpha\beta} & \sigma^2_{\beta\beta} \end{pmatrix}\).
\(R^2_{\text{trial}}\) close to 1 indicates that the treatment effect on the surrogate nearly perfectly predicts the treatment effect on the clinical endpoint — the surrogate is valid at the trial level.
Individual-level surrogacy is measured by the correlation between \(S\) and \(T\) after adjusting for treatment:
\[ R^2_{\text{indiv}} = \frac{\sigma^2_{ST}}{\sigma^2_S \times \sigma^2_T} \]
where \(\Sigma_\varepsilon = \begin{pmatrix} \sigma^2_S & \sigma^2_{ST} \\ \sigma^2_{ST} & \sigma^2_T \end{pmatrix}\).
Both \(R^2_{\text{trial}}\) and \(R^2_{\text{indiv}}\) should be high for a surrogate to be considered valid.
A meta-analysis of 12 randomised trials (total n = 48,000 patients) evaluates whether the treatment effect on systolic blood pressure (SBP) reduction is a valid surrogate for the treatment effect on cardiovascular events (stroke, myocardial infarction, or cardiovascular death).
Multi-trial analysis results:
At the trial level, the treatment effects on SBP (\(\alpha_i\)) and on CV events (\(\beta_i\)) are estimated for each trial. The linear regression of \(\hat{\beta}_i\) on \(\hat{\alpha}_i\) shows:
\[ \hat{\beta}_i = 0.025 + 0.012 \times \hat{\alpha}_i \quad (R^2_{\text{trial}} = 0.72) \]
At the individual level:
\[ R^2_{\text{indiv}} = 0.31 \]
Interpretation:
Based on this and similar meta-analyses, the FDA accepts SBP reduction as a validated surrogate for accelerated approval of antihypertensive drugs, while requiring post-marketing studies to confirm CV benefit. The strong trial-level surrogacy (\(R^2_{\text{trial}} = 0.72\)) supports this regulatory approach.
A meta-analysis of 8 Phase III randomised trials (total n = 6,200 patients) evaluates whether progression-free survival (PFS) is a valid surrogate for overall survival (OS) in advanced colorectal cancer treated with first-line chemotherapy.
Multi-trial analysis (both endpoints are time-to-event):
For each trial, the log hazard ratios for PFS (\(\hat{\lambda}_i^{PFS}\)) and OS (\(\hat{\lambda}_i^{OS}\)) are estimated using Cox models. The trial-level association is:
\[ \hat{\lambda}_i^{OS} = 0.04 + 0.72 \times \hat{\lambda}_i^{PFS} \]
\[ R^2_{\text{trial}} = 0.48 \quad (95\% \text{ CI: } 0.18, 0.78) \]
The individual-level Kendall's τ between PFS and OS = 0.65 (strong individual-level association).
Interpretation:
Practical implication: PFS can be used as a primary endpoint in Phase II trials (screening for activity) but should be interpreted cautiously in Phase III. Regulatory authorities in this setting would likely require OS data or at least a post-marketing commitment to confirm the OS benefit. This analysis highlights that surrogate validation is disease-specific — PFS may be a stronger surrogate for OS in some cancers (e.g., ovarian) than others (e.g., colorectal, NSCLC).
Meta-analysis is the statistical synthesis of results from multiple independent studies addressing the same research question. In clinical trials, meta-analysis combines evidence from several randomised controlled trials to obtain a more precise estimate of the treatment effect, resolve uncertainty when individual trials are inconclusive, assess consistency (heterogeneity) of effects across studies, and identify patient subgroups that may benefit differentially. When properly conducted, a meta-analysis provides the highest level of evidence in the evidence hierarchy.
The fixed-effect model assumes that all studies share a common true effect size. The observed differences between study results are entirely due to within-study sampling variability.
Model: \(\hat{\theta}_i = \theta + \varepsilon_i\), where \(\theta\) is the common effect and \(\varepsilon_i \sim N(0, \sigma_i^2)\).
Pooled estimate:
\[ \hat{\theta}_{FE} = \frac{\sum_{i=1}^k w_i \hat{\theta}_i}{\sum_{i=1}^k w_i} \]
where \(w_i = 1/\sigma_i^2\) is the inverse-variance weight (studies with smaller variance receive more weight).
Variance of the pooled estimate:
\[ \text{Var}(\hat{\theta}_{FE}) = \frac{1}{\sum_{i=1}^k w_i} \]
95% confidence interval:
\[ \hat{\theta}_{FE} \pm 1.96 \sqrt{\text{Var}(\hat{\theta}_{FE})} \]
For binary outcomes (Mantel-Haenszel method):
\[ \text{OR}_{MH} = \frac{\sum_{i=1}^k a_i d_i / N_i}{\sum_{i=1}^k b_i c_i / N_i} \]
where \(a_i, b_i, c_i, d_i\) are the cell counts of the 2×2 table in study \(i\) and \(N_i\) is the total sample size in study \(i\). The MH method is preferred when event rates are low or studies are small.
The random-effects model assumes that each study has its own true effect size \(\theta_i\), drawn from a distribution of effects with mean \(\theta\) and between-study variance \(\tau^2\). This is more realistic when studies vary in design, population, or intervention.
Model: \(\hat{\theta}_i = \theta_i + \varepsilon_i\), where \(\theta_i \sim N(\theta, \tau^2)\) and \(\varepsilon_i \sim N(0, \sigma_i^2)\).
The total variance of \(\hat{\theta}_i\) is \(\sigma_i^2 + \tau^2\).
Step 1: Estimate \(\tau^2\) (Cochran's Q and DerSimonian-Laird):
\[ Q = \sum_{i=1}^k w_i (\hat{\theta}_i - \hat{\theta}_{FE})^2 \]
\[ \hat{\tau}^2 = \max\left(0, \quad \frac{Q - (k-1)}{\sum w_i - \frac{\sum w_i^2}{\sum w_i}}\right) \]
If \(Q \leq k - 1\), then \(\hat{\tau}^2 = 0\) (no evidence of heterogeneity; equivalent to fixed-effect).
Step 2: Compute random-effects weights:
\[ w_i^* = \frac{1}{\sigma_i^2 + \hat{\tau}^2} \]
Step 3: Pooled estimate:
\[ \hat{\theta}_{RE} = \frac{\sum_{i=1}^k w_i^* \hat{\theta}_i}{\sum_{i=1}^k w_i^*} \]
Variance:
\[ \text{Var}(\hat{\theta}_{RE}) = \frac{1}{\sum_{i=1}^k w_i^*} \]
Key difference from fixed-effect: The random-effects weights are more balanced (smaller studies receive relatively more weight), and the confidence interval is wider, reflecting the additional uncertainty from between-study heterogeneity.
Cochran's Q statistic:
\[ Q = \sum_{i=1}^k w_i (\hat{\theta}_i - \hat{\theta}_{FE})^2 \]
Under \(H_0\) (homogeneity), \(Q \sim \chi^2_{k-1}\). Significant \(Q\) (p < 0.10) indicates heterogeneity. However, Q has low power with few studies and excessive power with many studies.
\(I^2\) statistic (Higgins & Thompson, 2002):
\[ I^2 = \frac{Q - (k-1)}{Q} \times 100\% = \frac{\hat{\tau}^2}{\hat{\tau}^2 + \tilde{\sigma}^2} \times 100\% \]
\(I^2\) describes the percentage of total variability due to between-study heterogeneity (rather than chance).
| \(I^2\) Value | Interpretation |
|---|---|
| 0–25% | Low heterogeneity |
| 25–50% | Moderate heterogeneity |
| 50–75% | Substantial heterogeneity |
| 75–100% | Considerable heterogeneity |
Prediction interval:
\[ \hat{\theta}_{RE} \pm t_{k-2, \, 0.025} \sqrt{\text{Var}(\hat{\theta}_{RE}) + \hat{\tau}^2} \]
The prediction interval estimates the range in which the true effect of a future study would fall, accounting for both the uncertainty in the mean and the between-study heterogeneity. A wide prediction interval (that crosses the null) indicates that the treatment effect may be beneficial in some settings and harmful in others.
Publication bias occurs when studies with significant or favourable results are more likely to be published than studies with non-significant or unfavourable results, leading to an inflated pooled estimate.
1. Funnel plot: A scatter plot of effect size (\(\hat{\theta}_i\)) vs. a measure of study size (SE or \(1/\sqrt{n_i}\)). In the absence of bias, the plot should be symmetric and funnel-shaped (larger studies near the top, smaller studies spreading at the bottom). Asymmetry suggests publication bias.
2. Egger's regression test:
\[ \frac{\hat{\theta}_i}{SE_i} = \alpha + \beta \times \frac{1}{SE_i} + \varepsilon_i \]
Under the null of no bias, the intercept \(\alpha = 0\). A significant \(\alpha \neq 0\) indicates asymmetry (potential publication bias).
3. Trim-and-fill method (Duval & Tweedie): Estimates the number of missing studies (due to publication bias) by iteratively trimming the most extreme small studies from the positive side, estimating the true centre, and then filling imputed studies symmetrically around the centre. The adjusted pooled estimate (after adding imputed studies) is compared with the original to assess sensitivity to bias.
4. Rosenthal's fail-safe N:
\[ N_{fs} = \frac{k \times \bar{Z}^2 - 2.706k}{2.706} \]
where \(k\) = number of studies and \(\bar{Z}\) = mean Z-score. \(N_{fs}\) is the number of unpublished null studies needed to make the overall result non-significant. If \(N_{fs} > 5k + 10\), the result is considered robust against publication bias.
A meta-analysis investigates whether statins reduce the risk of major cardiovascular events (MCE: non-fatal MI, non-fatal stroke, or cardiovascular death) in patients without prior cardiovascular disease (primary prevention). Five large randomised controlled trials are identified.
Study data (hazard ratios for MCE):
| Trial | n | Events (Treatment) | Events (Control) | HR | ln(HR) | SE(ln HR) |
|---|---|---|---|---|---|---|
| WOSCOPS | 6,595 | 174 | 248 | 0.69 | −0.371 | 0.094 |
| AFCAPS/TexCAPS | 6,605 | 116 | 183 | 0.63 | −0.462 | 0.109 |
| MEGA | 7,832 | 51 | 83 | 0.61 | −0.494 | 0.152 |
| JUPITER | 17,802 | 142 | 251 | 0.56 | −0.580 | 0.107 |
| HOPE-3 | 12,705 | 208 | 254 | 0.82 | −0.198 | 0.098 |
Fixed-effect meta-analysis:
Weights: \(w_i = 1/SE_i^2\). Pooled estimate:
\[ \hat{\theta}_{FE} = \frac{\sum w_i \hat{\theta}_i}{\sum w_i} = \frac{-0.371(113.2) + (-0.462)(84.2) + (-0.494)(43.3) + (-0.580)(87.4) + (-0.198)(104.1)}{432.2} \]
\[ = \frac{-42.0 - 38.9 - 21.4 - 50.7 - 20.6}{432.2} = \frac{-173.6}{432.2} = -0.402 \]
\[ SE(\hat{\theta}_{FE}) = \sqrt{1/432.2} = 0.0481 \]
95% CI: \(-0.402 \pm 1.96 \times 0.0481 = (-0.496, -0.308)\).
Pooled HR: \(\exp(-0.402) = 0.669\) (95% CI: 0.609–0.735), p < 0.001.
Heterogeneity:
\[ Q = 113.2(-0.371 + 0.402)^2 + 84.2(-0.462 + 0.402)^2 + 43.3(-0.494 + 0.402)^2 + 87.4(-0.580 + 0.402)^2 + 104.1(-0.198 + 0.402)^2 \]
\[ = 0.109 + 0.302 + 0.366 + 2.767 + 4.333 = 7.877 \]
\(df = 4\), \(p = 0.096\) (borderline significant). \(I^2 = (7.877 - 4)/7.877 \times 100 = 49.2\%\) (moderate heterogeneity).
Random-effects meta-analysis:
The DerSimonian-Laird denominator is \(C = \sum w_i - \sum w_i^2 / \sum w_i = 432.2 - 93.1 = 339.1\), so
\[ \hat{\tau}^2 = \frac{7.877 - 4}{339.1} = 0.0114 \]
Random-effects weights \(w_i^* = 1/(\sigma_i^2 + \hat{\tau}^2)\): 49.3, 42.9, 28.9, 43.7, 47.5 (total = 212.3).
\[ \hat{\theta}_{RE} = -0.410, \quad SE = \sqrt{1/212.3} = 0.0686 \]
Pooled HR: 0.664 (95% CI: 0.580–0.759), p < 0.001.
Prediction interval (using \(t_{k-2,\,0.025} = t_{3,\,0.025} = 3.182\)):
\[ -0.410 \pm 3.182 \times \sqrt{0.00471 + 0.0114} = -0.410 \pm 0.404 = (-0.814, -0.006) \]
On HR scale: (0.443, 0.994). The prediction interval stays just below 1.0, so a future trial is still expected to show benefit — but its upper limit lies close to the null, reflecting the moderate heterogeneity across populations.
Conclusion: Statins reduce the risk of major cardiovascular events by approximately 33% in primary prevention (pooled HR = 0.67, 95% CI: 0.60–0.75). The moderate heterogeneity (\(I^2 = 49\%\)) reflects differences in patient populations (e.g., JUPITER enrolled patients with elevated CRP; HOPE-3 included a broader population). The consistent direction of benefit across all 5 trials strengthens the evidence.
A meta-analysis examines the efficacy of selective serotonin reuptake inhibitors (SSRIs) vs. placebo in major depressive disorder. Seventeen randomised controlled trials are identified from a systematic literature search.
Pooled results (random-effects, standardised mean difference):
\[ \hat{\theta}_{RE} = 0.32 \quad (95\% \text{ CI: } 0.18, 0.46), \quad p < 0.001 \]
\[ \hat{\tau}^2 = 0.082, \quad I^2 = 62\% \quad (95\% \text{ CI: } 44\%, 74\%) \]
Interpretation of pooled effect: SSRIs produce a small-to-moderate improvement in depression scores compared to placebo (SMD = 0.32, approximately equivalent to a 3-point difference on the Hamilton Depression Rating Scale).
Publication bias assessment:
Funnel plot: Visual inspection reveals asymmetry — several small studies with large positive effects are present, but no small studies with null or negative effects, creating a gap in the lower-left portion of the funnel.
Egger's test: Intercept = 2.15 (SE = 0.82), t = 2.62, p = 0.020. Significant asymmetry, suggesting publication bias.
Trim-and-fill: The method estimates 4 missing studies. After imputation and re-analysis:
\[ \hat{\theta}_{\text{adjusted}} = 0.24 \quad (95\% \text{ CI: } 0.08, 0.40) \]
The adjusted estimate (SMD = 0.24) is smaller than the original (0.32), but remains statistically significant, suggesting that while publication bias inflates the estimated effect, SSRIs still have a genuine (though smaller) benefit over placebo.
Subgroup analysis:
| Subgroup | k | SMD | 95% CI |
|---|---|---|---|
| Severe depression (HAM-D ≥ 25) | 8 | 0.45 | (0.28, 0.62) |
| Moderate depression (HAM-D 18–24) | 9 | 0.19 | (0.02, 0.36) |
The effect is larger in severe depression (SMD = 0.45) compared to moderate depression (SMD = 0.19), with a test for subgroup difference yielding p = 0.035. This clinically meaningful finding — that antidepressants are more effective in severe depression — would not have been apparent from any single trial and illustrates a key strength of meta-analysis.
Key takeaway: This meta-analysis demonstrates the importance of assessing publication bias and exploring heterogeneity. Without the trim-and-fill adjustment, the pooled effect would be overstated by approximately 33% (0.32 vs. 0.24). The subgroup analysis provides clinically actionable information that individual trials could not deliver due to insufficient sample sizes within severity strata.