Skip to the content

Topics Covered

Categorical Outcomes Chi-square Tests Fisher's Exact Test Logistic Regression Odds Ratio and Relative Risk Survival Data Concepts Kaplan-Meier Estimator Log-rank Test Cox Proportional Hazards Model Reporting Standards
On this page
  1. 1.1 Contingency Tables and Chi-square Test
  2. 1.2 Fisher's Exact Test
  3. 1.3 Measures of Association for Categorical Data
  4. 1.4 Logistic Regression
  5. 1.5 Analysis by Trial Phase
  6. 2.1 Key Concepts in Survival Analysis
  7. 2.2 Kaplan-Meier Estimator
  8. 2.3 Log-rank Test
  9. 2.4 Cox Proportional Hazards Model
  10. 2.5 Reporting Survival Analyses

1. Analysis of Categorical Outcomes from Phase I–III Trials

KEY CONCEPT

Categorical outcomes are among the most common endpoints in clinical trials. These include binary responses (e.g., response/no response, cured/not cured), ordinal outcomes (e.g., disease severity: mild/moderate/severe), and nominal outcomes (e.g., type of adverse event: cardiac/gastrointestinal/neurological). The analysis of categorical data ranges from simple contingency table methods (chi-square test, Fisher's exact test) to more sophisticated regression models (logistic regression) that adjust for covariates and handle multi-category outcomes.

1.1 Contingency Tables and Chi-square Test

The simplest analysis of a binary or categorical outcome compares the distribution of the outcome across treatment groups using a contingency table.

2×2 CONTINGENCY TABLE
Response (+)No Response (−)Total
Treatmentaba + b
Controlcdc + d
Totala + cb + dN

Pearson chi-square test (without continuity correction):

\[ \chi^2 = \sum_{i=1}^{r} \sum_{j=1}^{c} \frac{(O_{ij} - E_{ij})^2}{E_{ij}} \]

where \(O_{ij}\) = observed frequency and \(E_{ij} = \frac{R_i \times C_j}{N}\) is the expected frequency under the null hypothesis of independence.

For a 2×2 table, this simplifies to:

\[ \chi^2 = \frac{N(ad - bc)^2}{(a+b)(c+d)(a+c)(b+d)}, \quad df = 1 \]

Yates' continuity correction:

\[ \chi^2_{\text{Yates}} = \frac{N(|ad - bc| - N/2)^2}{(a+b)(c+d)(a+c)(b+d)} \]

Applied when expected frequencies are small (some guidelines: when any \(E_{ij} < 5\)), though modern practice favours Fisher's exact test in such cases.

1.2 Fisher's Exact Test

When sample sizes are small and the chi-square approximation is unreliable, Fisher's exact test computes the exact probability of the observed or more extreme table configurations under the null hypothesis.

FISHER'S EXACT TEST

The probability of observing a specific 2×2 table with fixed marginals under \(H_0\):

\[ P(a, b, c, d) = \frac{\binom{a+b}{a} \binom{c+d}{c}}{\binom{N}{a+c}} = \frac{(a+b)!\,(c+d)!\,(a+c)!\,(b+d)!}{a!\,b!\,c!\,d!\,N!} \]

The p-value is the sum of probabilities of all tables as or more extreme than the observed one:

\[ p = \sum_{a' : P(a') \leq P(a)} P(a') \]

Fisher's exact test is recommended when any expected cell count is < 5, or when the total sample size is < 40.

1.3 Measures of Association for Categorical Data

ODDS RATIO AND RELATIVE RISK

Odds Ratio (OR):

\[ \text{OR} = \frac{a/b}{c/d} = \frac{ad}{bc} \]

Interpretation: OR = 1 means no association; OR > 1 means the treatment group has higher odds of response; OR < 1 means the treatment group has lower odds.

95% CI for ln(OR):

\[ \ln(\text{OR}) \pm 1.96 \sqrt{\frac{1}{a} + \frac{1}{b} + \frac{1}{c} + \frac{1}{d}} \]

Exponentiate the endpoints to obtain the CI on the OR scale.

Relative Risk (RR):

\[ \text{RR} = \frac{a/(a+b)}{c/(c+d)} \]

RR = 1 means no difference; RR > 1 means higher risk in the treatment group; RR < 1 means lower risk. For rare events, OR ≈ RR.

Risk Difference (RD):

\[ \text{RD} = \frac{a}{a+b} - \frac{c}{c+d} \]

The number needed to treat: \(\text{NNT} = 1/|\text{RD}|\) (when RD is positive and clinically meaningful).

1.4 Logistic Regression

Logistic regression extends the analysis of binary outcomes to include covariates, enabling adjusted treatment comparisons and identification of prognostic factors.

LOGISTIC REGRESSION MODEL

\[ \ln\left(\frac{p}{1-p}\right) = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + \cdots + \beta_k X_k \]

where \(p = P(Y = 1 \mid X_1, \ldots, X_k)\) is the probability of the outcome.

Interpretation: \(\exp(\beta_j)\) is the odds ratio for a one-unit increase in \(X_j\), holding all other covariates constant.

Maximum likelihood estimation: The log-likelihood function is:

\[ \ell(\boldsymbol{\beta}) = \sum_{i=1}^{n} \left[ Y_i \ln(\hat{p}_i) + (1 - Y_i) \ln(1 - \hat{p}_i) \right] \]

solved iteratively (Newton-Raphson or Fisher scoring).

Goodness of fit:

1.5 Analysis by Trial Phase

CATEGORICAL ANALYSIS ACROSS PHASES
PhaseTypical Categorical EndpointAnalysis Method
Phase IDLT (Yes/No) at each dose levelCRM, 3+3 rules, binomial CI for DLT rate
Phase IITumour response (CR+PR vs. SD+PD)One-sample binomial test, Simon two-stage, Fisher's exact
Phase IIIComposite clinical endpoint (event/no event)Chi-square, logistic regression, Cochran-Mantel-Haenszel

Cochran-Mantel-Haenszel (CMH) test: For stratified analysis (e.g., by study centre or baseline severity):

\[ \chi^2_{\text{CMH}} = \frac{\left[\sum_h (a_h - E(a_h))\right]^2}{\sum_h \text{Var}(a_h)} \]

where the sum is over strata \(h = 1, \ldots, H\), and \(a_h\) is the count in the treatment-response cell of stratum \(h\).

EXAMPLE 1 — Chi-square and Odds Ratio in a Phase III Trial

A Phase III trial compares a new antibiotic (Treatment) with the standard antibiotic (Control) for the treatment of community-acquired pneumonia. The primary endpoint is clinical cure at Day 14 (Yes/No). Three hundred patients are randomised 1:1.

Results:

CuredNot CuredTotal
Treatment10545150
Control8268150
Total187113300

Chi-square test:

Expected frequencies: \(E_{11} = 150 \times 187/300 = 93.5\), \(E_{12} = 56.5\), \(E_{21} = 93.5\), \(E_{22} = 56.5\).

\[ \chi^2 = \frac{(105 - 93.5)^2}{93.5} + \frac{(45 - 56.5)^2}{56.5} + \frac{(82 - 93.5)^2}{93.5} + \frac{(68 - 56.5)^2}{56.5} \]

\[ = 1.414 + 2.341 + 1.414 + 2.341 = 7.51 \]

With \(df = 1\), \(p = 0.006\). There is a statistically significant difference in cure rates.

Odds ratio:

\[ \text{OR} = \frac{105 \times 68}{45 \times 82} = \frac{7140}{3690} = 1.935 \]

95% CI for ln(OR):

\[ \ln(1.935) \pm 1.96 \sqrt{\frac{1}{105} + \frac{1}{45} + \frac{1}{82} + \frac{1}{68}} \]

\[ 0.660 \pm 1.96 \times 0.2422 = 0.660 \pm 0.475 = (0.185, 1.135) \]

95% CI for OR: \((\exp(0.185), \exp(1.135)) = (1.204, 3.110)\).

Interpretation: Patients receiving the new antibiotic have approximately twice the odds of clinical cure compared to the standard antibiotic (OR = 1.94, 95% CI: 1.20–3.11). The 95% CI does not include 1.0, confirming statistical significance at \(\alpha = 0.05\).

Relative risk: \(\text{RR} = (105/150) / (82/150) = 0.70/0.547 = 1.28\). The new antibiotic increases the cure probability by 28% relative to the standard.

Risk difference: \(\text{RD} = 0.70 - 0.547 = 0.153\) (15.3 percentage points). NNT = 1/0.153 ≈ 7 — about 7 patients need to be treated with the new antibiotic to achieve one additional cure.

EXAMPLE 2 — Logistic Regression in a Phase II Oncology Trial

A Phase II trial evaluates a new targeted therapy for non-small cell lung cancer. The primary endpoint is objective tumour response (complete or partial response vs. stable or progressive disease). One hundred patients are enrolled, and the investigators wish to assess the treatment effect while adjusting for known prognostic factors: performance status (ECOG 0–1 vs. 2) and smoking status (current/former vs. never).

Data summary:

VariableResponders (n=38)Non-responders (n=62)
Treatment (experimental vs. standard)24/3816/62
ECOG 0–130/3835/62
Never smoker18/3822/62

Univariate analysis (crude OR for treatment):

\[ \text{OR}_{\text{crude}} = \frac{24 \times 46}{14 \times 16} = \frac{1104}{224} = 4.93 \]

Multivariable logistic regression:

\[ \ln\left(\frac{p}{1-p}\right) = \beta_0 + \beta_1 \cdot \text{Treatment} + \beta_2 \cdot \text{ECOG} + \beta_3 \cdot \text{Smoking} \]

Results (maximum likelihood estimates):

Variable\(\hat{\beta}\)SEOR = \(\exp(\hat{\beta})\)95% CI for ORp-value
Intercept−2.150.58——<0.001
Treatment1.420.484.14(1.62, 10.58)0.003
ECOG 0–11.180.453.25(1.35, 7.84)0.009
Never smoker0.820.442.27(0.96, 5.38)0.062

Interpretation: After adjusting for performance status and smoking, the experimental treatment remains strongly associated with tumour response (adjusted OR = 4.14, 95% CI: 1.62–10.58, p = 0.003). The adjusted OR (4.14) is somewhat lower than the crude OR (4.93), indicating modest confounding — patients with better performance status and never-smokers were slightly over-represented in the treatment group by chance. Good performance status (ECOG 0–1) is also a significant independent predictor of response (OR = 3.25, p = 0.009). Smoking status shows a trend (OR = 2.27, p = 0.062) that may reach significance with a larger sample.

The model's AUC (c-statistic) = 0.79, indicating good discrimination. The Hosmer-Lemeshow goodness-of-fit test yields \(\chi^2 = 5.2\), \(df = 8\), \(p = 0.74\), suggesting adequate model fit. This logistic regression analysis provides a more reliable assessment of the treatment effect than the unadjusted chi-square test because it accounts for prognostic imbalance between the groups.

2. Analysis of Survival Data from Clinical Trials

KEY CONCEPT

Survival analysis deals with time-to-event data — the time from a defined origin (e.g., randomisation, diagnosis) to the occurrence of an event of interest (e.g., death, disease progression, relapse). A defining feature of survival data is censoring: some subjects have not experienced the event by the end of the study, so their exact survival time is unknown — we only know it exceeds their follow-up time. Standard methods (e.g., t-tests, linear regression) cannot handle censoring, necessitating specialised techniques: the Kaplan-Meier estimator, the log-rank test, and the Cox proportional hazards model.

2.1 Key Concepts in Survival Analysis

FUNDAMENTAL FUNCTIONS IN SURVIVAL ANALYSIS

Survival function: The probability of surviving beyond time \(t\):

\[ S(t) = P(T > t) = 1 - F(t) \]

where \(T\) is the random survival time and \(F(t)\) is the cumulative distribution function.

Hazard function: The instantaneous rate of the event at time \(t\), given survival to \(t\):

\[ h(t) = \lim_{\Delta t \to 0} \frac{P(t \leq T < t + \Delta t \mid T \geq t)}{\Delta t} \]

Cumulative hazard function:

\[ H(t) = \int_0^t h(u) \, du \]

Relationship:

\[ S(t) = \exp(-H(t)) \quad \text{and} \quad h(t) = -\frac{d}{dt} \ln S(t) \]

Types of censoring:

Key assumption: Censoring is non-informative — the reason for censoring is independent of the future risk of the event. If patients drop out because their disease is worsening, censoring is informative and standard methods are biased.

2.2 Kaplan-Meier Estimator

The Kaplan-Meier (KM) estimator is a non-parametric method for estimating the survival function from censored data. It computes the survival probability at each observed event time as a product of conditional survival probabilities.

KAPLAN-MEIER ESTIMATOR

Let \(t_1 < t_2 < \cdots < t_k\) be the distinct ordered event times. At time \(t_j\):

Kaplan-Meier estimator:

\[ \hat{S}(t) = \prod_{j: t_j \leq t} \left(1 - \frac{d_j}{n_j}\right) \]

Greenwood's formula for the variance:

\[ \text{Var}[\hat{S}(t)] = \left[\hat{S}(t)\right]^2 \sum_{j: t_j \leq t} \frac{d_j}{n_j(n_j - d_j)} \]

95% confidence interval:

\[ \hat{S}(t) \pm 1.96 \sqrt{\text{Var}[\hat{S}(t)]} \]

(More accurate: use the log-log transform for the CI.)

Median survival time: The smallest time \(t\) such that \(\hat{S}(t) \leq 0.50\).

2.3 Log-rank Test

The log-rank test compares the survival curves of two (or more) groups. It is a non-parametric test based on the observed vs. expected number of events at each event time.

LOG-RANK TEST

At each distinct event time \(t_j\), compute:

Test statistic:

\[ \chi^2 = \frac{\left(\sum_j (O_{1j} - E_{1j})\right)^2}{\sum_j \frac{n_{1j} \cdot n_{2j} \cdot d_j \cdot (n_j - d_j)}{n_j^2 \cdot (n_j - 1)}} \]

Under \(H_0\) (equal survival curves), \(\chi^2 \sim \chi^2_1\) for a two-group comparison.

Properties:

2.4 Cox Proportional Hazards Model

The Cox proportional hazards model is a semi-parametric regression model that relates the hazard function to covariates without specifying the baseline hazard.

COX PROPORTIONAL HAZARDS MODEL

\[ h(t \mid \mathbf{X}) = h_0(t) \exp(\beta_1 X_1 + \beta_2 X_2 + \cdots + \beta_k X_k) \]

where:

Key property — Proportional hazards:

\[ \frac{h(t \mid \mathbf{X}_1)}{h(t \mid \mathbf{X}_2)} = \exp\left(\sum_j \beta_j (X_{1j} - X_{2j})\right) \]

The hazard ratio is constant over time — this is the proportional hazards assumption.

Partial likelihood estimation:

\[ L(\boldsymbol{\beta}) = \prod_{i: \delta_i = 1} \frac{\exp(\boldsymbol{\beta}^T \mathbf{X}_i)} {\sum_{j \in R(t_i)} \exp(\boldsymbol{\beta}^T \mathbf{X}_j)} \]

where \(\delta_i = 1\) if subject \(i\) had an event, and \(R(t_i)\) is the risk set at time \(t_i\) (all subjects still at risk when subject \(i\) has the event). The baseline hazard \(h_0(t)\) cancels out, so it does not need to be specified — this is why the model is "semi-parametric."

Testing the proportional hazards assumption:

2.5 Reporting Survival Analyses

KEY ELEMENTS IN REPORTING SURVIVAL ANALYSES
ElementDescription
Median follow-up timeReported using the reverse Kaplan-Meier method (censoring events as the "event")
Number of eventsEvents and censoring patterns in each group
Median survival timeWith 95% CI from KM for each group
Survival ratesAt pre-specified time points (e.g., 1-year, 2-year) with 95% CI
Kaplan-Meier curvesWith number at risk table below the plot
Hazard ratioWith 95% CI from the Cox model
Log-rank p-valueUnadjusted comparison
Cox model resultsAdjusted HR with 95% CI, p-value, and covariates listed
PH assumption checkReport the test and results (Schoenfeld residuals test)
CONSORT GUIDELINE

For randomised clinical trials with time-to-event outcomes, the CONSORT statement recommends reporting: (1) the number of participants in each group, (2) the number of events, (3) the effect estimate (HR) with 95% CI, (4) the absolute risk at a clinically relevant time point, and (5) the p-value. Kaplan-Meier curves should always include the number at risk table and indicate censoring marks. Avoid reporting only median survival times without the full KM curves, as the shape of the curves (e.g., early separation vs. late separation) has important clinical implications.

EXAMPLE 1 — Kaplan-Meier and Log-rank Test in a Breast Cancer Trial

A Phase III trial compares adjuvant chemotherapy + new targeted agent (Treatment) with adjuvant chemotherapy alone (Control) in 200 patients with HER2-positive early breast cancer. The primary endpoint is disease-free survival (DFS) — the time from randomisation to disease recurrence or death from any cause. Median follow-up is 3.5 years.

Event summary:

Treatment (n=100)Control (n=100)
Recurrence or death2238
Censored7862

Kaplan-Meier estimates at key time points:

Time\(\hat{S}(t)\) Treatment95% CI\(\hat{S}(t)\) Control95% CI
1 year0.95(0.90, 0.98)0.91(0.84, 0.95)
2 years0.88(0.80, 0.93)0.75(0.65, 0.82)
3 years0.82(0.73, 0.89)0.64(0.53, 0.73)
4 years0.79(0.69, 0.86)0.60(0.49, 0.70)

Median DFS: not reached in either arm — over 50% of patients remain disease-free at 4 years in both groups (Treatment Ŝ = 0.79, Control Ŝ = 0.60 at 4 years). The benefit is therefore quantified by the log-rank test and hazard ratio below rather than by a difference in medians.

Log-rank test:

Sum of \((O - E)\) for Treatment group = \(22 - 30.8 = -8.8\). The variance term = 12.4.

\[ \chi^2 = \frac{(-8.8)^2}{12.4} = \frac{77.44}{12.4} = 6.24 \]

With \(df = 1\), \(p = 0.012\). The difference in DFS curves is statistically significant.

Cox model (unadjusted):

\[ \hat{HR} = \exp(\hat{\beta}) = 0.51 \quad (95\% \text{ CI: } 0.30, 0.87), \quad p = 0.013 \]

The targeted agent reduces the hazard of recurrence or death by 49% compared to chemotherapy alone. The log-log survival plots show approximately parallel curves, supporting the proportional hazards assumption (Schoenfeld residual test: \(p = 0.68\)).

Number needed to treat at 3 years: The 3-year DFS is 0.82 (Treatment) vs. 0.64 (Control), so the absolute risk reduction is 0.18 (18 percentage points), and \(\text{NNT} = 1/0.18 ≈ 6\). About 6 patients need to receive the targeted agent to prevent one additional recurrence or death at 3 years.

0.50 0.00 0.25 0.50 0.75 1.00 0 1 2 3 4 Time since randomisation (years) Disease-free survival Ŝ(t) Treatment (n=100) Control (n=100)
Fig. 4.1 — Kaplan–Meier disease-free survival curves for the breast-cancer trial (Example 1). Each curve is a step function that holds until the next yearly estimate and drops at it; markers show the tabulated Ŝ(t). The Treatment arm (blue) stays above the Control arm (red) throughout, and neither crosses the dashed 0.50 median line within 4 years (log-rank p = 0.012, HR = 0.51).
EXAMPLE 2 — Cox Regression with Multiple Covariates in a Heart Failure Trial

A randomised trial compares a new drug (sacubitril/valsartan) with enalapril in 1,800 patients with heart failure with reduced ejection fraction (HFrEF). The primary endpoint is cardiovascular death or heart failure hospitalisation. The investigators perform a Cox regression adjusting for important prognostic factors.

Unadjusted log-rank test: HR = 0.80 (95% CI: 0.68–0.94), log-rank p = 0.007.

Multivariable Cox model:

\[ h(t \mid \mathbf{X}) = h_0(t) \exp(\beta_1 \cdot \text{Treatment} + \beta_2 \cdot \text{Age} + \beta_3 \cdot \text{EF} + \beta_4 \cdot \text{NYHA} + \beta_5 \cdot \text{Diabetes} + \beta_6 \cdot \text{eGFR}) \]

Variable\(\hat{\beta}\)SEAdjusted HR95% CI for HRp-value
Treatment−0.240.0840.79(0.67, 0.93)0.004
Age (per 10 yr)0.150.0511.16(1.05, 1.28)0.003
Ejection fraction (per 5%)−0.090.0350.91(0.85, 0.98)0.010
NYHA III–IV0.480.1021.62(1.32, 1.98)<0.001
Diabetes0.280.0951.32(1.10, 1.59)0.003
eGFR (per 10 mL/min)−0.100.0400.90(0.83, 0.98)0.013

Interpretation: After adjusting for age, ejection fraction, NYHA class, diabetes, and kidney function, the new drug reduces the hazard of the composite endpoint by 21% (adjusted HR = 0.79, 95% CI: 0.67–0.93, p = 0.004). The adjusted HR (0.79) is very close to the unadjusted HR (0.80), indicating that the randomisation produced well-balanced groups and confounding is minimal. However, the multivariable model provides a more precise estimate and identifies independent prognostic factors: worse NYHA class (HR = 1.62), diabetes (HR = 1.32), older age (HR = 1.16 per 10 years), lower eGFR, and lower ejection fraction all independently predict worse outcomes.

PH assumption check: Schoenfeld residual test for the treatment variable: \(\rho = -0.04\), \(p = 0.52\), confirming no violation of proportional hazards. The treatment effect is consistent over the follow-up period.

Subgroup analysis: The treatment effect is consistent across subgroups (interaction p-values all > 0.10), with no evidence that the benefit varies by age, sex, diabetes status, or baseline ejection fraction. This supports the generalisability of the treatment benefit across the heart failure population studied.