Categorical outcomes are among the most common endpoints in clinical trials. These include binary responses (e.g., response/no response, cured/not cured), ordinal outcomes (e.g., disease severity: mild/moderate/severe), and nominal outcomes (e.g., type of adverse event: cardiac/gastrointestinal/neurological). The analysis of categorical data ranges from simple contingency table methods (chi-square test, Fisher's exact test) to more sophisticated regression models (logistic regression) that adjust for covariates and handle multi-category outcomes.
The simplest analysis of a binary or categorical outcome compares the distribution of the outcome across treatment groups using a contingency table.
| Response (+) | No Response (−) | Total | |
|---|---|---|---|
| Treatment | a | b | a + b |
| Control | c | d | c + d |
| Total | a + c | b + d | N |
Pearson chi-square test (without continuity correction):
\[ \chi^2 = \sum_{i=1}^{r} \sum_{j=1}^{c} \frac{(O_{ij} - E_{ij})^2}{E_{ij}} \]
where \(O_{ij}\) = observed frequency and \(E_{ij} = \frac{R_i \times C_j}{N}\) is the expected frequency under the null hypothesis of independence.
For a 2×2 table, this simplifies to:
\[ \chi^2 = \frac{N(ad - bc)^2}{(a+b)(c+d)(a+c)(b+d)}, \quad df = 1 \]
Yates' continuity correction:
\[ \chi^2_{\text{Yates}} = \frac{N(|ad - bc| - N/2)^2}{(a+b)(c+d)(a+c)(b+d)} \]
Applied when expected frequencies are small (some guidelines: when any \(E_{ij} < 5\)), though modern practice favours Fisher's exact test in such cases.
When sample sizes are small and the chi-square approximation is unreliable, Fisher's exact test computes the exact probability of the observed or more extreme table configurations under the null hypothesis.
The probability of observing a specific 2×2 table with fixed marginals under \(H_0\):
\[ P(a, b, c, d) = \frac{\binom{a+b}{a} \binom{c+d}{c}}{\binom{N}{a+c}} = \frac{(a+b)!\,(c+d)!\,(a+c)!\,(b+d)!}{a!\,b!\,c!\,d!\,N!} \]
The p-value is the sum of probabilities of all tables as or more extreme than the observed one:
\[ p = \sum_{a' : P(a') \leq P(a)} P(a') \]
Fisher's exact test is recommended when any expected cell count is < 5, or when the total sample size is < 40.
Odds Ratio (OR):
\[ \text{OR} = \frac{a/b}{c/d} = \frac{ad}{bc} \]
Interpretation: OR = 1 means no association; OR > 1 means the treatment group has higher odds of response; OR < 1 means the treatment group has lower odds.
95% CI for ln(OR):
\[ \ln(\text{OR}) \pm 1.96 \sqrt{\frac{1}{a} + \frac{1}{b} + \frac{1}{c} + \frac{1}{d}} \]
Exponentiate the endpoints to obtain the CI on the OR scale.
Relative Risk (RR):
\[ \text{RR} = \frac{a/(a+b)}{c/(c+d)} \]
RR = 1 means no difference; RR > 1 means higher risk in the treatment group; RR < 1 means lower risk. For rare events, OR ≈ RR.
Risk Difference (RD):
\[ \text{RD} = \frac{a}{a+b} - \frac{c}{c+d} \]
The number needed to treat: \(\text{NNT} = 1/|\text{RD}|\) (when RD is positive and clinically meaningful).
Logistic regression extends the analysis of binary outcomes to include covariates, enabling adjusted treatment comparisons and identification of prognostic factors.
\[ \ln\left(\frac{p}{1-p}\right) = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + \cdots + \beta_k X_k \]
where \(p = P(Y = 1 \mid X_1, \ldots, X_k)\) is the probability of the outcome.
Interpretation: \(\exp(\beta_j)\) is the odds ratio for a one-unit increase in \(X_j\), holding all other covariates constant.
Maximum likelihood estimation: The log-likelihood function is:
\[ \ell(\boldsymbol{\beta}) = \sum_{i=1}^{n} \left[ Y_i \ln(\hat{p}_i) + (1 - Y_i) \ln(1 - \hat{p}_i) \right] \]
solved iteratively (Newton-Raphson or Fisher scoring).
Goodness of fit:
| Phase | Typical Categorical Endpoint | Analysis Method |
|---|---|---|
| Phase I | DLT (Yes/No) at each dose level | CRM, 3+3 rules, binomial CI for DLT rate |
| Phase II | Tumour response (CR+PR vs. SD+PD) | One-sample binomial test, Simon two-stage, Fisher's exact |
| Phase III | Composite clinical endpoint (event/no event) | Chi-square, logistic regression, Cochran-Mantel-Haenszel |
Cochran-Mantel-Haenszel (CMH) test: For stratified analysis (e.g., by study centre or baseline severity):
\[ \chi^2_{\text{CMH}} = \frac{\left[\sum_h (a_h - E(a_h))\right]^2}{\sum_h \text{Var}(a_h)} \]
where the sum is over strata \(h = 1, \ldots, H\), and \(a_h\) is the count in the treatment-response cell of stratum \(h\).
A Phase III trial compares a new antibiotic (Treatment) with the standard antibiotic (Control) for the treatment of community-acquired pneumonia. The primary endpoint is clinical cure at Day 14 (Yes/No). Three hundred patients are randomised 1:1.
Results:
| Cured | Not Cured | Total | |
|---|---|---|---|
| Treatment | 105 | 45 | 150 |
| Control | 82 | 68 | 150 |
| Total | 187 | 113 | 300 |
Chi-square test:
Expected frequencies: \(E_{11} = 150 \times 187/300 = 93.5\), \(E_{12} = 56.5\), \(E_{21} = 93.5\), \(E_{22} = 56.5\).
\[ \chi^2 = \frac{(105 - 93.5)^2}{93.5} + \frac{(45 - 56.5)^2}{56.5} + \frac{(82 - 93.5)^2}{93.5} + \frac{(68 - 56.5)^2}{56.5} \]
\[ = 1.414 + 2.341 + 1.414 + 2.341 = 7.51 \]
With \(df = 1\), \(p = 0.006\). There is a statistically significant difference in cure rates.
Odds ratio:
\[ \text{OR} = \frac{105 \times 68}{45 \times 82} = \frac{7140}{3690} = 1.935 \]
95% CI for ln(OR):
\[ \ln(1.935) \pm 1.96 \sqrt{\frac{1}{105} + \frac{1}{45} + \frac{1}{82} + \frac{1}{68}} \]
\[ 0.660 \pm 1.96 \times 0.2422 = 0.660 \pm 0.475 = (0.185, 1.135) \]
95% CI for OR: \((\exp(0.185), \exp(1.135)) = (1.204, 3.110)\).
Interpretation: Patients receiving the new antibiotic have approximately twice the odds of clinical cure compared to the standard antibiotic (OR = 1.94, 95% CI: 1.20–3.11). The 95% CI does not include 1.0, confirming statistical significance at \(\alpha = 0.05\).
Relative risk: \(\text{RR} = (105/150) / (82/150) = 0.70/0.547 = 1.28\). The new antibiotic increases the cure probability by 28% relative to the standard.
Risk difference: \(\text{RD} = 0.70 - 0.547 = 0.153\) (15.3 percentage points). NNT = 1/0.153 ≈ 7 — about 7 patients need to be treated with the new antibiotic to achieve one additional cure.
A Phase II trial evaluates a new targeted therapy for non-small cell lung cancer. The primary endpoint is objective tumour response (complete or partial response vs. stable or progressive disease). One hundred patients are enrolled, and the investigators wish to assess the treatment effect while adjusting for known prognostic factors: performance status (ECOG 0–1 vs. 2) and smoking status (current/former vs. never).
Data summary:
| Variable | Responders (n=38) | Non-responders (n=62) |
|---|---|---|
| Treatment (experimental vs. standard) | 24/38 | 16/62 |
| ECOG 0–1 | 30/38 | 35/62 |
| Never smoker | 18/38 | 22/62 |
Univariate analysis (crude OR for treatment):
\[ \text{OR}_{\text{crude}} = \frac{24 \times 46}{14 \times 16} = \frac{1104}{224} = 4.93 \]
Multivariable logistic regression:
\[ \ln\left(\frac{p}{1-p}\right) = \beta_0 + \beta_1 \cdot \text{Treatment} + \beta_2 \cdot \text{ECOG} + \beta_3 \cdot \text{Smoking} \]
Results (maximum likelihood estimates):
| Variable | \(\hat{\beta}\) | SE | OR = \(\exp(\hat{\beta})\) | 95% CI for OR | p-value |
|---|---|---|---|---|---|
| Intercept | −2.15 | 0.58 | — | — | <0.001 |
| Treatment | 1.42 | 0.48 | 4.14 | (1.62, 10.58) | 0.003 |
| ECOG 0–1 | 1.18 | 0.45 | 3.25 | (1.35, 7.84) | 0.009 |
| Never smoker | 0.82 | 0.44 | 2.27 | (0.96, 5.38) | 0.062 |
Interpretation: After adjusting for performance status and smoking, the experimental treatment remains strongly associated with tumour response (adjusted OR = 4.14, 95% CI: 1.62–10.58, p = 0.003). The adjusted OR (4.14) is somewhat lower than the crude OR (4.93), indicating modest confounding — patients with better performance status and never-smokers were slightly over-represented in the treatment group by chance. Good performance status (ECOG 0–1) is also a significant independent predictor of response (OR = 3.25, p = 0.009). Smoking status shows a trend (OR = 2.27, p = 0.062) that may reach significance with a larger sample.
The model's AUC (c-statistic) = 0.79, indicating good discrimination. The Hosmer-Lemeshow goodness-of-fit test yields \(\chi^2 = 5.2\), \(df = 8\), \(p = 0.74\), suggesting adequate model fit. This logistic regression analysis provides a more reliable assessment of the treatment effect than the unadjusted chi-square test because it accounts for prognostic imbalance between the groups.
Survival analysis deals with time-to-event data — the time from a defined origin (e.g., randomisation, diagnosis) to the occurrence of an event of interest (e.g., death, disease progression, relapse). A defining feature of survival data is censoring: some subjects have not experienced the event by the end of the study, so their exact survival time is unknown — we only know it exceeds their follow-up time. Standard methods (e.g., t-tests, linear regression) cannot handle censoring, necessitating specialised techniques: the Kaplan-Meier estimator, the log-rank test, and the Cox proportional hazards model.
Survival function: The probability of surviving beyond time \(t\):
\[ S(t) = P(T > t) = 1 - F(t) \]
where \(T\) is the random survival time and \(F(t)\) is the cumulative distribution function.
Hazard function: The instantaneous rate of the event at time \(t\), given survival to \(t\):
\[ h(t) = \lim_{\Delta t \to 0} \frac{P(t \leq T < t + \Delta t \mid T \geq t)}{\Delta t} \]
Cumulative hazard function:
\[ H(t) = \int_0^t h(u) \, du \]
Relationship:
\[ S(t) = \exp(-H(t)) \quad \text{and} \quad h(t) = -\frac{d}{dt} \ln S(t) \]
Types of censoring:
Key assumption: Censoring is non-informative — the reason for censoring is independent of the future risk of the event. If patients drop out because their disease is worsening, censoring is informative and standard methods are biased.
The Kaplan-Meier (KM) estimator is a non-parametric method for estimating the survival function from censored data. It computes the survival probability at each observed event time as a product of conditional survival probabilities.
Let \(t_1 < t_2 < \cdots < t_k\) be the distinct ordered event times. At time \(t_j\):
Kaplan-Meier estimator:
\[ \hat{S}(t) = \prod_{j: t_j \leq t} \left(1 - \frac{d_j}{n_j}\right) \]
Greenwood's formula for the variance:
\[ \text{Var}[\hat{S}(t)] = \left[\hat{S}(t)\right]^2 \sum_{j: t_j \leq t} \frac{d_j}{n_j(n_j - d_j)} \]
95% confidence interval:
\[ \hat{S}(t) \pm 1.96 \sqrt{\text{Var}[\hat{S}(t)]} \]
(More accurate: use the log-log transform for the CI.)
Median survival time: The smallest time \(t\) such that \(\hat{S}(t) \leq 0.50\).
The log-rank test compares the survival curves of two (or more) groups. It is a non-parametric test based on the observed vs. expected number of events at each event time.
At each distinct event time \(t_j\), compute:
Test statistic:
\[ \chi^2 = \frac{\left(\sum_j (O_{1j} - E_{1j})\right)^2}{\sum_j \frac{n_{1j} \cdot n_{2j} \cdot d_j \cdot (n_j - d_j)}{n_j^2 \cdot (n_j - 1)}} \]
Under \(H_0\) (equal survival curves), \(\chi^2 \sim \chi^2_1\) for a two-group comparison.
Properties:
The Cox proportional hazards model is a semi-parametric regression model that relates the hazard function to covariates without specifying the baseline hazard.
\[ h(t \mid \mathbf{X}) = h_0(t) \exp(\beta_1 X_1 + \beta_2 X_2 + \cdots + \beta_k X_k) \]
where:
Key property — Proportional hazards:
\[ \frac{h(t \mid \mathbf{X}_1)}{h(t \mid \mathbf{X}_2)} = \exp\left(\sum_j \beta_j (X_{1j} - X_{2j})\right) \]
The hazard ratio is constant over time — this is the proportional hazards assumption.
Partial likelihood estimation:
\[ L(\boldsymbol{\beta}) = \prod_{i: \delta_i = 1} \frac{\exp(\boldsymbol{\beta}^T \mathbf{X}_i)} {\sum_{j \in R(t_i)} \exp(\boldsymbol{\beta}^T \mathbf{X}_j)} \]
where \(\delta_i = 1\) if subject \(i\) had an event, and \(R(t_i)\) is the risk set at time \(t_i\) (all subjects still at risk when subject \(i\) has the event). The baseline hazard \(h_0(t)\) cancels out, so it does not need to be specified — this is why the model is "semi-parametric."
Testing the proportional hazards assumption:
| Element | Description |
|---|---|
| Median follow-up time | Reported using the reverse Kaplan-Meier method (censoring events as the "event") |
| Number of events | Events and censoring patterns in each group |
| Median survival time | With 95% CI from KM for each group |
| Survival rates | At pre-specified time points (e.g., 1-year, 2-year) with 95% CI |
| Kaplan-Meier curves | With number at risk table below the plot |
| Hazard ratio | With 95% CI from the Cox model |
| Log-rank p-value | Unadjusted comparison |
| Cox model results | Adjusted HR with 95% CI, p-value, and covariates listed |
| PH assumption check | Report the test and results (Schoenfeld residuals test) |
For randomised clinical trials with time-to-event outcomes, the CONSORT statement recommends reporting: (1) the number of participants in each group, (2) the number of events, (3) the effect estimate (HR) with 95% CI, (4) the absolute risk at a clinically relevant time point, and (5) the p-value. Kaplan-Meier curves should always include the number at risk table and indicate censoring marks. Avoid reporting only median survival times without the full KM curves, as the shape of the curves (e.g., early separation vs. late separation) has important clinical implications.
A Phase III trial compares adjuvant chemotherapy + new targeted agent (Treatment) with adjuvant chemotherapy alone (Control) in 200 patients with HER2-positive early breast cancer. The primary endpoint is disease-free survival (DFS) — the time from randomisation to disease recurrence or death from any cause. Median follow-up is 3.5 years.
Event summary:
| Treatment (n=100) | Control (n=100) | |
|---|---|---|
| Recurrence or death | 22 | 38 |
| Censored | 78 | 62 |
Kaplan-Meier estimates at key time points:
| Time | \(\hat{S}(t)\) Treatment | 95% CI | \(\hat{S}(t)\) Control | 95% CI |
|---|---|---|---|---|
| 1 year | 0.95 | (0.90, 0.98) | 0.91 | (0.84, 0.95) |
| 2 years | 0.88 | (0.80, 0.93) | 0.75 | (0.65, 0.82) |
| 3 years | 0.82 | (0.73, 0.89) | 0.64 | (0.53, 0.73) |
| 4 years | 0.79 | (0.69, 0.86) | 0.60 | (0.49, 0.70) |
Median DFS: not reached in either arm — over 50% of patients remain disease-free at 4 years in both groups (Treatment Ŝ = 0.79, Control Ŝ = 0.60 at 4 years). The benefit is therefore quantified by the log-rank test and hazard ratio below rather than by a difference in medians.
Log-rank test:
Sum of \((O - E)\) for Treatment group = \(22 - 30.8 = -8.8\). The variance term = 12.4.
\[ \chi^2 = \frac{(-8.8)^2}{12.4} = \frac{77.44}{12.4} = 6.24 \]
With \(df = 1\), \(p = 0.012\). The difference in DFS curves is statistically significant.
Cox model (unadjusted):
\[ \hat{HR} = \exp(\hat{\beta}) = 0.51 \quad (95\% \text{ CI: } 0.30, 0.87), \quad p = 0.013 \]
The targeted agent reduces the hazard of recurrence or death by 49% compared to chemotherapy alone. The log-log survival plots show approximately parallel curves, supporting the proportional hazards assumption (Schoenfeld residual test: \(p = 0.68\)).
Number needed to treat at 3 years: The 3-year DFS is 0.82 (Treatment) vs. 0.64 (Control), so the absolute risk reduction is 0.18 (18 percentage points), and \(\text{NNT} = 1/0.18 ≈ 6\). About 6 patients need to receive the targeted agent to prevent one additional recurrence or death at 3 years.
A randomised trial compares a new drug (sacubitril/valsartan) with enalapril in 1,800 patients with heart failure with reduced ejection fraction (HFrEF). The primary endpoint is cardiovascular death or heart failure hospitalisation. The investigators perform a Cox regression adjusting for important prognostic factors.
Unadjusted log-rank test: HR = 0.80 (95% CI: 0.68–0.94), log-rank p = 0.007.
Multivariable Cox model:
\[ h(t \mid \mathbf{X}) = h_0(t) \exp(\beta_1 \cdot \text{Treatment} + \beta_2 \cdot \text{Age} + \beta_3 \cdot \text{EF} + \beta_4 \cdot \text{NYHA} + \beta_5 \cdot \text{Diabetes} + \beta_6 \cdot \text{eGFR}) \]
| Variable | \(\hat{\beta}\) | SE | Adjusted HR | 95% CI for HR | p-value |
|---|---|---|---|---|---|
| Treatment | −0.24 | 0.084 | 0.79 | (0.67, 0.93) | 0.004 |
| Age (per 10 yr) | 0.15 | 0.051 | 1.16 | (1.05, 1.28) | 0.003 |
| Ejection fraction (per 5%) | −0.09 | 0.035 | 0.91 | (0.85, 0.98) | 0.010 |
| NYHA III–IV | 0.48 | 0.102 | 1.62 | (1.32, 1.98) | <0.001 |
| Diabetes | 0.28 | 0.095 | 1.32 | (1.10, 1.59) | 0.003 |
| eGFR (per 10 mL/min) | −0.10 | 0.040 | 0.90 | (0.83, 0.98) | 0.013 |
Interpretation: After adjusting for age, ejection fraction, NYHA class, diabetes, and kidney function, the new drug reduces the hazard of the composite endpoint by 21% (adjusted HR = 0.79, 95% CI: 0.67–0.93, p = 0.004). The adjusted HR (0.79) is very close to the unadjusted HR (0.80), indicating that the randomisation produced well-balanced groups and confounding is minimal. However, the multivariable model provides a more precise estimate and identifies independent prognostic factors: worse NYHA class (HR = 1.62), diabetes (HR = 1.32), older age (HR = 1.16 per 10 years), lower eGFR, and lower ejection fraction all independently predict worse outcomes.
PH assumption check: Schoenfeld residual test for the treatment variable: \(\rho = -0.04\), \(p = 0.52\), confirming no violation of proportional hazards. The treatment effect is consistent over the follow-up period.
Subgroup analysis: The treatment effect is consistent across subgroups (interaction p-values all > 0.10), with no evidence that the benefit varies by age, sex, diabetes status, or baseline ejection fraction. This supports the generalisability of the treatment benefit across the heart failure population studied.