WHEN SURVIVAL CURVES MISLEAD INTERPRETATION: NON-IDENTIFIABILITY OF TIME-TO-EVENT EVIDENCE IN CROSS-STUDY HEOR COMPARISONS
Author(s)
Federico Manevy, MSc1, Mercedeh Ghadessi, MSc2, Noman Paracha, MSc1.
1Bayer Consumer Care AG, Basel, Switzerland, 2Bayer Healthcare LLC, Whippany, NJ, USA.
1Bayer Consumer Care AG, Basel, Switzerland, 2Bayer Healthcare LLC, Whippany, NJ, USA.
OBJECTIVES: Many HEOR submissions to HTA bodies and payers rely on cross-study comparisons, including externally controlled single-arm packages and unanchored indirect comparisons. These lack design-based identification and require unverifiable cross-population assumptions. Similarities or differences between Kaplan-Meier (KM) curves are often interpreted as evidence of treatment effects or population comparability. We stress-test this interpretation using a well-established result: the non-identifiability of mixture and frailty survival models.
METHODS: We built a minimal discrete-time, individual-level survival model in which each subject has two binary latent attributes reflecting unobserved heterogeneity: frailty, fixing the starting-position range relative to an absorbing "death" state, and progression speed, fixing the per-cycle step toward it. Given a starting position drawn uniformly within that range, time-to-event is deterministic. We specified two data-generating processes (DGPs) differing in their joint distribution, with identical progression dynamics and, by construction, equivalent theoretical population-level survival functions. Sample composition was held fixed across 1,000 replicates, so between-replicate variation arose from individual-level starting positions. Excluding treatment effects, censoring, and observed covariates isolated two identifiability failures: (i) distinct DGPs producing similar KM curves; (ii) one DGP producing divergent KM curves across replicates.
RESULTS: Distinct DGPs produced KM curves overlapping throughout follow-up, including at common landmarks and the median, while repeated samples from the same DGP yielded divergent trajectories from individual-level starting positions alone. KM similarity does not establish comparability of underlying DGPs, and KM differences alone—absent a controlled design—do not establish a treatment effect.
CONCLUSIONS: KM curves do not identify the underlying process, so cross-study conclusions about comparability or treatment effect cannot rest on curve comparison. Identification must come from design—via a pre-specified estimand and target-trial emulation. Population and heterogeneity assumptions should be stated (e.g. using causal diagrams), and analyses should report the range of effects compatible with the same curves under unmeasured heterogeneity, rather than a single estimate.
METHODS: We built a minimal discrete-time, individual-level survival model in which each subject has two binary latent attributes reflecting unobserved heterogeneity: frailty, fixing the starting-position range relative to an absorbing "death" state, and progression speed, fixing the per-cycle step toward it. Given a starting position drawn uniformly within that range, time-to-event is deterministic. We specified two data-generating processes (DGPs) differing in their joint distribution, with identical progression dynamics and, by construction, equivalent theoretical population-level survival functions. Sample composition was held fixed across 1,000 replicates, so between-replicate variation arose from individual-level starting positions. Excluding treatment effects, censoring, and observed covariates isolated two identifiability failures: (i) distinct DGPs producing similar KM curves; (ii) one DGP producing divergent KM curves across replicates.
RESULTS: Distinct DGPs produced KM curves overlapping throughout follow-up, including at common landmarks and the median, while repeated samples from the same DGP yielded divergent trajectories from individual-level starting positions alone. KM similarity does not establish comparability of underlying DGPs, and KM differences alone—absent a controlled design—do not establish a treatment effect.
CONCLUSIONS: KM curves do not identify the underlying process, so cross-study conclusions about comparability or treatment effect cannot rest on curve comparison. Identification must come from design—via a pre-specified estimand and target-trial emulation. Population and heterogeneity assumptions should be stated (e.g. using causal diagrams), and analyses should report the range of effects compatible with the same curves under unmeasured heterogeneity, rather than a single estimate.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
CO204
Topic
Clinical Outcomes, Methodological & Statistical Research
Topic Subcategory
Comparative Effectiveness or Efficacy
Disease
Oncology, Rare & Orphan Diseases