VALIDATION OF A GENERAL-PURPOSE AI AGENT (CLAUDE) FOR KAPLAN-MEIER CURVE DIGITIZATION AND PARAMETRIC SURVIVAL EXTRAPOLATION
Author(s)
Iwona Zerda, MSc1, Ghassen Dhiab, Msc2, Tomasz Stanisz, Msc, PhD1, Emilie Clay, Msc, PhD3, Michal Pochopien, Msc, PhD1, Samuel Aballea, MSc, PhD4, Mondher Toumi, MSc, PhD, MD5.
1Clever-Access, Kraków, Poland, 2Clever-Access, Tunis, Tunisia, 3Clever-Access, Paris, France, 4InovIntell, Rotterdam, Netherlands, 5Aix-Marseille University, Marseille, France.
1Clever-Access, Kraków, Poland, 2Clever-Access, Tunis, Tunisia, 3Clever-Access, Paris, France, 4InovIntell, Rotterdam, Netherlands, 5Aix-Marseille University, Marseille, France.
OBJECTIVES: Survival extrapolation from published Kaplan-Meier (KM) curves is widely used in health-technology assessment (HTA) when individual patient data are unavailable. However, digitisation of KM curves, reconstruction of pseudo-individual patient data (pIPD), and parametric extrapolation are time-consuming and may be affected by analyst-dependent variability. This study evaluated whether a general-purpose artificial intelligence (AI) agent (Claude, Anthropic) could perform this end-to-end workflow and reproduce results comparable to the conventional manual workflow.
METHODS: We analysed 30 trials (60 figures, 118 KM curves). Using Claude, KM curves were digitised, pIPD were reconstructed using the Guyot algorithm, and six standard parametric survival models were fitted and extrapolated over an 80-year horizon. AI-generated results were compared with a reference workflow consisting of manual KM curve digitization, pIPD reconstruction, and parametric survival modelling in R. Agreement was assessed using RMSE, MAE, and maximum vertical deviation (MVD, worst-point survival difference). We additionally compared fitted curves and concordance in best-fitting model selection (AIC, BIC).
RESULTS: AI-derived KM reconstructions closely matched the reference workflow, although visual review remained necessary. Median RMSE and MAE were 0.018 and 0.014 survival probability units, respectively. The median MVD was 0.040, indicating that AI-reconstructed curves typically remained within four percentage points of the reference throughout observed follow-up. Agreement remained high for parametric extrapolations (median RMSE 0.004; median MAE 0.001; median MVD 0.028), although divergence increased over longer extrapolation horizons. The largest discrepancies occurred for monochrome figures in which treatment arms were distinguished only by line style.
CONCLUSIONS: An AI agent was able to reproduce a standard KM digitisation and survival extrapolation workflow with mostly good agreement to a validated R-based approach. While results support the potential of AI to accelerate routine HTA modelling tasks, sensitivity in long-term extrapolation and model selection highlights the need for expert oversight, particularly when source figures have low visual discriminability.
METHODS: We analysed 30 trials (60 figures, 118 KM curves). Using Claude, KM curves were digitised, pIPD were reconstructed using the Guyot algorithm, and six standard parametric survival models were fitted and extrapolated over an 80-year horizon. AI-generated results were compared with a reference workflow consisting of manual KM curve digitization, pIPD reconstruction, and parametric survival modelling in R. Agreement was assessed using RMSE, MAE, and maximum vertical deviation (MVD, worst-point survival difference). We additionally compared fitted curves and concordance in best-fitting model selection (AIC, BIC).
RESULTS: AI-derived KM reconstructions closely matched the reference workflow, although visual review remained necessary. Median RMSE and MAE were 0.018 and 0.014 survival probability units, respectively. The median MVD was 0.040, indicating that AI-reconstructed curves typically remained within four percentage points of the reference throughout observed follow-up. Agreement remained high for parametric extrapolations (median RMSE 0.004; median MAE 0.001; median MVD 0.028), although divergence increased over longer extrapolation horizons. The largest discrepancies occurred for monochrome figures in which treatment arms were distinguished only by line style.
CONCLUSIONS: An AI agent was able to reproduce a standard KM digitisation and survival extrapolation workflow with mostly good agreement to a validated R-based approach. While results support the potential of AI to accelerate routine HTA modelling tasks, sensitivity in long-term extrapolation and model selection highlights the need for expert oversight, particularly when source figures have low visual discriminability.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
P36
Topic
Clinical Outcomes, Economic Evaluation, Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas, Oncology