IMPACT OF IMAGE QUALITY ON AUTOMATED KAPLAN-MEIER CURVE DIGITISATION USING SYNTHETIC KNOWN-TRUTH SURVIVAL CURVES
Author(s)
Robin Philip, MSci , MSc, Nadine D. Younan, MSc, PhD, Necdet Gunsoy, MPH, PhD.
Evimed Solutions Ltd, Amersham, United Kingdom.
Evimed Solutions Ltd, Amersham, United Kingdom.
OBJECTIVES: Reconstructing patient-level data from published Kaplan-Meier (KM) curves is routine in health economics, but reliability depends on figure quality, a recognised but insufficiently quantified error source. Using synthetic KM curves with known truth, we explored how controlled image degradations affect automated digitisation accuracy.
METHODS: A synthetic two-arm KM truth was generated from Weibull models and held fixed (200 patients/arm, 36-month follow-up, 12/18-month control/treatment median survival, censoring at months 30-36, censor ticks rendered). The 300 DPI blue/red baseline generated 25 degraded images reflecting real-world acquisition: resolution loss, blur, JPEG compression, recompression, chroma subsampling and colour separability. Each image was processed three times using byte-identical inputs through a two-stage pipeline. First, a gpt-5-mini multimodal vision model read each plot image and extracted axis limits, curve count, and rotation. Second, SurvdigitizeR (2024) performed colour-based arm separation and tracing. Primary outcome was restricted mean survival time (RMST) error at 36 months. Runs required two validly labelled arms and at least 150 points/arm (baseline approximately 955 points/arm; failures 27-104 points).
RESULTS: Overall, 39 (of 75) runs met acceptance criteria and were scored; 36 (of 75) failed acceptance. The baseline produced a low mean absolute arm-level RMST error of 0.050 months, so no baseline adjustment was applied. Among scored conditions, error ranged from 0.030 months with chroma 4:2:0 subsampling to 0.291 months at 75 DPI. The largest scoreable treatment-effect distortion occurred under heaviest blur, with 0.451-month RMST-difference error. The dominant failure mode was arm-separation collapse across JPEG compression, recompression, and low-contrast colour. Across all factors and runs, the multimodal vision stage consistently identified axes, curve count, and rotation, indicating downstream algorithm-level failure.
CONCLUSIONS: Visually plausible KM digitisation can carry measurable error, while common image stresses can cause abrupt curve-separation failure. Because real-world KM images come from heterogeneous sources, reliable AI-assisted digitisation requires arm-separation methods less dependent on colour distinctness.
METHODS: A synthetic two-arm KM truth was generated from Weibull models and held fixed (200 patients/arm, 36-month follow-up, 12/18-month control/treatment median survival, censoring at months 30-36, censor ticks rendered). The 300 DPI blue/red baseline generated 25 degraded images reflecting real-world acquisition: resolution loss, blur, JPEG compression, recompression, chroma subsampling and colour separability. Each image was processed three times using byte-identical inputs through a two-stage pipeline. First, a gpt-5-mini multimodal vision model read each plot image and extracted axis limits, curve count, and rotation. Second, SurvdigitizeR (2024) performed colour-based arm separation and tracing. Primary outcome was restricted mean survival time (RMST) error at 36 months. Runs required two validly labelled arms and at least 150 points/arm (baseline approximately 955 points/arm; failures 27-104 points).
RESULTS: Overall, 39 (of 75) runs met acceptance criteria and were scored; 36 (of 75) failed acceptance. The baseline produced a low mean absolute arm-level RMST error of 0.050 months, so no baseline adjustment was applied. Among scored conditions, error ranged from 0.030 months with chroma 4:2:0 subsampling to 0.291 months at 75 DPI. The largest scoreable treatment-effect distortion occurred under heaviest blur, with 0.451-month RMST-difference error. The dominant failure mode was arm-separation collapse across JPEG compression, recompression, and low-contrast colour. Across all factors and runs, the multimodal vision stage consistently identified axes, curve count, and rotation, indicating downstream algorithm-level failure.
CONCLUSIONS: Visually plausible KM digitisation can carry measurable error, while common image stresses can cause abrupt curve-separation failure. Because real-world KM images come from heterogeneous sources, reliable AI-assisted digitisation requires arm-separation methods less dependent on colour distinctness.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR176
Topic
Methodological & Statistical Research, Real World Data & Information Systems, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas