DEVELOPMENT OF A LOCAL NEURAL NETWORK-ASSISTED WORKFLOW FOR DETERMINISTIC RECONSTRUCTION OF MULTI-ARM KAPLAN-MEIER CURVES
Author(s)
Tom Ward1, Alex van Doornewaard, MSc2, Oliver Darlington, MSc2.
1Director, Horizon Health Economics Ltd., Poole, United Kingdom, 2Horizon Health Economics Ltd., Poole, United Kingdom.
1Director, Horizon Health Economics Ltd., Poole, United Kingdom, 2Horizon Health Economics Ltd., Poole, United Kingdom.
OBJECTIVES: Manual digitisation of Kaplan-Meier (KM) curves remains common in survival evidence synthesis but is time-consuming, difficult to audit, and subject to analyst variability. We developed a local neural-segmentation workflow in Python for deterministic KM reconstruction without generative AI.
METHODS: Image and text heuristics identify KM figures, which are then processed by two de novo U-Net-style convolutional neural networks. A plot-anatomy detector identifies and segments 21 image regions, including KM and non-KM plot areas, axes, tick labels, axis titles, labels, legends, risk tables, captions, and boundaries. A K-query multitask U-Net, where K denotes curve count, predicts shared curve, curve-junction, axis and plot-region masks, direct arm-specific masks, and curve-instance embeddings to separate overlapping arms. Models were trained on complex, variable, synthetic publication-like KM figures with exact ground-truth masks (50,000 training and 5,000 validation images). Figures varied in size, aspect ratio, arm count, censor marks, confidence intervals, risk tables, legends, gridlines, axis spacing, contrast, colouring, compression, layout, and journal-page context. Post-processing maps K-query arm masks and segmented plot pixels to time-survival coordinates using plot geometry, axis evidence, plot labels, and numbers-at-risk constraints. Predictions were assessed using Dice similarity coefficients, measuring overlap between predictions and KM ground-truths (0: no correspondence; 1: perfect correspondence).
RESULTS: The anatomy detector achieved a mean validation Dice of 0.924. For the K-query model, the curve, axis, plot region, curve-junction, and matched arm Dice scores were 0.967, 0.971, 0.980, 0.926, and 0.901, respectively. Reconstruction against known ground-truths showed arm count accuracy of 98.1%, arm-matching accuracy of 90.5% and a mean KM curve coverage ratio, measuring how far the extracted KM curve extends compared to the true curve, of 0.909.
CONCLUSIONS: This neural network-assisted workflow enables deterministic, auditable KM curve reconstruction from published studies. The approach may improve accuracy, reproducibility and scalability of survival evidence synthesis for HEOR without reliance on generative AI models.
METHODS: Image and text heuristics identify KM figures, which are then processed by two de novo U-Net-style convolutional neural networks. A plot-anatomy detector identifies and segments 21 image regions, including KM and non-KM plot areas, axes, tick labels, axis titles, labels, legends, risk tables, captions, and boundaries. A K-query multitask U-Net, where K denotes curve count, predicts shared curve, curve-junction, axis and plot-region masks, direct arm-specific masks, and curve-instance embeddings to separate overlapping arms. Models were trained on complex, variable, synthetic publication-like KM figures with exact ground-truth masks (50,000 training and 5,000 validation images). Figures varied in size, aspect ratio, arm count, censor marks, confidence intervals, risk tables, legends, gridlines, axis spacing, contrast, colouring, compression, layout, and journal-page context. Post-processing maps K-query arm masks and segmented plot pixels to time-survival coordinates using plot geometry, axis evidence, plot labels, and numbers-at-risk constraints. Predictions were assessed using Dice similarity coefficients, measuring overlap between predictions and KM ground-truths (0: no correspondence; 1: perfect correspondence).
RESULTS: The anatomy detector achieved a mean validation Dice of 0.924. For the K-query model, the curve, axis, plot region, curve-junction, and matched arm Dice scores were 0.967, 0.971, 0.980, 0.926, and 0.901, respectively. Reconstruction against known ground-truths showed arm count accuracy of 98.1%, arm-matching accuracy of 90.5% and a mean KM curve coverage ratio, measuring how far the extracted KM curve extends compared to the true curve, of 0.909.
CONCLUSIONS: This neural network-assisted workflow enables deterministic, auditable KM curve reconstruction from published studies. The approach may improve accuracy, reproducibility and scalability of survival evidence synthesis for HEOR without reliance on generative AI models.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR180
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas, Oncology