THE BEST OF BOTH WORLDS? HYBRID PROGRAMMATIC-AI WORKFLOWS FOR HTA-GRADE EVIDENCE GENERATION FOR TIME-TO-EVENT ANALYSIS INTERPRETATION AND REPORTING
Author(s)
Simone Rivolo, BSc, MSc, PhD1, Tereza Lanitis, MSc2, George Bungey, MSc3, Michalis Galanakis, MSc4, Ivan Houisse, MSc5, Jack Ishak, PhD6.
1Thermo Fisher Scientific, San Felice Segrate, Italy, 2Thermo Fisher Scientific, Limassol, Cyprus, 3Thermo Fisher Scientific, London, United Kingdom, 4Thermo Fisher Scientific, Chania, Greece, 5Thermo Fisher Scientific, Budapest, Hungary, 6Thermo Fisher Scientific, Waltham, MA, USA.
1Thermo Fisher Scientific, San Felice Segrate, Italy, 2Thermo Fisher Scientific, Limassol, Cyprus, 3Thermo Fisher Scientific, London, United Kingdom, 4Thermo Fisher Scientific, Chania, Greece, 5Thermo Fisher Scientific, Budapest, Hungary, 6Thermo Fisher Scientific, Waltham, MA, USA.
OBJECTIVES: Generative artificial intelligence (GenAI) can streamline research activities, including interpretation and reporting of statistical analyses for health technology assessment (HTA). This study demonstrates, through a time-to-event (TTE) survival analysis case study, how HTA-grade interpretation and technical report drafting can be achieved using a hybrid programmatic-AI workflow.
METHODS: The TTE interpretation and reporting workflow (TTERep-AI) combined a programmatic layer with a GenAI layer based on a frontier large language model. The programmatic layer used a templated report structure that automatically incorporated TTE analysis outputs, including tables and figures, across the following sections: Kaplan-Meier (KM) description, hazard-shape interpretation, proportional hazards (PH) assessment, fit statistics and visual model-fit assessments. The AI layer generated narrative descriptions and interpretations using section-specific prompts iteratively developed by subject matter experts (SMEs) on two datasets. The workflow was implemented in R and integrated with ChatGPT 5.2 through application programming interfaces. Two independent SMEs then evaluated the output across six datasets and rated report quality and time savings versus manual drafting using 5-point Likert scales.
RESULTS: TTERep-AI generated draft reports for SME review within minutes. The programmatic layer ensured reproducible handling of statistical outputs, including automated interpretation of fit statistics across parametric models. The AI layer provided relevant scientific summaries, interpretations and conclusions for each section. Across six datasets, SMEs reported minor edits were required for KM and hazard-shape sections (scores ≥4). PH assessments required intermediate-to-minor revisions (scores ≥3). Visual model-fit interpretation required intermediate-to-major revisions depending on dataset complexity. Estimated time savings specific to the AI layer were up to 80% across most sections, but below 30% for visual-fit interpretation. All draft reports required human editing for finalization.
CONCLUSIONS: Hybrid programmatic-AI workflows can enhance consistency, transparency, and efficiency while maintaining alignment with HTA reporting requirements. This approach bridges automation and scientific oversight, supporting HTA-grade reporting of TTE survival analyses.
METHODS: The TTE interpretation and reporting workflow (TTERep-AI) combined a programmatic layer with a GenAI layer based on a frontier large language model. The programmatic layer used a templated report structure that automatically incorporated TTE analysis outputs, including tables and figures, across the following sections: Kaplan-Meier (KM) description, hazard-shape interpretation, proportional hazards (PH) assessment, fit statistics and visual model-fit assessments. The AI layer generated narrative descriptions and interpretations using section-specific prompts iteratively developed by subject matter experts (SMEs) on two datasets. The workflow was implemented in R and integrated with ChatGPT 5.2 through application programming interfaces. Two independent SMEs then evaluated the output across six datasets and rated report quality and time savings versus manual drafting using 5-point Likert scales.
RESULTS: TTERep-AI generated draft reports for SME review within minutes. The programmatic layer ensured reproducible handling of statistical outputs, including automated interpretation of fit statistics across parametric models. The AI layer provided relevant scientific summaries, interpretations and conclusions for each section. Across six datasets, SMEs reported minor edits were required for KM and hazard-shape sections (scores ≥4). PH assessments required intermediate-to-minor revisions (scores ≥3). Visual model-fit interpretation required intermediate-to-major revisions depending on dataset complexity. Estimated time savings specific to the AI layer were up to 80% across most sections, but below 30% for visual-fit interpretation. All draft reports required human editing for finalization.
CONCLUSIONS: Hybrid programmatic-AI workflows can enhance consistency, transparency, and efficiency while maintaining alignment with HTA reporting requirements. This approach bridges automation and scientific oversight, supporting HTA-grade reporting of TTE survival analyses.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR232
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas