AUTOMATING R SHINY COST EFFECTIVENESS MODELS FOR STROKE USING A MULTI AGENT AI FRAMEWORK: DEVELOPMENT AND VALIDATION
Author(s)
Namita Tundia, MS, PhD1, Timon Schicht, MSc2, Schiffon L. Wong, MPH3, Sameer Mansoori, MSc4, Barinder Singh, RPh4, Shubhram Pandey, MSc4.
1EMD Serono, Billerica, MA, USA, 2Merck Healthcare KGaA, Darmstadt, Germany, 3Schiffon Wong Strategic Advisory, Greater Boston, MA, USA, 4Pharmacoevidence, Mohali, India.
1EMD Serono, Billerica, MA, USA, 2Merck Healthcare KGaA, Darmstadt, Germany, 3Schiffon Wong Strategic Advisory, Greater Boston, MA, USA, 4Pharmacoevidence, Mohali, India.
OBJECTIVES: Early-stage cost-effectiveness models are structurally complex, time consuming, requiring iterative analyses, making R a preferred platform for its efficiency and flexibility. This study develops and evaluates a multi-agent artificial intelligence (AI) framework to automate a de novo R Shiny cost-effectiveness model for stroke.
METHODS: A multi-agent large language model architecture was developed to automate the development of a R Shiny cost-effectiveness model. A super-agent parsed an economic model protocol, dividing it into domain-specific tasks for specialised sub-agents: model structure and Markov trace; input parameterisation and calculations; cost-effectiveness outputs (incremental cost-effectiveness ratio (ICER), cost/quality-adjusted life year (QALY), cost/life year gained); uncertainty analyses (deterministic, probabilistic sensitivity analysis, cost-effectiveness acceptability curves, economically justifiable price, willingness-to-pay threshold analysis), and Shiny-UI. The super-agent then synthesised all sub-agent codebases into a deployable R-Shiny model. A dedicated validation agent flagged discrepancies at each integration step for automated resolution. Final outputs were independently reviewed by two senior health economists against a manually developed reference model using a pre-defined specification-compliance checklist.
RESULTS: The framework fully automated R-Shiny model code generation. Structured code review against the economic model protocol demonstrated 95.8% specification concordance across model structure, parameterisation, and analytical outputs; ICER deviation was <0.4% from the manually developed reference model. Incremental costs, QALYs, and transition probabilities showed high concordance across deterministic and probabilistic sensitivity analysis. Uncertainty analyses produced consistent outputs, directionally aligned with base case results. Model development time was reduced by ~65% (21 days vs 60 days) compared to conventional manual workflow.
CONCLUSIONS: This study demonstrates that multi-agent AI can generate robust R Shiny cost-effectiveness models with high concordance to expert benchmarks, significantly reducing development time while preserving model integrity, parameterisation, and uncertainty analysis. These results establish AI-driven multi-agent modelling as a credible, scalable complement to expert-led modelling, accelerating manual workflows, enabling faster and more reliable decision-making in health technology assessment.
METHODS: A multi-agent large language model architecture was developed to automate the development of a R Shiny cost-effectiveness model. A super-agent parsed an economic model protocol, dividing it into domain-specific tasks for specialised sub-agents: model structure and Markov trace; input parameterisation and calculations; cost-effectiveness outputs (incremental cost-effectiveness ratio (ICER), cost/quality-adjusted life year (QALY), cost/life year gained); uncertainty analyses (deterministic, probabilistic sensitivity analysis, cost-effectiveness acceptability curves, economically justifiable price, willingness-to-pay threshold analysis), and Shiny-UI. The super-agent then synthesised all sub-agent codebases into a deployable R-Shiny model. A dedicated validation agent flagged discrepancies at each integration step for automated resolution. Final outputs were independently reviewed by two senior health economists against a manually developed reference model using a pre-defined specification-compliance checklist.
RESULTS: The framework fully automated R-Shiny model code generation. Structured code review against the economic model protocol demonstrated 95.8% specification concordance across model structure, parameterisation, and analytical outputs; ICER deviation was <0.4% from the manually developed reference model. Incremental costs, QALYs, and transition probabilities showed high concordance across deterministic and probabilistic sensitivity analysis. Uncertainty analyses produced consistent outputs, directionally aligned with base case results. Model development time was reduced by ~65% (21 days vs 60 days) compared to conventional manual workflow.
CONCLUSIONS: This study demonstrates that multi-agent AI can generate robust R Shiny cost-effectiveness models with high concordance to expert benchmarks, significantly reducing development time while preserving model integrity, parameterisation, and uncertainty analysis. These results establish AI-driven multi-agent modelling as a credible, scalable complement to expert-led modelling, accelerating manual workflows, enabling faster and more reliable decision-making in health technology assessment.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR192
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
Cardiovascular Disorders (including MI, Stroke, Circulatory), Neurological Disorders