DEVELOPING AND EVALUATING AN AI-ASSISTED AGENTIC PIPELINE FOR EXCEL-TO-R MIGRATION OF HEALTH ECONOMIC MODELS
Author(s)
Varyam Memon, MSc, BSc.
Health Economics Team, Roche Diagnostics UK & Ireland, Burgess Hill, United Kingdom.
Health Economics Team, Roche Diagnostics UK & Ireland, Burgess Hill, United Kingdom.
OBJECTIVES: Migrating health economic models from Excel to R improves long-term speed, reproducibility, and scalability; validating R methods against established Excel logic is a critical step toward full migration. We developed and evaluated an AI-assisted agentic pipeline for rebuilding these models in R, assessing numerical fidelity and reporting completeness against the ISPOR ELEVATE-GenAI framework
METHODS: A four-stage Claude Code agent pipeline executed Excel extraction, modular R rebuilding, cell-level numerical validation, and automated compliance checking (28-item CHEERS 2022 and 10-domain ELEVATE-GenAI) alongside analyst review. The pipeline flags discrepancies instead of fixing them silently; each build iterated against Excel-derived targets until cells matched, with the reusable pipeline logic refined across successive models. We tested four Excel model archetypes: a decision-tree cost-analysis; two 5-year budget-impact analyses; and a 40-year semi-Markov cost-utility model featuring PSA, OWSA, severity modifiers, and 14-pathway blended-comparator replication..
RESULTS:
Numerical fidelity (cell-level agreement within ≤0.01% relative tolerance) was achieved in 338 of 343 comparisons (98.5%). The remaining five deviated by ≤0.025% (Excel rounding or R precision). The pipeline's discrepancy log exposed a four-cell formula misreference in a model that had already passed human review which substantially altered the budget impact result once corrected. As the AI-assisted workflow learned from each build, ELEVATE-GenAI reporting scores rose from 18 to 26/30, highest for the model with purpose-built AI governance documentation, indicating remaining gaps are documentation, not method. CHEERS items met ranged from 16 to 23/28, with unmet items strictly procedural (abstract, HEAP, COI). Build time fell from two weeks by hand to two hours per model.
CONCLUSIONS: Structured agentic pipelines deliver rapid, faithful Excel-to-R migrations. Replication paired with discrepancy logging strengthens oversight rather than diluting it, surfacing errors that manual review missed. The human role shifts from line-by-line transcriber to strategic architect and validator, where domain expertise adds the most value
METHODS: A four-stage Claude Code agent pipeline executed Excel extraction, modular R rebuilding, cell-level numerical validation, and automated compliance checking (28-item CHEERS 2022 and 10-domain ELEVATE-GenAI) alongside analyst review. The pipeline flags discrepancies instead of fixing them silently; each build iterated against Excel-derived targets until cells matched, with the reusable pipeline logic refined across successive models. We tested four Excel model archetypes: a decision-tree cost-analysis; two 5-year budget-impact analyses; and a 40-year semi-Markov cost-utility model featuring PSA, OWSA, severity modifiers, and 14-pathway blended-comparator replication..
RESULTS:
Numerical fidelity (cell-level agreement within ≤0.01% relative tolerance) was achieved in 338 of 343 comparisons (98.5%). The remaining five deviated by ≤0.025% (Excel rounding or R precision). The pipeline's discrepancy log exposed a four-cell formula misreference in a model that had already passed human review which substantially altered the budget impact result once corrected. As the AI-assisted workflow learned from each build, ELEVATE-GenAI reporting scores rose from 18 to 26/30, highest for the model with purpose-built AI governance documentation, indicating remaining gaps are documentation, not method. CHEERS items met ranged from 16 to 23/28, with unmet items strictly procedural (abstract, HEAP, COI). Build time fell from two weeks by hand to two hours per model.
CONCLUSIONS: Structured agentic pipelines deliver rapid, faithful Excel-to-R migrations. Replication paired with discrepancy logging strengthens oversight rather than diluting it, surfacing errors that manual review missed. The human role shifts from line-by-line transcriber to strategic architect and validator, where domain expertise adds the most value
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR64
Topic
Economic Evaluation, Health Technology Assessment, Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas