A STRUCTURED AI CODE-REVIEW FRAMEWORK FOR VALIDATING R-BASED COST-EFFECTIVENESS MODELS: A HEAD-TO-HEAD COMPARISON WITH HACKATHON-STYLE HUMAN REVIEW
Author(s)
Benjamin P. Geisler, MPH, MD.
PhD Student, University of Oslo, Oslo, Norway.
PhD Student, University of Oslo, Oslo, Norway.
OBJECTIVES: Verification and validation of model-based economic evaluation (MBEE) is essential, yet depends on scarce expert time and is rarely reproducible. Agentic AI coding tools could provide scalable, consistent review, but their reliability for MBEE is unestablished. We aimed to develop a structured, transparent AI code-review framework for R-based MBEEs and evaluate it against hackathon-style human review, addressing whether an AI agent match experienced modelers in detecting errors and assessing coding-standard adherence?
METHODS: We encoded validation checklists (CADTH, TECH-Ver, AdViSHE, CHAOS) as machine-readable item specifications restricted to white-box (verification) and black-box (validation) testing. Deployed through an agentic assistant (Claude Code) on local repository clones, the framework executed each model, ran its tests, applied a standardized black-box battery (e.g., null treatment effect, degenerate transition probabilities, cost scaling, cohort conservation), and checked naming conventions (modified from the DARTH Working Group), modularity, documentation, and testthat coverage. It produced an evidence-cited pass/fail report and severity-tiered action list. The same repositories underwent independent review by human modelers in a hackathon. We compared approaches by issue category and severity, agreement against a consensus reference set with seeded errors (sensitivity, specificity, false positives), and reviewer time.
RESULTS: For each approach (human review pending), we will report the number and type of issues detected, those identified uniquely versus jointly, concordance with the reference set, false-positive rates, and time expended. We expect both approaches will detect issues the other missed. Full findings, with discordant cases, will be presented at the session.
CONCLUSIONS: We position AI-assisted review of MBEE as a fast, reproducible complement to (not a replacement for) human validation. A hybrid workflow in which AI performs first-pass triage and generates re-runnable tests, concentrating expert time on judgement-based validation, could be envisioned. To support adoption, we will share a public GitHub repository with the checklists and framework as a reusable Claude skill.
METHODS: We encoded validation checklists (CADTH, TECH-Ver, AdViSHE, CHAOS) as machine-readable item specifications restricted to white-box (verification) and black-box (validation) testing. Deployed through an agentic assistant (Claude Code) on local repository clones, the framework executed each model, ran its tests, applied a standardized black-box battery (e.g., null treatment effect, degenerate transition probabilities, cost scaling, cohort conservation), and checked naming conventions (modified from the DARTH Working Group), modularity, documentation, and testthat coverage. It produced an evidence-cited pass/fail report and severity-tiered action list. The same repositories underwent independent review by human modelers in a hackathon. We compared approaches by issue category and severity, agreement against a consensus reference set with seeded errors (sensitivity, specificity, false positives), and reviewer time.
RESULTS: For each approach (human review pending), we will report the number and type of issues detected, those identified uniquely versus jointly, concordance with the reference set, false-positive rates, and time expended. We expect both approaches will detect issues the other missed. Full findings, with discordant cases, will be presented at the session.
CONCLUSIONS: We position AI-assisted review of MBEE as a fast, reproducible complement to (not a replacement for) human validation. A hybrid workflow in which AI performs first-pass triage and generates re-runnable tests, concentrating expert time on judgement-based validation, could be envisioned. To support adoption, we will share a public GitHub repository with the checklists and framework as a reusable Claude skill.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
EE146
Topic
Economic Evaluation, Methodological & Statistical Research, Organizational Practices
Disease
No Additional Disease & Conditions/Specialized Treatment Areas