DECISION-CONTEXT METADATA IMPROVES AI-GUIDED ECONOMIC MODEL-STRUCTURE TRIAGE: A 15-CASE NON-ONCOLOGY HTA VALIDATION
Author(s)
Ankit Dhaundiyal, MS1, Jagriti Prasad, MPH2, Tejesh S, MPH2, Mona Thangamma, MSc Health Economics2, Megha Tharad, PhD3.
1Evalueserve GmbH, Rheinbach, Germany, 2Evalueserve Pvt Ltd, Bengaluru, India, 3Evalueserve SEZ Pvt Ltd, Gurugram, India.
1Evalueserve GmbH, Rheinbach, Germany, 2Evalueserve Pvt Ltd, Bengaluru, India, 3Evalueserve SEZ Pvt Ltd, Gurugram, India.
OBJECTIVES: To evaluate whether structured HEOR decision-context metadata improves AI-guided health economic model-structure triage, defined as early model-family selection to focus economic modeller review, across non-oncology HTA scenarios.
METHODS: An exploratory blinded validation was conducted across 15 non-oncology NICE technology appraisal scenarios. Reference model structures were extracted from public NICE appraisal materials and mapped using a prespecified economic-modelling taxonomy. Broader HEOR acceptability was adjudicated using predefined rules when the selected approach was clinically and economically reasonable despite not exactly matching the reference structure. Three input strategies were compared: clinical vignettes alone; vignettes plus a generic cost-comparison guardrail; and vignettes plus structured decision-context metadata. Metadata captured pre-modelling inputs including economic question, comparator relationship, QALY requirement, cost-comparison plausibility, model-reuse feasibility, decision driver, and expected complexity, without revealing reference labels. Final HTA conclusions and reference structures were withheld during AI runs. Strict agreement required exact concordance between the AI-selected model family and reference structure.
RESULTS: The clinical-vignette-only workflow achieved 8/15 strict agreement (53.3%) and 13/15 broad acceptability (86.7%). Adding a generic guardrail did not improve performance: 7/15 strict agreement (46.7%) and 12/15 broad acceptability (80.0%). The structured decision-context workflow increased strict agreement to 14/15 (93.3%; Wilson 95% CI: 70.2%-98.8%, wide interval reflecting small sample size) and broad acceptability to 15/15 (100.0%). Both cost-comparison reference cases were correctly selected, interpreted directionally given the small number of cases.
CONCLUSIONS: Structured decision-context metadata improved AI-guided model-family selection in this controlled non-oncology validation. The workflow may help HEOR teams focus modeller review earlier by clarifying when cost-utility modelling, cost comparison, risk-equation, responder-based, or other structures are appropriate. Findings are limited by sample size, NICE-centric references, absence of a human-modeller baseline, and lack of inter-rater reliability or workflow-efficiency measurement. Broader prospective validation is warranted before dossier-critical use.
METHODS: An exploratory blinded validation was conducted across 15 non-oncology NICE technology appraisal scenarios. Reference model structures were extracted from public NICE appraisal materials and mapped using a prespecified economic-modelling taxonomy. Broader HEOR acceptability was adjudicated using predefined rules when the selected approach was clinically and economically reasonable despite not exactly matching the reference structure. Three input strategies were compared: clinical vignettes alone; vignettes plus a generic cost-comparison guardrail; and vignettes plus structured decision-context metadata. Metadata captured pre-modelling inputs including economic question, comparator relationship, QALY requirement, cost-comparison plausibility, model-reuse feasibility, decision driver, and expected complexity, without revealing reference labels. Final HTA conclusions and reference structures were withheld during AI runs. Strict agreement required exact concordance between the AI-selected model family and reference structure.
RESULTS: The clinical-vignette-only workflow achieved 8/15 strict agreement (53.3%) and 13/15 broad acceptability (86.7%). Adding a generic guardrail did not improve performance: 7/15 strict agreement (46.7%) and 12/15 broad acceptability (80.0%). The structured decision-context workflow increased strict agreement to 14/15 (93.3%; Wilson 95% CI: 70.2%-98.8%, wide interval reflecting small sample size) and broad acceptability to 15/15 (100.0%). Both cost-comparison reference cases were correctly selected, interpreted directionally given the small number of cases.
CONCLUSIONS: Structured decision-context metadata improved AI-guided model-family selection in this controlled non-oncology validation. The workflow may help HEOR teams focus modeller review earlier by clarifying when cost-utility modelling, cost comparison, risk-equation, responder-based, or other structures are appropriate. Findings are limited by sample size, NICE-centric references, absence of a human-modeller baseline, and lack of inter-rater reliability or workflow-efficiency measurement. Broader prospective validation is warranted before dossier-critical use.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
EE605
Topic
Economic Evaluation, Health Technology Assessment, Methodological & Statistical Research
Disease
No Additional Disease & Conditions/Specialized Treatment Areas