CAN GENERATIVE AI RELIABLY IDENTIFY PAYER DECISIONS AND REPLACE MANUAL HTA REVIEW? A COMPARISON OF AI AND HUMAN EXTRACTION OF A FRENCH HTA REPORT
Author(s)
Rachel Beckerman, PhD, Charles Davis, BS, Mirella Dudzic, MPharm, James Horscroft, PhD, Csilla Kinyik-Merena, MSc, Sylwia Lach, MPH, Whitney Longstaff, MSc, Janice Stricker-Shaver, PhD, Roisin Wynne, MHE.
Maple Health Group LLC, New York, NY, USA.
Maple Health Group LLC, New York, NY, USA.
OBJECTIVES: Despite growing interest in applying GenAI to analyse health technology assessment (HTA) reports which inform market access decision-making, its ability to accurately identify payer decision drivers and the rationale underpinning these decisions remains uncertain.
METHODS: We assessed the performance of three GenAI models (ChatGPT 5.5 Pro, Claude Opus 4.8 Max, and Gemini 3.1 Pro extended) in extracting key payer insights from the French HTA of valoctocogene roxaparvovec in severe haemophilia A using a standardized extraction prompt run twice per model to assess reproducibility. Outputs were assessed against a human reference extraction. Evaluation focused on the accuracy of HTA recommendations, submitted evidence, and identification of decision drivers and their associated rationale. Both qualitative and quantitative differences between GenAI and human extraction approaches were documented.
RESULTS: Compared with human-only extraction, GenAI models offered substantial time savings; however, performance differed in the interpretation of payer reasoning. GenAI models were generally effective at identifying overall HTA conclusions but were less consistent in distinguishing supporting evidence from payer critique and in accurately attributing the rationale underlying positive and negative payer commentary. Variability was observed across models in the completeness and precision of extracted findings. Human review identified omissions, occasional misclassification of payer perspectives, loss of contextual detail, and some instances of hallucinations, all of which could influence the downstream interpretation of HTA outcomes. While GenAI substantially reduced extraction time, human-led extraction remained necessary to ensure accuracy and distinguish high-level conclusions from the specific rationale informing HTA decisions.
CONCLUSIONS: While future use of GenAI-assisted HTA extraction can accelerate evidence review, current models did not consistently capture the full depth of payer reasoning and decision-making context. Human validation remains essential to ensure accurate identification and interpretation of HTA commentary, decision drivers, and supporting rationale.
METHODS: We assessed the performance of three GenAI models (ChatGPT 5.5 Pro, Claude Opus 4.8 Max, and Gemini 3.1 Pro extended) in extracting key payer insights from the French HTA of valoctocogene roxaparvovec in severe haemophilia A using a standardized extraction prompt run twice per model to assess reproducibility. Outputs were assessed against a human reference extraction. Evaluation focused on the accuracy of HTA recommendations, submitted evidence, and identification of decision drivers and their associated rationale. Both qualitative and quantitative differences between GenAI and human extraction approaches were documented.
RESULTS: Compared with human-only extraction, GenAI models offered substantial time savings; however, performance differed in the interpretation of payer reasoning. GenAI models were generally effective at identifying overall HTA conclusions but were less consistent in distinguishing supporting evidence from payer critique and in accurately attributing the rationale underlying positive and negative payer commentary. Variability was observed across models in the completeness and precision of extracted findings. Human review identified omissions, occasional misclassification of payer perspectives, loss of contextual detail, and some instances of hallucinations, all of which could influence the downstream interpretation of HTA outcomes. While GenAI substantially reduced extraction time, human-led extraction remained necessary to ensure accuracy and distinguish high-level conclusions from the specific rationale informing HTA decisions.
CONCLUSIONS: While future use of GenAI-assisted HTA extraction can accelerate evidence review, current models did not consistently capture the full depth of payer reasoning and decision-making context. Human validation remains essential to ensure accurate identification and interpretation of HTA commentary, decision drivers, and supporting rationale.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
P33
Topic
Clinical Outcomes, Health Technology Assessment, Methodological & Statistical Research
Topic Subcategory
Decision & Deliberative Processes
Disease
Genetic, Regenerative & Curative Therapies, Rare & Orphan Diseases