READY FOR PRIME TIME? EVALUATING LARGE LANGUAGE MODELS FOR AUTOMATED EXTRACTION OF COST-EFFECTIVENESS MODELS FROM HTA SUBMISSIONS

Author(s)

Gábor Szabó, MSc1, Amy Pinsent, PhD2, Mariana Farraia, PhD3, Akvile Chapman, MSc2, Iga Piasecka, MSc4, Caroline von Wilamowitz-Moellendorff, PhD2, Simone Rivolo, PhD5.
1Thermo Fisher Scientific, Budapest, Hungary, 2Thermo Fisher Scientific, London, United Kingdom, 3Thermo Fischer Scientific, Ede, Netherlands, 4Thermo Fisher Scientific, Lódz, Poland, 5Thermo Fisher Scientific, Milan, Italy.
OBJECTIVES: Large language models (LLMs) can automate extraction of health economic data (HED) from published technology assessments (TAs). Previous evaluations (ChatGPT-4) demonstrated suboptimal performance despite evaluating multiple prompting strategies. This study assessed whether a reasoning-enabled LLM, combined with domain-specific extraction guidance, can support accurate and automated extraction of HED from TAs.
METHODS: A custom GPT agent based on ChatGPT-5.5 (Thinking-Light) was developed to extract structured HED exclusively from uploaded TA documents, without access to external sources. The agent was guided by sixteen detailed domain-specific knowledge files, covering both simple domains (e.g. population, model structure) and advanced domains (e.g. committee critiques, modelling assumptions). For each domain, the agent reported the extracted information, supporting text excerpt, and source location to facilitate subject matter expert (SME) verification. Nine TAs across NICE (UK), CADTH/CDA (Canada), and ICER (US) were reviewed and extracted using the custom GPT agent. Outputs were compared against an SME reference standard and scored using domain-specific Likert scales (four-points for simple domains, five-point for advanced domains). Automated extraction and scoring of nine additional TAs are ongoing.
RESULTS: Each TA was fully extracted within minutes, compared with 60-90 minutes of SME extraction. The GPT-agent extractions demonstrated excellent agreement to SME extractions across simple domains (all extractions rated as correct or partially correct). Across advanced domains, more than 85% of extractions were rated as Good or Excellent, with only the committee critiques (55.6% Good/Excellent; 44.4% Fair) and modelling assumptions (66.7% Good/Excellent; 33.3% Fair/Poor) domains showing lower performance. No hallucinated extractions were identified.
CONCLUSIONS: A reasoning-enabled GPT, guided by domain-specific knowledge, achieved high agreement with SME reference extractions across most domains, while substantially reducing extraction time. These findings demonstrate the feasibility of AI-assisted, traceable extraction workflows for HTA evidence synthesis, with human validation remaining essential, especially for advanced domains.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

P20

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×