READY FOR PRIME TIME? EVALUATING LARGE LANGUAGE MODELS FOR AUTOMATED EXTRACTION OF COST-EFFECTIVENESS MODELS FROM HTA SUBMISSIONS
Author(s)
Gábor Szabó, MSc1, Amy Pinsent, PhD2, Mariana Farraia, PhD3, Akvile Chapman, MSc2, Iga Piasecka, MSc4, Caroline von Wilamowitz-Moellendorff, PhD2, Simone Rivolo, PhD5.
1Thermo Fisher Scientific, Budapest, Hungary, 2Thermo Fisher Scientific, London, United Kingdom, 3Thermo Fischer Scientific, Ede, Netherlands, 4Thermo Fisher Scientific, Lódz, Poland, 5Thermo Fisher Scientific, Milan, Italy.
1Thermo Fisher Scientific, Budapest, Hungary, 2Thermo Fisher Scientific, London, United Kingdom, 3Thermo Fischer Scientific, Ede, Netherlands, 4Thermo Fisher Scientific, Lódz, Poland, 5Thermo Fisher Scientific, Milan, Italy.
OBJECTIVES: Large language models (LLMs) can automate extraction of health economic data (HED) from published technology assessments (TAs). Previous evaluations (ChatGPT-4) demonstrated suboptimal performance despite evaluating multiple prompting strategies. This study assessed whether a reasoning-enabled LLM, combined with domain-specific extraction guidance, can support accurate and automated extraction of HED from TAs.
METHODS: A custom GPT agent based on ChatGPT-5.5 (Thinking-Light) was developed to extract structured HED exclusively from uploaded TA documents, without access to external sources. The agent was guided by sixteen detailed domain-specific knowledge files, covering both simple domains (e.g. population, model structure) and advanced domains (e.g. committee critiques, modelling assumptions). For each domain, the agent reported the extracted information, supporting text excerpt, and source location to facilitate subject matter expert (SME) verification. Nine TAs across NICE (UK), CADTH/CDA (Canada), and ICER (US) were reviewed and extracted using the custom GPT agent. Outputs were compared against an SME reference standard and scored using domain-specific Likert scales (four-points for simple domains, five-point for advanced domains). Automated extraction and scoring of nine additional TAs are ongoing.
RESULTS: Each TA was fully extracted within minutes, compared with 60-90 minutes of SME extraction. The GPT-agent extractions demonstrated excellent agreement to SME extractions across simple domains (all extractions rated as correct or partially correct). Across advanced domains, more than 85% of extractions were rated as Good or Excellent, with only the committee critiques (55.6% Good/Excellent; 44.4% Fair) and modelling assumptions (66.7% Good/Excellent; 33.3% Fair/Poor) domains showing lower performance. No hallucinated extractions were identified.
CONCLUSIONS: A reasoning-enabled GPT, guided by domain-specific knowledge, achieved high agreement with SME reference extractions across most domains, while substantially reducing extraction time. These findings demonstrate the feasibility of AI-assisted, traceable extraction workflows for HTA evidence synthesis, with human validation remaining essential, especially for advanced domains.
METHODS: A custom GPT agent based on ChatGPT-5.5 (Thinking-Light) was developed to extract structured HED exclusively from uploaded TA documents, without access to external sources. The agent was guided by sixteen detailed domain-specific knowledge files, covering both simple domains (e.g. population, model structure) and advanced domains (e.g. committee critiques, modelling assumptions). For each domain, the agent reported the extracted information, supporting text excerpt, and source location to facilitate subject matter expert (SME) verification. Nine TAs across NICE (UK), CADTH/CDA (Canada), and ICER (US) were reviewed and extracted using the custom GPT agent. Outputs were compared against an SME reference standard and scored using domain-specific Likert scales (four-points for simple domains, five-point for advanced domains). Automated extraction and scoring of nine additional TAs are ongoing.
RESULTS: Each TA was fully extracted within minutes, compared with 60-90 minutes of SME extraction. The GPT-agent extractions demonstrated excellent agreement to SME extractions across simple domains (all extractions rated as correct or partially correct). Across advanced domains, more than 85% of extractions were rated as Good or Excellent, with only the committee critiques (55.6% Good/Excellent; 44.4% Fair) and modelling assumptions (66.7% Good/Excellent; 33.3% Fair/Poor) domains showing lower performance. No hallucinated extractions were identified.
CONCLUSIONS: A reasoning-enabled GPT, guided by domain-specific knowledge, achieved high agreement with SME reference extractions across most domains, while substantially reducing extraction time. These findings demonstrate the feasibility of AI-assisted, traceable extraction workflows for HTA evidence synthesis, with human validation remaining essential, especially for advanced domains.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
P20
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas