A CONFIGURATION-DRIVEN RAG EVALUATION FRAMEWORK FOR HEALTH ECONOMICS AND OUTCOMES RESEARCH (HEOR)

Author(s)

Finlay McIntyre, PhD1, Ahmad Hecham Alani, PharmD2, Mackenzie Mills, PhD2, Diana Rebeca Acosta Focil, MD2, Panos Kanavos, BSc, MSc, PhD3, Sandhya Alagan, PhD2.
1Data Scientist, HTA-Hive, London, United Kingdom, 2HTA-Hive, London, United Kingdom, 3London School of Economics and Political Science, London, United Kingdom.
OBJECTIVES: Large language models (LLMs) are increasingly used for data extraction, evidence synthesis, and analysis within HEOR, but validated approaches for evaluating model accuracy and performance are lacking. This study develops a systematic Retrieval-Augmented Generation (RAG) evaluation framework for data synthesis applied to the context of HTA reports.
METHODS: A RAG system was designed specifically to extract nuanced and aggregate information from collections of publicly available HTA reports in the UK, Canada and Australia. The pipeline was built using OpenAI embeddings, PGVector for vector storage, VoyageAI reranking, and Gemini 2.5 Flash as the core LLM. The workflow followed a two-phase architecture: data collection and statistical scoring. During collection, structured extraction tasks of varied complexity were performed with targeted context injection, ensuring evaluation against predefined ground-truth HTA documents, and minimizing retrieval bias. RAG-generated answers were manually labelled with binary truthfulness and completeness scores, and a hallucination proxy metric tracked unsupported predictions lacking source grounding. Performance was quantified using confusion matrices and derived metrics, measuring both the identification of evidence existence and its impact on the HTA outcome.
RESULTS: Performance varied by task complexity. Simple evidence identification tasks (e.g. categorising types of evidence considered) achieved high truthfulness and completeness (F1 > 0.85). More nuanced evidence appraisal tasks, such as determining evidence acceptability, were more challenging, with F1 scores ranging from 0.65 to 0.80. Targeted context injection and advanced reranking significantly improved precision, while the hallucination proxy metric flagged unsupported claims and reduced false-positive risk.
CONCLUSIONS: A systematic evaluation framework is essential for safely integrating RAG into HEOR. Judging whether AI-generated answers are faithful and complete is inherently difficult, and without robust validation, changes to the retrieval strategy, prompts, or language model risk introducing undetected errors. Our configuration-driven approach supports continuous quality improvement, enabling more reliable and transparent use of LLMs for complex HTA evidence synthesis.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

HTA301

Topic

Health Technology Assessment, Methodological & Statistical Research

Topic Subcategory

Decision & Deliberative Processes

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×