A CONFIGURATION-DRIVEN RAG EVALUATION FRAMEWORK FOR HEALTH ECONOMICS AND OUTCOMES RESEARCH (HEOR)
Author(s)
Finlay McIntyre, PhD1, Ahmad Hecham Alani, PharmD2, Mackenzie Mills, PhD2, Diana Rebeca Acosta Focil, MD2, Panos Kanavos, BSc, MSc, PhD3, Sandhya Alagan, PhD2.
1Data Scientist, HTA-Hive, London, United Kingdom, 2HTA-Hive, London, United Kingdom, 3London School of Economics and Political Science, London, United Kingdom.
1Data Scientist, HTA-Hive, London, United Kingdom, 2HTA-Hive, London, United Kingdom, 3London School of Economics and Political Science, London, United Kingdom.
OBJECTIVES: Large language models (LLMs) are increasingly used for data extraction, evidence synthesis, and analysis within HEOR, but validated approaches for evaluating model accuracy and performance are lacking. This study develops a systematic Retrieval-Augmented Generation (RAG) evaluation framework for data synthesis applied to the context of HTA reports.
METHODS: A RAG system was designed specifically to extract nuanced and aggregate information from collections of publicly available HTA reports in the UK, Canada and Australia. The pipeline was built using OpenAI embeddings, PGVector for vector storage, VoyageAI reranking, and Gemini 2.5 Flash as the core LLM. The workflow followed a two-phase architecture: data collection and statistical scoring. During collection, structured extraction tasks of varied complexity were performed with targeted context injection, ensuring evaluation against predefined ground-truth HTA documents, and minimizing retrieval bias. RAG-generated answers were manually labelled with binary truthfulness and completeness scores, and a hallucination proxy metric tracked unsupported predictions lacking source grounding. Performance was quantified using confusion matrices and derived metrics, measuring both the identification of evidence existence and its impact on the HTA outcome.
RESULTS: Performance varied by task complexity. Simple evidence identification tasks (e.g. categorising types of evidence considered) achieved high truthfulness and completeness (F1 > 0.85). More nuanced evidence appraisal tasks, such as determining evidence acceptability, were more challenging, with F1 scores ranging from 0.65 to 0.80. Targeted context injection and advanced reranking significantly improved precision, while the hallucination proxy metric flagged unsupported claims and reduced false-positive risk.
CONCLUSIONS: A systematic evaluation framework is essential for safely integrating RAG into HEOR. Judging whether AI-generated answers are faithful and complete is inherently difficult, and without robust validation, changes to the retrieval strategy, prompts, or language model risk introducing undetected errors. Our configuration-driven approach supports continuous quality improvement, enabling more reliable and transparent use of LLMs for complex HTA evidence synthesis.
METHODS: A RAG system was designed specifically to extract nuanced and aggregate information from collections of publicly available HTA reports in the UK, Canada and Australia. The pipeline was built using OpenAI embeddings, PGVector for vector storage, VoyageAI reranking, and Gemini 2.5 Flash as the core LLM. The workflow followed a two-phase architecture: data collection and statistical scoring. During collection, structured extraction tasks of varied complexity were performed with targeted context injection, ensuring evaluation against predefined ground-truth HTA documents, and minimizing retrieval bias. RAG-generated answers were manually labelled with binary truthfulness and completeness scores, and a hallucination proxy metric tracked unsupported predictions lacking source grounding. Performance was quantified using confusion matrices and derived metrics, measuring both the identification of evidence existence and its impact on the HTA outcome.
RESULTS: Performance varied by task complexity. Simple evidence identification tasks (e.g. categorising types of evidence considered) achieved high truthfulness and completeness (F1 > 0.85). More nuanced evidence appraisal tasks, such as determining evidence acceptability, were more challenging, with F1 scores ranging from 0.65 to 0.80. Targeted context injection and advanced reranking significantly improved precision, while the hallucination proxy metric flagged unsupported claims and reduced false-positive risk.
CONCLUSIONS: A systematic evaluation framework is essential for safely integrating RAG into HEOR. Judging whether AI-generated answers are faithful and complete is inherently difficult, and without robust validation, changes to the retrieval strategy, prompts, or language model risk introducing undetected errors. Our configuration-driven approach supports continuous quality improvement, enabling more reliable and transparent use of LLMs for complex HTA evidence synthesis.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
HTA301
Topic
Health Technology Assessment, Methodological & Statistical Research
Topic Subcategory
Decision & Deliberative Processes
Disease
No Additional Disease & Conditions/Specialized Treatment Areas