REDUCING HALLUCINATIONS IN HEOR APPLICATIONS THROUGH RETRIEVAL OPTIMIZATION AND CONTEXT ENGINEERING TECHNIQUES
Author(s)
Barinder Singh, RPh, Gagandeep Kaur, M Pharma, Shubhram Pandey, MSc, Rajdeep Kaur, PhD.
Pharmacoevidence Pvt. Ltd., Mohali, India.
Pharmacoevidence Pvt. Ltd., Mohali, India.
OBJECTIVES: Generative artificial intelligence (GenAI) has significant potential to accelerate health economics and outcomes research (HEOR); however, hallucinated content, unsupported claims, and citation inaccuracies remain critical barriers to reliable GenAI adoption for evidence generation. This study compared outputs generated using a standalone large language model (LLM) and an LLM integrated with Retrieval-Augmented Generation (RAG), retrieval optimization, and context engineering to evaluate their comparative performance across hallucinations, factual accuracy and evidence traceability
METHODS: A structured proof-of-concept (PoC) compared two AI pipelines: (1) a standalone LLM generating responses from pre-trained knowledge and (2) an LLM integrated with RAG, retrieval optimization, and context engineering. The enhanced pipeline used curated HEOR publications as the knowledge source, automatically generated structured metadata during document ingestion, and applied metadata-aware semantic retrieval to identify relevant evidence before response generation. Outputs from both pipelines were independently evaluated by subject matter experts using predefined scoring framework
RESULTS: A total of 40 HEOR publications, including peer-reviewed journal articles, conference abstracts, and other evidence sources, were used to generate structured evidence reports. SME evaluation identified hallucinated content, unsupported claims, citation inaccuracies, and limited source traceability in responses generated by the standalone LLM. In contrast, no hallucinated content, unsupported claims, or fabricated citations were identified in outputs generated by the RAG-based framework. Minor refinements were limited to one or two relevant references not retrieved during evidence synthesis. SMEs agreed that reports generated by the RAG-based framework were accurate, relevant, fully traceable to the source publications, and represented a high-quality first draft requiring only targeted expert review and refinement
CONCLUSIONS: Reliable HEOR evidence generation depends on grounding AI-generated responses in source publications with complete citation and traceability. Compared with a standalone LLM, integrating RAG eliminated hallucinations. These findings support RAG-integrated GenAI as a reliable foundation for evidence synthesis, subject to continued expert oversight
METHODS: A structured proof-of-concept (PoC) compared two AI pipelines: (1) a standalone LLM generating responses from pre-trained knowledge and (2) an LLM integrated with RAG, retrieval optimization, and context engineering. The enhanced pipeline used curated HEOR publications as the knowledge source, automatically generated structured metadata during document ingestion, and applied metadata-aware semantic retrieval to identify relevant evidence before response generation. Outputs from both pipelines were independently evaluated by subject matter experts using predefined scoring framework
RESULTS: A total of 40 HEOR publications, including peer-reviewed journal articles, conference abstracts, and other evidence sources, were used to generate structured evidence reports. SME evaluation identified hallucinated content, unsupported claims, citation inaccuracies, and limited source traceability in responses generated by the standalone LLM. In contrast, no hallucinated content, unsupported claims, or fabricated citations were identified in outputs generated by the RAG-based framework. Minor refinements were limited to one or two relevant references not retrieved during evidence synthesis. SMEs agreed that reports generated by the RAG-based framework were accurate, relevant, fully traceable to the source publications, and represented a high-quality first draft requiring only targeted expert review and refinement
CONCLUSIONS: Reliable HEOR evidence generation depends on grounding AI-generated responses in source publications with complete citation and traceability. Compared with a standalone LLM, integrating RAG eliminated hallucinations. These findings support RAG-integrated GenAI as a reliable foundation for evidence synthesis, subject to continued expert oversight
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR214
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas