REDUCING HALLUCINATIONS IN HEOR APPLICATIONS THROUGH RETRIEVAL OPTIMIZATION AND CONTEXT ENGINEERING TECHNIQUES

Author(s)

Barinder Singh, RPh, Gagandeep Kaur, M Pharma, Shubhram Pandey, MSc, Rajdeep Kaur, PhD.
Pharmacoevidence Pvt. Ltd., Mohali, India.
OBJECTIVES: Generative artificial intelligence (GenAI) has significant potential to accelerate health economics and outcomes research (HEOR); however, hallucinated content, unsupported claims, and citation inaccuracies remain critical barriers to reliable GenAI adoption for evidence generation. This study compared outputs generated using a standalone large language model (LLM) and an LLM integrated with Retrieval-Augmented Generation (RAG), retrieval optimization, and context engineering to evaluate their comparative performance across hallucinations, factual accuracy and evidence traceability
METHODS: A structured proof-of-concept (PoC) compared two AI pipelines: (1) a standalone LLM generating responses from pre-trained knowledge and (2) an LLM integrated with RAG, retrieval optimization, and context engineering. The enhanced pipeline used curated HEOR publications as the knowledge source, automatically generated structured metadata during document ingestion, and applied metadata-aware semantic retrieval to identify relevant evidence before response generation. Outputs from both pipelines were independently evaluated by subject matter experts using predefined scoring framework
RESULTS: A total of 40 HEOR publications, including peer-reviewed journal articles, conference abstracts, and other evidence sources, were used to generate structured evidence reports. SME evaluation identified hallucinated content, unsupported claims, citation inaccuracies, and limited source traceability in responses generated by the standalone LLM. In contrast, no hallucinated content, unsupported claims, or fabricated citations were identified in outputs generated by the RAG-based framework. Minor refinements were limited to one or two relevant references not retrieved during evidence synthesis. SMEs agreed that reports generated by the RAG-based framework were accurate, relevant, fully traceable to the source publications, and represented a high-quality first draft requiring only targeted expert review and refinement
CONCLUSIONS: Reliable HEOR evidence generation depends on grounding AI-generated responses in source publications with complete citation and traceability. Compared with a standalone LLM, integrating RAG eliminated hallucinations. These findings support RAG-integrated GenAI as a reliable foundation for evidence synthesis, subject to continued expert oversight

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR214

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×