DOES RETRIEVAL-AUGMENTED GENERATION IMPROVE ARTIFICIAL INTELLIGENCE (AI) WRITING QUALITY FOR HEALTH ECONOMICS AND OUTCOMES RESEARCH (HEOR) ENVIRONMENTS? A CONTROLLED MULTI-MODEL ASSESSMENT USING A CANADIAN HEALTH TECHNOLOGY ASSESSMENT (HTA) KNOWLEDGE...

Author(s)

Gabriel Tremblay, MSc, DPhil1, Sharada Harricharan, PharmD2.
1Untraceable AI, Levis, QC, Canada, 2Frontier HEOR, Toronto, ON, Canada.
OBJECTIVES: To quantify whether retrieval-augmented generation (RAG) paired with a Grounding Codex can improve the quality of payer-relevant content that is AI-generated, and to characterize benefits across dimensions.
METHODS: Four frontier large language models (Claude Opus 4.8, Claude Sonnet 4.6, GPT 5.5, Gemini 3.1 Pro) were prompted, with and without RAG, to develop a Canadian HTA strategy for sacituzumab govitecan in HR+/HER2- metastatic breast cancer in the second-line, based on existing third-line submission documents. RAG provides curated references instead of relying on training alone, combined with a Grounding Codex, a human-authored, reference-preserving layer of 200+ entries on Canadian HTA guidelines and core HEOR methods. Outputs were scored by a blinded panel of four AI judges (Opus 4.8, Claude Fable 5, GPT 5.5, GPT 5.4-thinking) comprising an AI Delphi, anchored to a human expert rating, and across weighted criteria spanning writing quality, market access, economic, clinical, comparative-effectiveness, burden-of-illness, completeness, and hallucination.
RESULTS: Incorporating RAG improved mean quality from 81.1 to 83.5/100 (+2.4), with improvements across all models. Gains were largest for GPT 5.5 (+4.1), followed by Sonnet 4.6 (+2.0), Gemini 3.1 Pro (+1.8), and smallest for Opus 4.8 (+1.5). Opus 4.8 scored highest overall (85.8 to 87.3), while Gemini 3.1 Pro remained lowest despite modest improvement (72.5 to 74.3).
Improvement varied by criterion. RAG delivered its strongest gains in writing quality (up to +10.0), comparative-effectiveness (up to +8.5), hallucination reduction, and clinical efficacy, however caused small declines in certain criteria rather than improving every dimension monotonically.
CONCLUSIONS: A domain-specific RAG paired with a Grounding Codex improved payer-relevant content in frontier AI models, with the largest gains in writing quality, evidence synthesis, and accuracy. Effects varied across criteria, therefore RAG benefits should not be assumed in regulated workflows. These findings support domain-grounded retrieval to enhance payer-relevant content and motivate dedicated guideline-aligned evaluation across indications and jurisdictions.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

HTA190

Topic

Health Technology Assessment, Methodological & Statistical Research, Organizational Practices

Topic Subcategory

Value Frameworks & Dossier Format

Disease

Oncology

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×