DOES RETRIEVAL-AUGMENTED GENERATION IMPROVE ARTIFICIAL INTELLIGENCE (AI) WRITING QUALITY FOR HEALTH ECONOMICS AND OUTCOMES RESEARCH (HEOR) ENVIRONMENTS? A CONTROLLED MULTI-MODEL ASSESSMENT USING A CANADIAN HEALTH TECHNOLOGY ASSESSMENT (HTA) KNOWLEDGE...
Author(s)
Gabriel Tremblay, MSc, DPhil1, Sharada Harricharan, PharmD2.
1Untraceable AI, Levis, QC, Canada, 2Frontier HEOR, Toronto, ON, Canada.
1Untraceable AI, Levis, QC, Canada, 2Frontier HEOR, Toronto, ON, Canada.
OBJECTIVES: To quantify whether retrieval-augmented generation (RAG) paired with a Grounding Codex can improve the quality of payer-relevant content that is AI-generated, and to characterize benefits across dimensions.
METHODS: Four frontier large language models (Claude Opus 4.8, Claude Sonnet 4.6, GPT 5.5, Gemini 3.1 Pro) were prompted, with and without RAG, to develop a Canadian HTA strategy for sacituzumab govitecan in HR+/HER2- metastatic breast cancer in the second-line, based on existing third-line submission documents. RAG provides curated references instead of relying on training alone, combined with a Grounding Codex, a human-authored, reference-preserving layer of 200+ entries on Canadian HTA guidelines and core HEOR methods. Outputs were scored by a blinded panel of four AI judges (Opus 4.8, Claude Fable 5, GPT 5.5, GPT 5.4-thinking) comprising an AI Delphi, anchored to a human expert rating, and across weighted criteria spanning writing quality, market access, economic, clinical, comparative-effectiveness, burden-of-illness, completeness, and hallucination.
RESULTS: Incorporating RAG improved mean quality from 81.1 to 83.5/100 (+2.4), with improvements across all models. Gains were largest for GPT 5.5 (+4.1), followed by Sonnet 4.6 (+2.0), Gemini 3.1 Pro (+1.8), and smallest for Opus 4.8 (+1.5). Opus 4.8 scored highest overall (85.8 to 87.3), while Gemini 3.1 Pro remained lowest despite modest improvement (72.5 to 74.3).
Improvement varied by criterion. RAG delivered its strongest gains in writing quality (up to +10.0), comparative-effectiveness (up to +8.5), hallucination reduction, and clinical efficacy, however caused small declines in certain criteria rather than improving every dimension monotonically.
CONCLUSIONS: A domain-specific RAG paired with a Grounding Codex improved payer-relevant content in frontier AI models, with the largest gains in writing quality, evidence synthesis, and accuracy. Effects varied across criteria, therefore RAG benefits should not be assumed in regulated workflows. These findings support domain-grounded retrieval to enhance payer-relevant content and motivate dedicated guideline-aligned evaluation across indications and jurisdictions.
METHODS: Four frontier large language models (Claude Opus 4.8, Claude Sonnet 4.6, GPT 5.5, Gemini 3.1 Pro) were prompted, with and without RAG, to develop a Canadian HTA strategy for sacituzumab govitecan in HR+/HER2- metastatic breast cancer in the second-line, based on existing third-line submission documents. RAG provides curated references instead of relying on training alone, combined with a Grounding Codex, a human-authored, reference-preserving layer of 200+ entries on Canadian HTA guidelines and core HEOR methods. Outputs were scored by a blinded panel of four AI judges (Opus 4.8, Claude Fable 5, GPT 5.5, GPT 5.4-thinking) comprising an AI Delphi, anchored to a human expert rating, and across weighted criteria spanning writing quality, market access, economic, clinical, comparative-effectiveness, burden-of-illness, completeness, and hallucination.
RESULTS: Incorporating RAG improved mean quality from 81.1 to 83.5/100 (+2.4), with improvements across all models. Gains were largest for GPT 5.5 (+4.1), followed by Sonnet 4.6 (+2.0), Gemini 3.1 Pro (+1.8), and smallest for Opus 4.8 (+1.5). Opus 4.8 scored highest overall (85.8 to 87.3), while Gemini 3.1 Pro remained lowest despite modest improvement (72.5 to 74.3).
Improvement varied by criterion. RAG delivered its strongest gains in writing quality (up to +10.0), comparative-effectiveness (up to +8.5), hallucination reduction, and clinical efficacy, however caused small declines in certain criteria rather than improving every dimension monotonically.
CONCLUSIONS: A domain-specific RAG paired with a Grounding Codex improved payer-relevant content in frontier AI models, with the largest gains in writing quality, evidence synthesis, and accuracy. Effects varied across criteria, therefore RAG benefits should not be assumed in regulated workflows. These findings support domain-grounded retrieval to enhance payer-relevant content and motivate dedicated guideline-aligned evaluation across indications and jurisdictions.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
HTA190
Topic
Health Technology Assessment, Methodological & Statistical Research, Organizational Practices
Topic Subcategory
Value Frameworks & Dossier Format
Disease
Oncology