SPECIALIZED RETRIEVAL-AUGMENTED GENERATION VERSUS A STANDALONE LARGE LANGUAGE MODEL FOR MARKET ACCESS QUESTIONS ON DRUG EVALUATION: A BENCHMARK AGAINST FRENCH TRANSPARENCY COMMITTEE OPINIONS
Author(s)
Ludovic Lamarsalle, MSc, PharmD1, Hugo De Oliveira, PhD2, Mona Rysak, MSc3, Martin PRODEL, PhD4.
1Founder, HEALSTRA, Lyon, France, 2Independent Researcher, Hamburg, Germany, 3DALI, Lyon, France, 4DALI, LYON, France.
1Founder, HEALSTRA, Lyon, France, 2Independent Researcher, Hamburg, Germany, 3DALI, Lyon, France, 4DALI, LYON, France.
OBJECTIVES: General-purpose large language models (LLMs) answer from parametric memory and may fabricate non-existent assessments or untraceable sources, unacceptable for market access, where every argument must be auditable. We tested whether a specialized retrieval-augmented generation (RAG) algorithm grounded in French Transparency Committee (TC) opinions outperforms the same LLM used alone for specialized drug-evaluation questions.
METHODS: We built a benchmark of 10 questions representative of market access practice (e.g., reimbursement rating, eligible population, cross-document counting). Reference answers and expected keywords were predefined as ground truth. A RAG pipeline combining metadata filtering, query rewriting, hybrid dense and sparse retrieval with reciprocal rank fusion, cross-encoder re-ranking, and removal of hallucinated URLs was connected to Mistral Medium 3.5 and queried over 6,000 TC opinions since January 2016. The same model without retrieval was the comparator. Each question ran 10 times per method (200 runs). An LLM judge scored adherence (percentage fidelity to the reference, 0-100), complemented by expected-keyword coverage.
RESULTS: Across 100 runs, the RAG system markedly outperformed the standalone LLM on mean adherence to reference answers (44.8% vs 7.6%; approximately six-fold) and retrieved substantially more expected key concepts (72.0% vs 53.5%). Performance peaked on targeted questions answerable from specific opinions (mean adherence around 71-76%). As anticipated, adherence was lower on questions requiring exhaustive counting across many opinions, a scenario deliberately included to probe a boundary outside the system's intended retrieval scope.
CONCLUSIONS: Beyond a near six-fold gain in fidelity, grounding the same LLM in TC opinions made answers verifiable: each statement links to an accessible source document, decisive for auditable market access and unattainable with a standalone LLM. The residual weakness on exhaustive, multi-opinion queries is addressable by an agentic design that queries document metadata (e.g., via SQL) before generation, extending precision beyond purely semantic retrieval.
METHODS: We built a benchmark of 10 questions representative of market access practice (e.g., reimbursement rating, eligible population, cross-document counting). Reference answers and expected keywords were predefined as ground truth. A RAG pipeline combining metadata filtering, query rewriting, hybrid dense and sparse retrieval with reciprocal rank fusion, cross-encoder re-ranking, and removal of hallucinated URLs was connected to Mistral Medium 3.5 and queried over 6,000 TC opinions since January 2016. The same model without retrieval was the comparator. Each question ran 10 times per method (200 runs). An LLM judge scored adherence (percentage fidelity to the reference, 0-100), complemented by expected-keyword coverage.
RESULTS: Across 100 runs, the RAG system markedly outperformed the standalone LLM on mean adherence to reference answers (44.8% vs 7.6%; approximately six-fold) and retrieved substantially more expected key concepts (72.0% vs 53.5%). Performance peaked on targeted questions answerable from specific opinions (mean adherence around 71-76%). As anticipated, adherence was lower on questions requiring exhaustive counting across many opinions, a scenario deliberately included to probe a boundary outside the system's intended retrieval scope.
CONCLUSIONS: Beyond a near six-fold gain in fidelity, grounding the same LLM in TC opinions made answers verifiable: each statement links to an accessible source document, decisive for auditable market access and unattainable with a standalone LLM. The residual weakness on exhaustive, multi-opinion queries is addressable by an agentic design that queries document metadata (e.g., via SQL) before generation, extending precision beyond purely semantic retrieval.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
HTA10
Topic
Health Technology Assessment, Methodological & Statistical Research
Topic Subcategory
Decision & Deliberative Processes