SPECIALIZED RETRIEVAL-AUGMENTED GENERATION VERSUS A STANDALONE LARGE LANGUAGE MODEL FOR MARKET ACCESS QUESTIONS ON DRUG EVALUATION: A BENCHMARK AGAINST FRENCH TRANSPARENCY COMMITTEE OPINIONS

Author(s)

Ludovic Lamarsalle, MSc, PharmD1, Hugo De Oliveira, PhD2, Mona Rysak, MSc3, Martin PRODEL, PhD4.
1Founder, HEALSTRA, Lyon, France, 2Independent Researcher, Hamburg, Germany, 3DALI, Lyon, France, 4DALI, LYON, France.
OBJECTIVES: General-purpose large language models (LLMs) answer from parametric memory and may fabricate non-existent assessments or untraceable sources, unacceptable for market access, where every argument must be auditable. We tested whether a specialized retrieval-augmented generation (RAG) algorithm grounded in French Transparency Committee (TC) opinions outperforms the same LLM used alone for specialized drug-evaluation questions.
METHODS: We built a benchmark of 10 questions representative of market access practice (e.g., reimbursement rating, eligible population, cross-document counting). Reference answers and expected keywords were predefined as ground truth. A RAG pipeline combining metadata filtering, query rewriting, hybrid dense and sparse retrieval with reciprocal rank fusion, cross-encoder re-ranking, and removal of hallucinated URLs was connected to Mistral Medium 3.5 and queried over 6,000 TC opinions since January 2016. The same model without retrieval was the comparator. Each question ran 10 times per method (200 runs). An LLM judge scored adherence (percentage fidelity to the reference, 0-100), complemented by expected-keyword coverage.
RESULTS: Across 100 runs, the RAG system markedly outperformed the standalone LLM on mean adherence to reference answers (44.8% vs 7.6%; approximately six-fold) and retrieved substantially more expected key concepts (72.0% vs 53.5%). Performance peaked on targeted questions answerable from specific opinions (mean adherence around 71-76%). As anticipated, adherence was lower on questions requiring exhaustive counting across many opinions, a scenario deliberately included to probe a boundary outside the system's intended retrieval scope.
CONCLUSIONS: Beyond a near six-fold gain in fidelity, grounding the same LLM in TC opinions made answers verifiable: each statement links to an accessible source document, decisive for auditable market access and unattainable with a standalone LLM. The residual weakness on exhaustive, multi-opinion queries is addressable by an agentic design that queries document metadata (e.g., via SQL) before generation, extending precision beyond purely semantic retrieval.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

HTA10

Topic

Health Technology Assessment, Methodological & Statistical Research

Topic Subcategory

Decision & Deliberative Processes

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×