DEVELOPMENT AND VALIDATION OF A VARIABLE-BASED AI PROMPT FRAMEWORK FOR AUTOMATED GENERATION OF AMNOG DOSSIER SECTIONS: A MULTI-DRUG SOURCE FIDELITY EVALUATION

Author(s)

Jonas Jost1, Lutz Michael Vollmer2, Monika Ficek, MSc3, Gunter Ladinek4, Daniela Simon, Dr.5, Jelle Luijten, BSc6, Stefan Walzer, MA, PhD7, Stefanie Maxion-Bergemann, Sr., Dr.8.
1MArS Market Access & Pricing Strategy GmbH, Weil am Rhein, Germany, 2MArS Market Access & Pricing Strategy GmbH, Tuebingen, Germany, 3F. Hoffmann-La Roche AG, Szczecin, Poland, 4Roche Pharma AG, Grenzach-Wyhlen, Germany, 5Roche Pharma AG, Grenzach, Switzerland, 6Erasmus School of Health Policy & Management, Rotterdam, Netherlands, 7MArS Market Access & Pricing Strategy GmbH, Germany, 8Hoffmann La Roche, Basel, Switzerland.
OBJECTIVES: German AMNOG (Act on the Reform of the Market for Medicinal Products) benefit assessment dossiers demand evidence-grounded medical writing under tight timelines. We developed a structured, variable-based AI prompt framework for automated generation of AMNOG dossier sections and evaluated its output quality, source fidelity, and prompt transferability across therapeutic areas.
METHODS: Using the large language model (LLM) Google Gemini, a variable-based prompt library mapped AMNOG template sections across Modules 2 and 3 to dedicated prompts with reusable placeholders. Prompts were applied to two active ingredients from distinct therapeutic areas (ophthalmology; neurology), including data-dependent epidemiology sections requiring external prevalence modelling. The LLM generated text exclusively from primary source documents; no pre-synthesised summaries or previously compiled dossiers were used as input. A senior HTA medical writer assessed output quality against 14 pre-specified criteria spanning clarity, coherence, conceptual completeness, regulatory terminology, and structural compliance. Source fidelity was assessed by independent claim-level verification against cited references.
RESULTS: 35 prompts were developed across Modules 2 and 3, producing complete sections or sub-components. Fourteen outputs were evaluated: overall pass rate 92.9% (13/14), with 100% pass rates across all other criteria. One disease-description output was rated indeterminate on accuracy/referencing, as some statements were unverifiable due to non-identifiable AI-generated references; this resolved for the next drug through prompt refinement. Across 64 claims in two sections, 96.9% (62/64) had full reference support, 2/64 partial, 0/64 unsupported. The data-dependent epidemiology section passed expert review after one cycle. Applied to the second therapeutic area, the library reduced refinement cycles from two to one, indicating transferability.
CONCLUSIONS: A structured, variable-based prompt framework can generate AMNOG dossier text with high quality and near-complete source fidelity. The reusable architecture reduces per-drug adaptation effort and enables time savings supporting scalable, auditable LLM-assisted HTA medical writing. Broader validation across additional modules, therapeutic areas, and LLM platforms is warranted.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

HTA223

Topic

Health Technology Assessment, Methodological & Statistical Research

Topic Subcategory

Value Frameworks & Dossier Format

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×