DEVELOPMENT AND VALIDATION OF A VARIABLE-BASED AI PROMPT FRAMEWORK FOR AUTOMATED GENERATION OF UK NICE DOSSIER SECTIONS: A MULTI-DRUG SOURCE FIDELITY EVALUATION
Author(s)
Jonas Jost1, Megan Lewis2, Stephen Bradley3, Monika Ficek, MSc4, Lutz Michael Vollmer5, Stefan Walzer, MA, PhD6, Stefanie Maxion-Bergemann, Sr., Dr.7.
1MArS Market Access & Pricing Strategy GmbH, Weil am Rhein, Germany, 2Roche Products Ltd, Welwyn Garden City, United Kingdom, 3Roche Products Ltd, Letchworth Garden City, United Kingdom, 4F. Hoffmann-La Roche AG, Basel, Switzerland, 5MArS Market Access & Pricing Strategy GmbH, Tuebingen, Germany, 6MArS Market Access & Pricing Strategy GmbH, Germany, 7Hoffmann La Roche, Basel, Switzerland.
1MArS Market Access & Pricing Strategy GmbH, Weil am Rhein, Germany, 2Roche Products Ltd, Welwyn Garden City, United Kingdom, 3Roche Products Ltd, Letchworth Garden City, United Kingdom, 4F. Hoffmann-La Roche AG, Basel, Switzerland, 5MArS Market Access & Pricing Strategy GmbH, Tuebingen, Germany, 6MArS Market Access & Pricing Strategy GmbH, Germany, 7Hoffmann La Roche, Basel, Switzerland.
OBJECTIVES: Technology appraisal submissions to the National Institute for Health and Care Excellence (NICE) require structured, evidence-grounded medical writing under demanding timelines. This study aimed to develop and validate a variable-based AI prompt framework for automated generation of draft NICE Single Technology Appraisal (STA) template sections and to evaluate output quality, source fidelity, and prompt transferability across indications.
METHODS: A variable-based prompt library was developed mapping each STA template section to dedicated prompts with reusable placeholders. Two active ingredients from two distinct haematological indications were tested across sections 1.2 (technology description) and 1.3 (disease background, clinical management, care pathway, unmet need). The large language model Google Gemini generated text exclusively from primary sources (clinical study reports, SmPCs, clinical guidelines). The prompts and nine resulting outputs were evaluated by the pharmaceutical company`s HTA leads against 14 pre-specified quality criteria covering clarity, accuracy, coherence, content coverage, and regulatory terminology.
RESULTS: Overall pass rate was 100% (9/9 outputs across all applicable criteria). Transfer to a second indication required minor adaptation only: five of seven section prompts were modified through variable-level changes, one new placeholder was introduced, and refinement cycles decreased from two to one, reducing framework development time by 49% (78 to 40 days). Reviewer analysis identified a recurring tendency of the model to draw inferential conclusions rather than strictly reproducing cited content, indicating citation-level quality gates as a priority optimisation target.
CONCLUSIONS: A variable-based AI prompt framework can generate NICE technology appraisal dossier content with high quality across distinct haematological indications; the reusable architecture enables rapid cross-indication transfer. The framework functions as an evidence compilation and drafting tool, consistent with NICE guidance that AI should augment rather than replace human involvement: clinical validation, narrative development, and regulatory judgement remain expert-led. Prospective validation across additional document sections and therapeutic areas is warranted.
METHODS: A variable-based prompt library was developed mapping each STA template section to dedicated prompts with reusable placeholders. Two active ingredients from two distinct haematological indications were tested across sections 1.2 (technology description) and 1.3 (disease background, clinical management, care pathway, unmet need). The large language model Google Gemini generated text exclusively from primary sources (clinical study reports, SmPCs, clinical guidelines). The prompts and nine resulting outputs were evaluated by the pharmaceutical company`s HTA leads against 14 pre-specified quality criteria covering clarity, accuracy, coherence, content coverage, and regulatory terminology.
RESULTS: Overall pass rate was 100% (9/9 outputs across all applicable criteria). Transfer to a second indication required minor adaptation only: five of seven section prompts were modified through variable-level changes, one new placeholder was introduced, and refinement cycles decreased from two to one, reducing framework development time by 49% (78 to 40 days). Reviewer analysis identified a recurring tendency of the model to draw inferential conclusions rather than strictly reproducing cited content, indicating citation-level quality gates as a priority optimisation target.
CONCLUSIONS: A variable-based AI prompt framework can generate NICE technology appraisal dossier content with high quality across distinct haematological indications; the reusable architecture enables rapid cross-indication transfer. The framework functions as an evidence compilation and drafting tool, consistent with NICE guidance that AI should augment rather than replace human involvement: clinical validation, narrative development, and regulatory judgement remain expert-led. Prospective validation across additional document sections and therapeutic areas is warranted.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
HTA140
Topic
Health Technology Assessment, Methodological & Statistical Research
Topic Subcategory
Value Frameworks & Dossier Format
Disease
No Additional Disease & Conditions/Specialized Treatment Areas