GOVERNANCE AND SAFEGUARDS TO EVALUATE GENERATIVE AI TOOLS AT NICE
Author(s)
Raphael Sonabend-Friend1, Michael John Merchant, PhD2, Robert Willans, PhD2, Vandana Ayyar Gupta, PhD3, Monica Casey, BSc2, Emma McFarlane, PhD2, Pall Jonsson, BSc, PhD4.
1Scientific Adviser, NICE, United Kingdom, 2NICE, Manchester, United Kingdom, 3National Institute for Health and Care Excellence, London, United Kingdom, 4National Institute for Health and Care Excellence (NICE), Manchester, United Kingdom.
1Scientific Adviser, NICE, United Kingdom, 2NICE, Manchester, United Kingdom, 3National Institute for Health and Care Excellence, London, United Kingdom, 4National Institute for Health and Care Excellence (NICE), Manchester, United Kingdom.
OBJECTIVES: AI has the potential to accelerate health technology assessment (HTA) and guideline development. Adoption requires organisations to move from fully human processes towards workflows in which AI augments decision making. Robust validation is essential to ensure tools are safe, reliable, and suitable for high-stakes environments. Our objective is to describe the internal governance and validation guide developed at NICE to evaluate AI tools, particularly large language models.
METHODS: NICE have developed an internal validation guide supported by a Python-based testing infrastructure hosted within an environment independent from NICE production servers. The guide operates across three levels. First, testing progresses through defined phases from proof-of-concept to shadow testing against real NICE workflows. Second, experiments follow a validation protocol in which development and evaluation datasets are separated, success metrics are predefined, and experiments are executed and documented in reproducible environments using Jupyter notebooks. Third, model optimisation follows a structured iterative process, including prompt design, k-shot prompting, reasoning prompt engineering where appropriate, and parameter setting. Configurations are recorded in a structured format and evaluated using in-house software. Technical safeguards include system prompts to prevent misuse and limiting analyst interaction to only adjusting predefined settings.
RESULTS: Several use cases have been identified that are using this guide, with initial testing evaluating Committee Discussion and Interpretation of the Evidence. The supporting infrastructure includes automated testing, version control and tracking software issues, enabling transparency and technical robustness. The guide was independently reviewed by an external partner to focus on sociotechnical perspectives; in light of the review the guide was updated to include more detail on accountability, responsibility, and consideration for the use of a holdout dataset as part of a prompt test suite.
CONCLUSIONS: NICE's evaluation guide provides a structured approach for safely evaluating AI tools and may offer a model for organisations considering responsible adoption of AI-assisted methods.
METHODS: NICE have developed an internal validation guide supported by a Python-based testing infrastructure hosted within an environment independent from NICE production servers. The guide operates across three levels. First, testing progresses through defined phases from proof-of-concept to shadow testing against real NICE workflows. Second, experiments follow a validation protocol in which development and evaluation datasets are separated, success metrics are predefined, and experiments are executed and documented in reproducible environments using Jupyter notebooks. Third, model optimisation follows a structured iterative process, including prompt design, k-shot prompting, reasoning prompt engineering where appropriate, and parameter setting. Configurations are recorded in a structured format and evaluated using in-house software. Technical safeguards include system prompts to prevent misuse and limiting analyst interaction to only adjusting predefined settings.
RESULTS: Several use cases have been identified that are using this guide, with initial testing evaluating Committee Discussion and Interpretation of the Evidence. The supporting infrastructure includes automated testing, version control and tracking software issues, enabling transparency and technical robustness. The guide was independently reviewed by an external partner to focus on sociotechnical perspectives; in light of the review the guide was updated to include more detail on accountability, responsibility, and consideration for the use of a holdout dataset as part of a prompt test suite.
CONCLUSIONS: NICE's evaluation guide provides a structured approach for safely evaluating AI tools and may offer a model for organisations considering responsible adoption of AI-assisted methods.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
P34
Topic
Health Technology Assessment, Methodological & Statistical Research, Organizational Practices
Disease
No Additional Disease & Conditions/Specialized Treatment Areas