REPORTING OF ARTIFICIAL INTELLIGENCE (AI) USE IN NICE CLINICAL GUIDELINES AND TECHNOLOGY APPRAISALS: A STRUCTURED REVIEW

Author(s)

Elzbieta Olewinska, MSc1, Achraf Benayed, Engineer2, Beata Smela, PhD1, Emilie Clay, PhD3, Samuel Aballea, MSc, PhD4.
1Clever-Access, Cracow, Poland, 2Clever-Access, Tunis, Tunisia, 3Clever-Access, Paris, France, 4InovIntell, Paris, France.
OBJECTIVES: NICE’s Statement of Intent for Artificial Intelligence (November 2024) emphasizes the importance of transparency regarding the use of AI in health technology assessment. This study assessed AI reporting in NICE clinical guidelines (CGs) and technology appraisals (TAs), focusing on AI tools, intended use, human oversight, validation, performance, and limitations.
METHODS: A structured review of NICE documents published between January 2023 and 25 May 2026 was conducted. Eligible sources included CGs, published TAs, and TAs under development. For TAs, committee papers, company submissions, and ERG/EAG reports were reviewed. Documents were screened using predefined AI-related keywords. Data were extracted on AI tools, use, human oversight, validation, performance, and limitations. The review assessed reported rather than actual AI use.
RESULTS: Twenty-nine CGs, five published TAs, and three TAs under development were included. All CGs reported use of EPPI-Reviewer; however, 21 (72.4%) provided no information on whether AI-enabled functionalities were used. Eight CGs reported whether priority screening was used; seven reported using it and one reported not using it. Reported AI applications in CGs were limited to literature screening prioritisation. Published TAs reported AI applications in literature screening, data extraction, natural language processing, and treatment effect modifier identification. Reporting focused on intended use, with limited information on validation, performance, or limitations. All TAs under development reported human verification, validation, benchmarking against reviewers, performance metrics, and methodological limitations. Limited commentary was identified in ERG/EAG reports. One EAG questioned the use of unsupervised machine learning to define model health states.
CONCLUSIONS: AI-assisted methods were identified across NICE evidence reviews; however, reporting was inconsistent and often lacked information on validation, performance, and limitations. More detailed reporting was observed in appraisals under development; however, differences in document availability and the small number of appraisals preclude conclusions regarding reporting trends. Standardised reporting could improve transparency and support responsible AI adoption within HTA.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

SA104

Topic

Health Technology Assessment, Methodological & Statistical Research, Study Approaches

Topic Subcategory

Literature Review & Synthesis

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×