EVALUATION OF ARTIFICIAL INTELLIGENCE TOOLS FOR THE CONSIDERATION OF AUTOMATION OR AUGMENTATION OF LITERATURE REVIEW PROCESSES AT NICE: A SCOPING REVIEW

Author(s)

Emma McFarlane, PhD1, Monica Casey, PhD2, Chris Carmona, PhD2, Raphael Sonabend-Friend3, David Andrew Nicholls, BSc2, Stephen Duffield, PhD, MD1, Aye Paing, MD2, Ahmed Yosef, MSc2, Michael John Merchant, PhD1, Pall Jonsson, BSc, PhD1, Vanessa D. Nunes, MSc2.
1NICE, Manchester, United Kingdom, 2NICE, London, United Kingdom, 3Scientific Adviser, NICE, United Kingdom.
OBJECTIVES: The National Institute for Health and Care Excellence (NICE) uses literature reviews to inform health technology assessment (HTA) and guideline development. Artificial intelligence (AI) may enhance the efficiency, consistency, and scalability of literature reviews. This study focuses on producing actionable insights into the implementation of AI in literature review, advancing beyond existing systematic reviews of task‑specific performance to guide workflow‑level evaluation and real-world adoption.
METHODS: This review focused on publicly accessible AI tools built on foundation large language models (e.g., ChatGPT, Claude), published since 2024 and identified through MEDLINE, forward and backward citation searching, and reference-list screening of studies included in literature reviews. Studies evaluating fine-tuned and obsolete models were excluded. AI use was categorised by search, title/abstract screening, full-text screening, data extraction, risk-of-bias assessment, and analysis. Outcomes included performance measures, the role of AI (e.g., augmentation, replacement), and implementation approaches (e.g., API, web-based). Findings were synthesised using narrative summaries.
RESULTS: A total of 107 studies were identified, evaluating AI use in searching (N=13), title/abstract screening (N=37), full-text screening (N=9), data extraction (N=27), risk-of-bias assessment (N=18), and analysis (N=3). AI-assisted title/abstract screening could reduce reviewer workload while maintaining consistently high sensitivity. Tools for full-text screening, data extraction, and risk-of-bias assessment demonstrated inconsistent performance, indicating the need for more thorough human oversight. Implementation varied across tasks, with mixed interface use, heterogeneous prompting practices, and adequate reporting of AI systems but inconsistent model-version detail.
CONCLUSIONS: AI has the potential to support literature review in HTA, particularly in screening. This work supports real-world adoption of AI in HTA bodies by clarifying where AI might already be viable. We additionally identified requirements for standardisation and alignment with existing processes and methodology prior to sector adoption.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR164

Topic

Health Technology Assessment, Methodological & Statistical Research, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×