CAN AI-ASSISTED SLRS MEET HTA STANDARDS? EVIDENCE FROM A NICE SINGLE TECHNOLOGY APPRAISAL

Author(s)

Grace E. Fox, PhD1, Gavin W. Stewart, MSc2, Peter Morten, MBiochem3, Philippa Murphy, MSc2, Darsh Devani, MS1, Stephan N. Martin, MPH1, Bengt Liljas, PhD4, Hannah Harrington, BA2.
1OPEN Health, New York, NY, USA, 2AstraZeneca, London, United Kingdom, 3AstraZeneca, Cambridge, United Kingdom, 4AstraZeneca, Gaithersburg, MD, USA.
OBJECTIVES: Artificial intelligence (AI) is increasingly proposed to improve the efficiency of systematic literature reviews (SLRs), but real-world examples of acceptance by health technology assessment (HTA) bodies remain limited. This case study describes the methods and outcome of 2 AI-assisted SLRs submitted to NICE.
METHODS: Clinical and economic SLRs supporting the NICE single technology appraisal of perioperative durvalumab plus fluorouracil, leucovorin, oxaliplatin, and docetaxel for resectable gastric and gastro-esophageal junction cancer (MATTERHORN trial; TA1160) were updated in June 2025 using AI for title/abstract screening (DistillerSR DAISY), full-text review (DistillerSR SEE), and data extraction/critical appraisal (ChatGPT Pro v5.0, Agent mode). At full-text review, SEE generated suggested PICOS eligibility assessments, although human reviewers made all inclusion decisions. AI performance was evaluated against the adjudicated reference standard for screening and data extraction/critical appraisal, with screening sensitivity prioritised to avoid omissions. Methods were aligned with NICE’s position statement on using AI in evidence generation. At the decision-problem meeting and throughout the appraisal, the company engaged NICE proactively.
RESULTS: AI title/abstract screening achieved 100% sensitivity in both SLRs (specificity, 84.8%-87.4%; accuracy, 88.7%-92.6%; precision, 69.0%-85.0%). Data extraction was completely correct or satisfactory for 92.1%-94.3% of data points (hallucination, 3.2%-4.3%), qualifying as a second reviewer; critical appraisal concordance was 86%-94%. At clarification, NICE asked minor methodological questions and requested that all AI tools be reported using the RAISE reporting standard, whose publication post-dated the project’s start. The External Assessment Group concluded that the SLRs’ methods were appropriate and that AI-assisted study selection was in line with accepted practice, identifying data extraction/critical appraisal as areas for further methodological transparency. The final draft guidance raised no concerns regarding the SLRs’ methodology.
CONCLUSIONS: Both AI-assisted SLRs were accepted within a NICE appraisal. Proactive engagement, sensitivity-prioritised screening, consistent human oversight of AI outputs, and transparent RAISE-aligned reporting supported acceptance.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

HTA98

Topic

Health Technology Assessment, Methodological & Statistical Research, Study Approaches

Topic Subcategory

Decision & Deliberative Processes

Disease

Oncology

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×