PERFORMANCE OF ARTIFICIAL INTELLIGENCE (AI) AS A SECOND REVIEWER IN TITLE/ABSTRACT SCREENING: EVIDENCE FROM 30 HEALTH TECHNOLOGY ASSESSMENT-FOCUSED SYSTEMATIC LITERATURE REVIEWS (SLRS)

Author(s)

Theodora Oikonomidi, PhD1, Tanushree Chaudhary Pavithran, MSc2, Ines Guerra, MSc3, Ketevan Rtveladze, MSc3, Lirong Zhang, MSc4, Georgie Weston, PhD5.
1IQVIA, Athens, Greece, 2IQVIA, Trivandrum, India, 3IQVIA, London, United Kingdom, 4Daiichi-Sankyo, Munich, Germany, 5AstraZeneca Pharmaceuticals, London, United Kingdom.
OBJECTIVES: To evaluate the use of AI as a second reviewer in title/abstract screening across 30 SLRs in advanced/metastatic solid tumours.
METHODS: Screening was conducted on DistillerSR. Two human reviewers screened a proportion of retrieved abstracts per SLR, until ≥90% of predicted, total includes in the SLR had been identified, according to the DistillerSR algorithm. After development of the training set, the second human reviewer was replaced by AI. AI assigned a probability of inclusion to each remaining abstract (range: 0.0-1.0), which was dichotomised into a binary inclusion/exclusion decision, using a stepped approach. First, an exclusion threshold of ≤0.2 was used, which led to the exclusion of most records. For the remaining records, the probability threshold for exclusion was incrementally increased up to 0.6. Human/AI decision conflicts were resolved by a third reviewer.
RESULTS: 117,667 abstracts were retrieved across 30 SLRs (342-18,381 per SLR). The prespecified 90% threshold was reached in 18/30 SLRs; AI was not used in 12 SLRs where the threshold was not reached. Across 18 SLRs, AI was used to make screening decisions as reviewer 2 for 28,676 abstracts (range: 46-4,055 per SLR); this represented a maximum of 72% of total abstracts per SLR (median: 37.5%, interquartile range: 28.7%-51.0%, indicating heterogeneity). Reasons for variation included SLR size, type and indication. Greater AI use was observed in mid-to-large SLRs and in economic or clinical SLRs compared to utility SLRs. Limited AI deployment was observed in indications that included multiple subcategories (e.g. rare cancers). Regarding accuracy, ≥99% agreement was achieved in 16/18 SLRs between human and AI decisions; in the remaining SLRs there were 14 (9%) and 153 (14%) human/AI conflicts, respectively.
CONCLUSIONS: Findings suggest that AI can be used as a second reviewer (human-in-the-loop approach), without compromising accuracy. However, AI may not reduce effort in smaller and utility SLRs.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

HTA344

Topic

Health Technology Assessment, Methodological & Statistical Research

Disease

Oncology

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×