COMPARING ARTIFICIAL INTELLIGENCE SCREENING TOOLS IN HEALTH ECONOMIC SYSTEMATIC REVIEWS

Author(s)

Lotte K. Staal, MSc1, Sopany Saing, BSc, MPH, PhD1, Olivier Manintveld, PhD, MD2, Erik Koffijberg, MSc, PhD1.
1University of Twente, Enschede, Netherlands, 2Erasmus MC, University Medical Center Rotterdam, Rotterdamn, Netherlands.
OBJECTIVES: Systematic reviews involve screening of large volumes of literature, making the process labor intensive and time-consuming. Modern Artificial Intelligence (AI) tools have enabled automated or semi-automated screening. While increasingly used, the comparative performance of such AI tools in systematic reviews of health economic evaluations (SR-HEEs) has not yet been studied. This study evaluates several open-source AI screening tools using previously completed SR-HEEs and assesses whether the choice of AI tool could affect review outcomes and subsequent decision-making in healthcare.
METHODS: ASReview, Catchii, DoCTER, RobotAnalyst and SWIFT-review were evaluated using three completed SR-HEEs. Comparative performance was assessed on (i) screening burden, defined as the proportion of articles requiring manual screening, (ii) accuracy in identifying studies included in the original reviews, and (iii) concordance between AI tools regarding missed articles. Additionally, the effect of tool settings (prior knowledge and stopping rules) was evaluated.
RESULTS: All evaluated AI screening tools reduced manual screening burden while generally retaining high accuracy, although performance differed considerably across tools and settings. ASReview consistently achieved high accuracy (91%-100%) with a low screening burden (11%-45%) whereas DoCTER and RobotAnalyst required screening a larger proportion of articles (DoCTER 27%-58%; RobotAnalyst 11%-66%) to reach similar accuracy (DoCTER 98%-100%; RobotAnalyst 92%-100%). The performance of SWIFT-review and Catchii was more dependent on tool settings, showing greater variability in screening burden (SWIFT 12%-100%; Catchii 3%-32%) and accuracy (SWIFT 75%-100%; Catchii 50%-97%). Providing additional prior knowledge and applying more conservative stopping rules improved overall accuracy in all tools.
CONCLUSIONS: Our findings demonstrate substantial variation in performance among AI screening tools, suggesting that both tool selection and setting influence which studies are identified in SR-HEEs. Consequently, review conclusions and healthcare decision-making may be affected by the chosen screening approach. Furthermore, maximizing screening performance may require careful optimization of tool settings and user input.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR170

Topic

Economic Evaluation, Methodological & Statistical Research, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×