COMPARING ARTIFICIAL INTELLIGENCE SCREENING TOOLS IN HEALTH ECONOMIC SYSTEMATIC REVIEWS
Author(s)
Lotte K. Staal, MSc1, Sopany Saing, BSc, MPH, PhD1, Olivier Manintveld, PhD, MD2, Erik Koffijberg, MSc, PhD1.
1University of Twente, Enschede, Netherlands, 2Erasmus MC, University Medical Center Rotterdam, Rotterdamn, Netherlands.
1University of Twente, Enschede, Netherlands, 2Erasmus MC, University Medical Center Rotterdam, Rotterdamn, Netherlands.
OBJECTIVES: Systematic reviews involve screening of large volumes of literature, making the process labor intensive and time-consuming. Modern Artificial Intelligence (AI) tools have enabled automated or semi-automated screening. While increasingly used, the comparative performance of such AI tools in systematic reviews of health economic evaluations (SR-HEEs) has not yet been studied. This study evaluates several open-source AI screening tools using previously completed SR-HEEs and assesses whether the choice of AI tool could affect review outcomes and subsequent decision-making in healthcare.
METHODS: ASReview, Catchii, DoCTER, RobotAnalyst and SWIFT-review were evaluated using three completed SR-HEEs. Comparative performance was assessed on (i) screening burden, defined as the proportion of articles requiring manual screening, (ii) accuracy in identifying studies included in the original reviews, and (iii) concordance between AI tools regarding missed articles. Additionally, the effect of tool settings (prior knowledge and stopping rules) was evaluated.
RESULTS: All evaluated AI screening tools reduced manual screening burden while generally retaining high accuracy, although performance differed considerably across tools and settings. ASReview consistently achieved high accuracy (91%-100%) with a low screening burden (11%-45%) whereas DoCTER and RobotAnalyst required screening a larger proportion of articles (DoCTER 27%-58%; RobotAnalyst 11%-66%) to reach similar accuracy (DoCTER 98%-100%; RobotAnalyst 92%-100%). The performance of SWIFT-review and Catchii was more dependent on tool settings, showing greater variability in screening burden (SWIFT 12%-100%; Catchii 3%-32%) and accuracy (SWIFT 75%-100%; Catchii 50%-97%). Providing additional prior knowledge and applying more conservative stopping rules improved overall accuracy in all tools.
CONCLUSIONS: Our findings demonstrate substantial variation in performance among AI screening tools, suggesting that both tool selection and setting influence which studies are identified in SR-HEEs. Consequently, review conclusions and healthcare decision-making may be affected by the chosen screening approach. Furthermore, maximizing screening performance may require careful optimization of tool settings and user input.
METHODS: ASReview, Catchii, DoCTER, RobotAnalyst and SWIFT-review were evaluated using three completed SR-HEEs. Comparative performance was assessed on (i) screening burden, defined as the proportion of articles requiring manual screening, (ii) accuracy in identifying studies included in the original reviews, and (iii) concordance between AI tools regarding missed articles. Additionally, the effect of tool settings (prior knowledge and stopping rules) was evaluated.
RESULTS: All evaluated AI screening tools reduced manual screening burden while generally retaining high accuracy, although performance differed considerably across tools and settings. ASReview consistently achieved high accuracy (91%-100%) with a low screening burden (11%-45%) whereas DoCTER and RobotAnalyst required screening a larger proportion of articles (DoCTER 27%-58%; RobotAnalyst 11%-66%) to reach similar accuracy (DoCTER 98%-100%; RobotAnalyst 92%-100%). The performance of SWIFT-review and Catchii was more dependent on tool settings, showing greater variability in screening burden (SWIFT 12%-100%; Catchii 3%-32%) and accuracy (SWIFT 75%-100%; Catchii 50%-97%). Providing additional prior knowledge and applying more conservative stopping rules improved overall accuracy in all tools.
CONCLUSIONS: Our findings demonstrate substantial variation in performance among AI screening tools, suggesting that both tool selection and setting influence which studies are identified in SR-HEEs. Consequently, review conclusions and healthcare decision-making may be affected by the chosen screening approach. Furthermore, maximizing screening performance may require careful optimization of tool settings and user input.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR170
Topic
Economic Evaluation, Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas