A HYBRID AI-ASSISTED SCREENING FRAMEWORK FOR SYSTEMATIC LITERATURE REVIEWS IN ACUTE ISCHAEMIC STROKE-LARGE VESSEL OCCLUSION: EFFICIENCY AND ACCURACY
Author(s)
Namita Tundia, MS, PhD1, Timon Schicht, MSc2, Schiffon L. Wong, MPH3, Sumeet Attri, MPharm4, Shubhram Pandey, MSc4, Barinder Singh, RPh4.
1EMD Serono, Billerica, MA, USA, 2Merck Healthcare KGaA, Darmstadt, Germany, 3Schiffon Wong Strategic Advisory, Greater Boston, MA, USA, 4Pharmacoevidence, Mohali, India.
1EMD Serono, Billerica, MA, USA, 2Merck Healthcare KGaA, Darmstadt, Germany, 3Schiffon Wong Strategic Advisory, Greater Boston, MA, USA, 4Pharmacoevidence, Mohali, India.
OBJECTIVES: Systematic literature reviews (SLRs) are labour-intensive, foundational research, that require extensive manual screening. Building on the first NICE-accepted artificial intelligence (AI) assisted SLR, this study implements a hybrid AI-human review framework to evaluate accuracy, time, and efficiency across SLRs in acute ischaemic stroke with large vessel occlusion.
METHODS: EMBASE®, MEDLINE®, CENTRAL, and CDSR were searched from inception to August 2025, with inclusion/exclusion criteria guided by the PICOS framework. Citations were screened by a human reviewer and metaSLR AI tool (Claude Sonnet 3.7), with conflicts resolved by a subject matter expert. Performance of the framework was evaluated using accuracy and sensitivity metrics. The overall methodology was aligned with the AI-enabled SLR approach previously applied in a NICE submission by Makhija et al.
RESULTS: A total of 11,192 citations were screened using the hybrid framework, with an initial pilot of 100 citations used to optimize AI prompts. In title/abstract screening, the overall average AI-human agreement was 95.6%, ranging from 94% in healthcare resource utilization/cost review to 96% across economic evaluations, health utility, humanistic burden, and clinical reviews. Overall sensitivity was 93.8%, and quality assurance confirmed that no relevant citations were excluded by the hybrid workflow. The hybrid approach substantially improved efficiency, achieving nearly twice the speed of traditional manual review (14.5 days vs 29.5 days) while maintaining accuracy above the 95% benchmark, exceeding traditional dual human reviewer agreement. Full-text screening further demonstrated an accuracy of 92.8%.
CONCLUSIONS: Aligned with the NICE submission use case reported by Makhija et al., this study demonstrated an overall accuracy of >90%. Agreement across first- and second-stage screening exceeded that of the traditional dual human review approach, while delivering substantial efficiency gains and enhanced scalability for future HTA submissions. Notably, the AI-based approach reduced screening time and associated costs by approximately 50%, underscoring its potential as a reliable, resource-efficient alternative.
METHODS: EMBASE®, MEDLINE®, CENTRAL, and CDSR were searched from inception to August 2025, with inclusion/exclusion criteria guided by the PICOS framework. Citations were screened by a human reviewer and metaSLR AI tool (Claude Sonnet 3.7), with conflicts resolved by a subject matter expert. Performance of the framework was evaluated using accuracy and sensitivity metrics. The overall methodology was aligned with the AI-enabled SLR approach previously applied in a NICE submission by Makhija et al.
RESULTS: A total of 11,192 citations were screened using the hybrid framework, with an initial pilot of 100 citations used to optimize AI prompts. In title/abstract screening, the overall average AI-human agreement was 95.6%, ranging from 94% in healthcare resource utilization/cost review to 96% across economic evaluations, health utility, humanistic burden, and clinical reviews. Overall sensitivity was 93.8%, and quality assurance confirmed that no relevant citations were excluded by the hybrid workflow. The hybrid approach substantially improved efficiency, achieving nearly twice the speed of traditional manual review (14.5 days vs 29.5 days) while maintaining accuracy above the 95% benchmark, exceeding traditional dual human reviewer agreement. Full-text screening further demonstrated an accuracy of 92.8%.
CONCLUSIONS: Aligned with the NICE submission use case reported by Makhija et al., this study demonstrated an overall accuracy of >90%. Agreement across first- and second-stage screening exceeded that of the traditional dual human review approach, while delivering substantial efficiency gains and enhanced scalability for future HTA submissions. Notably, the AI-based approach reduced screening time and associated costs by approximately 50%, underscoring its potential as a reliable, resource-efficient alternative.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR249
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
Cardiovascular Disorders (including MI, Stroke, Circulatory), Neurological Disorders