REPRODUCIBILITY OF AI RECOMMENDATIONS DURING TITLE-ABSTRACT SCREENING: QUANTIFYING DECISION STOCHASTICITY
Author(s)
Shainki Sharma, MPharm1, Geetank Kamboj, MPharm1, Abhishek Malik, MSc2, Hemant Rathi, MSc2.
1Skyward Analytics, Gurugram, India, 2EasySLR, Gurugram, India.
1Skyward Analytics, Gurugram, India, 2EasySLR, Gurugram, India.
OBJECTIVES: To evaluate the reproducibility of title-abstract screening decisions by artificial intelligence (AI) across repeated runs using the same dataset using the OpenAI GPT-5.0 mini model.
METHODS: A previously completed systematic literature review dataset was used to investigate the reproducibility of title-abstract screening decisions by AI within the EasySLR™ platform. The same dataset was screened on five separate occasions using two distinct execution approaches. In the first approach, five screening runs were conducted consecutively. In the second approach, runs were conducted at four-hour intervals to reduce potential session-level memory effects. Concordance of screening decisions across runs, frequency of decision changes, and patterns of variability at the individual-record level were evaluated for both approaches. The analysis also examined records that exhibited one or more changes in decisions and transitions between inclusion and exclusion decisions.
RESULTS: A total of 1,523 records were screened. Under the first approach, 90.2% of records demonstrated complete concordance across all five runs, 5.1% showed agreement in four out of five runs, and the remaining 4.7% showed agreement in three out of five runs. Under the second approach, 90.9% of records demonstrated complete concordance across all five runs, 5.4% showed agreement in four out of five runs, and 3.6% showed agreement in three out of five runs. Notably, three of the five runs generated identical screening decisions for every record under both approaches. Bidirectional decision transitions were observed, including both inclusion-to-exclusion and exclusion-to-inclusion changes.
CONCLUSIONS: Title-abstract screening by AI demonstrated high reproducibility across repeated runs, indicating substantial decision stability. However, a small proportion of records exhibited variability in screening outcomes, suggesting stochastic effects that may influence individual inclusion and exclusion decisions. These findings highlight the importance of understanding decision variability when implementing AI-based screening workflows.
METHODS: A previously completed systematic literature review dataset was used to investigate the reproducibility of title-abstract screening decisions by AI within the EasySLR™ platform. The same dataset was screened on five separate occasions using two distinct execution approaches. In the first approach, five screening runs were conducted consecutively. In the second approach, runs were conducted at four-hour intervals to reduce potential session-level memory effects. Concordance of screening decisions across runs, frequency of decision changes, and patterns of variability at the individual-record level were evaluated for both approaches. The analysis also examined records that exhibited one or more changes in decisions and transitions between inclusion and exclusion decisions.
RESULTS: A total of 1,523 records were screened. Under the first approach, 90.2% of records demonstrated complete concordance across all five runs, 5.1% showed agreement in four out of five runs, and the remaining 4.7% showed agreement in three out of five runs. Under the second approach, 90.9% of records demonstrated complete concordance across all five runs, 5.4% showed agreement in four out of five runs, and 3.6% showed agreement in three out of five runs. Notably, three of the five runs generated identical screening decisions for every record under both approaches. Bidirectional decision transitions were observed, including both inclusion-to-exclusion and exclusion-to-inclusion changes.
CONCLUSIONS: Title-abstract screening by AI demonstrated high reproducibility across repeated runs, indicating substantial decision stability. However, a small proportion of records exhibited variability in screening outcomes, suggesting stochastic effects that may influence individual inclusion and exclusion decisions. These findings highlight the importance of understanding decision variability when implementing AI-based screening workflows.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR220
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas