CAN AI GENERATE COMPREHENSIVE, ROBUST, AND EFFICIENT SEARCH STRATEGIES FOR SYSTEMATIC LITERATURE REVIEWS?
Author(s)
Surabhi Aggarwal, MPharm1, Geetank Kamboj, MPharm1, Abhishek Malik, MSc2, Hemant Rathi, MSc2.
1Skyward Analytics, Gurugram, India, 2EasySLR, Gurugram, India.
1Skyward Analytics, Gurugram, India, 2EasySLR, Gurugram, India.
OBJECTIVES: To validate artificial intelligence (AI)-generated search strategies using the EasySLR™ platform by evaluating their ability to retrieve studies included in previously completed systematic literature reviews (SLRs). A secondary objective was to assess search efficiency by comparing changes in citation volume relative to conventional human-developed search strategies.
METHODS: AI-generated search strategies were evaluated against three completed SLRs by assessing their ability to retrieve included studies using the PubMed database. For each SLR, a search strategy was generated using the OpenAI GPT-5.2 model. The model was provided with a previously validated research question, including all Population, Intervention, Comparator, Outcomes, and Study design (PICOS) elements. Search performance was assessed using recall, defined as the proportion of studies from predefined reference sets that were successfully retrieved. Human expert review was subsequently undertaken; no new search terms were added, and all refinements were limited to selecting or deselecting AI-suggested terms. Search efficiency was evaluated by comparing citation volumes and calculating the percentage reduction in the number of retrieved records relative to the corresponding human-developed search strategy.
RESULTS: Across all three SLRs, the AI-generated search strategies achieved 100% recall, successfully identifying all studies included in the reference sets. The initial AI-generated strategies reduced citation volume by 48%, 73%, and 72% in Projects 1, 2, and 3, respectively, while maintaining complete recall. Subsequent human review and refinement of the AI-suggested search terms further improved search efficiency. Human-guided refinement, limited to excluding AI-suggested search terms, reduced citation volume by 78%, 79%, and 87%, respectively, relative to the corresponding human-developed search strategies, without loss of recall.
CONCLUSIONS: AI-generated search strategies provided a sensitive starting point for evidence identification. These findings support the use of AI-generated search strategies as decision-support tools within human-centred SLR workflows. However, human oversight remains essential to refine search strategies, ensure conceptual relevance, and minimise the risk of over-optimisation.
METHODS: AI-generated search strategies were evaluated against three completed SLRs by assessing their ability to retrieve included studies using the PubMed database. For each SLR, a search strategy was generated using the OpenAI GPT-5.2 model. The model was provided with a previously validated research question, including all Population, Intervention, Comparator, Outcomes, and Study design (PICOS) elements. Search performance was assessed using recall, defined as the proportion of studies from predefined reference sets that were successfully retrieved. Human expert review was subsequently undertaken; no new search terms were added, and all refinements were limited to selecting or deselecting AI-suggested terms. Search efficiency was evaluated by comparing citation volumes and calculating the percentage reduction in the number of retrieved records relative to the corresponding human-developed search strategy.
RESULTS: Across all three SLRs, the AI-generated search strategies achieved 100% recall, successfully identifying all studies included in the reference sets. The initial AI-generated strategies reduced citation volume by 48%, 73%, and 72% in Projects 1, 2, and 3, respectively, while maintaining complete recall. Subsequent human review and refinement of the AI-suggested search terms further improved search efficiency. Human-guided refinement, limited to excluding AI-suggested search terms, reduced citation volume by 78%, 79%, and 87%, respectively, relative to the corresponding human-developed search strategies, without loss of recall.
CONCLUSIONS: AI-generated search strategies provided a sensitive starting point for evidence identification. These findings support the use of AI-generated search strategies as decision-support tools within human-centred SLR workflows. However, human oversight remains essential to refine search strategies, ensure conceptual relevance, and minimise the risk of over-optimisation.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR221
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas