AN AGENTIC WORKFLOW FOR DEVELOPING HIGH-SENSITIVITY PUBMED SEARCH STRATEGIES FOR EVIDENCE SYNTHESIS
Author(s)
Artur Nowak, MSc, Ewelina Sadowska, MPharm, Ewa Borowiack, MSc, Monika Opalek, PhD.
Evidence Prime, Krakow, Poland.
Evidence Prime, Krakow, Poland.
OBJECTIVES: High-sensitivity PubMed strategies are critical for evidence synthesis but require iterative concept selection, vocabulary expansion, retrieval testing, and refinement. We developed an agentic large language model workflow to generate PubMed search strategies from review eligibility criteria alone.
METHODS: Eligibility criteria were extracted verbatim from published reviews and used as the only review-level input. To evaluate performance under a difficult, early-stage search-development setting, known included studies were deliberately withheld and were not used as examples, labels, or seed records during strategy generation. The workflow identified search-relevant concepts while deferring criteria better suited to screening. It then discovered candidate seed records de novo through PubMed, Semantic Scholar, and internal literature-discovery workflows; screened seed relevance; mined titles, abstracts, MeSH terms, entry terms, and indexing patterns; generated candidate Boolean strategies; tested candidates directly in PubMed; and revised strategies using query-structure diagnostics and a structured search-critic step. The agent also had access to a library of validated search filters and used them to increase robustness and interpretability of the query.
RESULTS: In a validation set of 13 reviews, the workflow retrieved 287 of 291 PubMed-indexed included studies, missing 4, for recall of 98.6%. The strategies also retrieved 52,402 non-included records, corresponding to approximately 183 records requiring screening per included study retrieved. Missed studies were concentrated in 3 reviews, suggesting topic-specific term or concept coverage gaps rather than broad workflow failure.
CONCLUSIONS: An agentic workflow combining de novo seed discovery, MeSH and text-word mining, PubMed testing, validated filters, query diagnostics, and structured critique achieved high recall, despite receiving only verbatim eligibility criteria and no known included-study examples. These findings support further evaluation but do not establish generalizability; a held-out test-set evaluation is underway to estimate performance.
METHODS: Eligibility criteria were extracted verbatim from published reviews and used as the only review-level input. To evaluate performance under a difficult, early-stage search-development setting, known included studies were deliberately withheld and were not used as examples, labels, or seed records during strategy generation. The workflow identified search-relevant concepts while deferring criteria better suited to screening. It then discovered candidate seed records de novo through PubMed, Semantic Scholar, and internal literature-discovery workflows; screened seed relevance; mined titles, abstracts, MeSH terms, entry terms, and indexing patterns; generated candidate Boolean strategies; tested candidates directly in PubMed; and revised strategies using query-structure diagnostics and a structured search-critic step. The agent also had access to a library of validated search filters and used them to increase robustness and interpretability of the query.
RESULTS: In a validation set of 13 reviews, the workflow retrieved 287 of 291 PubMed-indexed included studies, missing 4, for recall of 98.6%. The strategies also retrieved 52,402 non-included records, corresponding to approximately 183 records requiring screening per included study retrieved. Missed studies were concentrated in 3 reviews, suggesting topic-specific term or concept coverage gaps rather than broad workflow failure.
CONCLUSIONS: An agentic workflow combining de novo seed discovery, MeSH and text-word mining, PubMed testing, validated filters, query diagnostics, and structured critique achieved high recall, despite receiving only verbatim eligibility criteria and no known included-study examples. These findings support further evaluation but do not establish generalizability; a held-out test-set evaluation is underway to estimate performance.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR276
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas