CLOSER, BUT NOT THERE YET: EVALUATING LARGE LANGUAGE MODELS (LLMS) FOR SCOPING SEARCHES IN ONCOLOGY SYSTEMATIC LITERATURE REVIEWS (SLRS)
Author(s)
Theodora Oikonomidi, PhD1, Dimitra Gerna, MSc1, Adam Asteriadis, MSc1, Yogesh Punekar, MSc2.
1IQVIA, Athens, Greece, 2IQVIA, London, United Kingdom.
1IQVIA, Athens, Greece, 2IQVIA, London, United Kingdom.
OBJECTIVES: Scoping searches are critical in SLR planning and estimating anticipated workload. Earlier LLMs generated scoping searches with invalid subject headings, unsupported characters, and syntax errors. This study assessed Claude Opus 4.7 in developing SLR scoping searches.
METHODS: Searches previously developed by subject matter experts (SMEs) for three SLRs in oncology were compared with scoping searches generated by Opus. The LLM was prompted to develop a search for Embase® via OVID SP®, based on two inputs: the research question and the population, intervention, comparison, outcome, and study design (PICOS) criteria. LLM-generated searches were run and results reviewed by SMEs. Evaluation criteria were: (1) whether the search code executed on OVID SP®; (2) total number of results versus the SME search; and (3) reliability (i.e. whether rerunning the same prompt produced strategies yielding similar hits).
RESULTS: In all test cases, LLM searches yielded more hits than SME searches, by 370 to 24,413 records. One of three LLM searches required SME correction to run on OVID SP®; otherwise, outputs were free from major errors. Opus correctly selected free-text search fields (ti,ab,kf,kw), exploded indexing terms appropriately, and applied correct Boolean logic and limit syntax for conference abstracts, language, and publication year. Two factors drove results overestimation: (1) use of non-indication-specific terms (e.g. “UC” for urothelial carcinoma capturing ulcerative colitis records); and (2) failure to exclude ineligible study types (e.g. animal studies, case reports, in vitro studies). Rerunning the same prompt with the same input produced a different search strategy with added limits, yielding approximately 40% of the results of the first LLM-produced search strategy, indicating low reliability.
CONCLUSIONS: Based on findings from this study, LLM performance in estimating search yield was inconsistent and potentially misleading for SLR scoping, although strategies were largely free from major syntax errors and indexing term hallucinations.
METHODS: Searches previously developed by subject matter experts (SMEs) for three SLRs in oncology were compared with scoping searches generated by Opus. The LLM was prompted to develop a search for Embase® via OVID SP®, based on two inputs: the research question and the population, intervention, comparison, outcome, and study design (PICOS) criteria. LLM-generated searches were run and results reviewed by SMEs. Evaluation criteria were: (1) whether the search code executed on OVID SP®; (2) total number of results versus the SME search; and (3) reliability (i.e. whether rerunning the same prompt produced strategies yielding similar hits).
RESULTS: In all test cases, LLM searches yielded more hits than SME searches, by 370 to 24,413 records. One of three LLM searches required SME correction to run on OVID SP®; otherwise, outputs were free from major errors. Opus correctly selected free-text search fields (ti,ab,kf,kw), exploded indexing terms appropriately, and applied correct Boolean logic and limit syntax for conference abstracts, language, and publication year. Two factors drove results overestimation: (1) use of non-indication-specific terms (e.g. “UC” for urothelial carcinoma capturing ulcerative colitis records); and (2) failure to exclude ineligible study types (e.g. animal studies, case reports, in vitro studies). Rerunning the same prompt with the same input produced a different search strategy with added limits, yielding approximately 40% of the results of the first LLM-produced search strategy, indicating low reliability.
CONCLUSIONS: Based on findings from this study, LLM performance in estimating search yield was inconsistent and potentially misleading for SLR scoping, although strategies were largely free from major syntax errors and indexing term hallucinations.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR300
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas