CLOSER, BUT NOT THERE YET: EVALUATING LARGE LANGUAGE MODELS (LLMS) FOR SCOPING SEARCHES IN ONCOLOGY SYSTEMATIC LITERATURE REVIEWS (SLRS)

Author(s)

Theodora Oikonomidi, PhD1, Dimitra Gerna, MSc1, Adam Asteriadis, MSc1, Yogesh Punekar, MSc2.
1IQVIA, Athens, Greece, 2IQVIA, London, United Kingdom.
OBJECTIVES: Scoping searches are critical in SLR planning and estimating anticipated workload. Earlier LLMs generated scoping searches with invalid subject headings, unsupported characters, and syntax errors. This study assessed Claude Opus 4.7 in developing SLR scoping searches.
METHODS: Searches previously developed by subject matter experts (SMEs) for three SLRs in oncology were compared with scoping searches generated by Opus. The LLM was prompted to develop a search for Embase® via OVID SP®, based on two inputs: the research question and the population, intervention, comparison, outcome, and study design (PICOS) criteria. LLM-generated searches were run and results reviewed by SMEs. Evaluation criteria were: (1) whether the search code executed on OVID SP®; (2) total number of results versus the SME search; and (3) reliability (i.e. whether rerunning the same prompt produced strategies yielding similar hits).
RESULTS: In all test cases, LLM searches yielded more hits than SME searches, by 370 to 24,413 records. One of three LLM searches required SME correction to run on OVID SP®; otherwise, outputs were free from major errors. Opus correctly selected free-text search fields (ti,ab,kf,kw), exploded indexing terms appropriately, and applied correct Boolean logic and limit syntax for conference abstracts, language, and publication year. Two factors drove results overestimation: (1) use of non-indication-specific terms (e.g. “UC” for urothelial carcinoma capturing ulcerative colitis records); and (2) failure to exclude ineligible study types (e.g. animal studies, case reports, in vitro studies). Rerunning the same prompt with the same input produced a different search strategy with added limits, yielding approximately 40% of the results of the first LLM-produced search strategy, indicating low reliability.
CONCLUSIONS: Based on findings from this study, LLM performance in estimating search yield was inconsistent and potentially misleading for SLR scoping, although strategies were largely free from major syntax errors and indexing term hallucinations.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR300

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×