EVALUATING CHATGPT AND CLAUDE FOR AI-ASSISTED TITLE/ABSTRACT SCREENING: PERFORMANCE GAPS, ITERATIVE PROMPT REFINEMENT REQUIREMENTS, AND RESOURCE BURDEN

Author(s)

Darsh Devani, MS, Grace E. Fox, PhD.
OPEN Health, New York, NY, USA.
OBJECTIVES: General-purpose large language models (LLMs) are increasingly used for title/abstract screening, yet how these models fail (risking missed evidence or incorrect inclusions) and the prompt refinement burden required to improve performance remain poorly understood. This study benchmarks ChatGPT and Claude against a human reference standard and quantifies the refinement cycles and researcher time required
METHODS: ChatGPT (GPT-5.5) and Claude (Claude Sonnet 4.6) independently screened titles/abstracts of 992 cardiovascular disease literature review references using a PICOS-based 5-question algorithm for full-text screening inclusion. Performance was assessed against 98 human-reviewed inclusions, with 100% sensitivity as the target threshold. 3 refinement cycles were conducted and researcher time recorded.
RESULTS: At first-pass title/abstract screening, ChatGPT missed 84 of 98 references (sensitivity=0.14), 80 (95%) due to misclassifying registry studies and databases as non-human studies. Claude missed 12 (sensitivity=0.88), not recognizing national cardiac registries and administrative databases as eligible designs. ChatGPT produced identical results across repeat runs, while Claude produced slightly different results across sessions on the same prompt without modification. After 3 cycles, ChatGPT sensitivity reached 0.94, with 575 references incorrectly advanced for review (false positives) and 6 references still missed; 3 involved younger adults misclassified as non-adult and 3 involved national health database studies not recognized as eligible. Claude maintained sensitivity=1.00 throughout, with false positives reducing from 406 to 312; however, 312 persisted despite targeted exclusion criteria. Failure pattern identification and criteria redesign required ~6 combined researcher hours.
CONCLUSIONS: ChatGPT excluded most relevant studies at first pass and shifted output unpredictably with prompt changes. Claude over-included despite targeted exclusion instructions. Neither model resolved these issues without refinement cycles. While Claude achieved the threshold for HTA-grade SLR, ChatGPT failed to reach this threshold even after 3 cycles. These failure modes and the iterative refinement burden of these general-purpose LLMs must factor into assessments of AI-assisted screening tools.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

SA44

Topic

Methodological & Statistical Research, Study Approaches

Topic Subcategory

Literature Review & Synthesis

Disease

Cardiovascular Disorders (including MI, Stroke, Circulatory)

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×