EVALUATING CHATGPT AND CLAUDE FOR AI-ASSISTED TITLE/ABSTRACT SCREENING: PERFORMANCE GAPS, ITERATIVE PROMPT REFINEMENT REQUIREMENTS, AND RESOURCE BURDEN
Author(s)
Darsh Devani, MS, Grace E. Fox, PhD.
OPEN Health, New York, NY, USA.
OPEN Health, New York, NY, USA.
OBJECTIVES: General-purpose large language models (LLMs) are increasingly used for title/abstract screening, yet how these models fail (risking missed evidence or incorrect inclusions) and the prompt refinement burden required to improve performance remain poorly understood. This study benchmarks ChatGPT and Claude against a human reference standard and quantifies the refinement cycles and researcher time required
METHODS: ChatGPT (GPT-5.5) and Claude (Claude Sonnet 4.6) independently screened titles/abstracts of 992 cardiovascular disease literature review references using a PICOS-based 5-question algorithm for full-text screening inclusion. Performance was assessed against 98 human-reviewed inclusions, with 100% sensitivity as the target threshold. 3 refinement cycles were conducted and researcher time recorded.
RESULTS: At first-pass title/abstract screening, ChatGPT missed 84 of 98 references (sensitivity=0.14), 80 (95%) due to misclassifying registry studies and databases as non-human studies. Claude missed 12 (sensitivity=0.88), not recognizing national cardiac registries and administrative databases as eligible designs. ChatGPT produced identical results across repeat runs, while Claude produced slightly different results across sessions on the same prompt without modification. After 3 cycles, ChatGPT sensitivity reached 0.94, with 575 references incorrectly advanced for review (false positives) and 6 references still missed; 3 involved younger adults misclassified as non-adult and 3 involved national health database studies not recognized as eligible. Claude maintained sensitivity=1.00 throughout, with false positives reducing from 406 to 312; however, 312 persisted despite targeted exclusion criteria. Failure pattern identification and criteria redesign required ~6 combined researcher hours.
CONCLUSIONS: ChatGPT excluded most relevant studies at first pass and shifted output unpredictably with prompt changes. Claude over-included despite targeted exclusion instructions. Neither model resolved these issues without refinement cycles. While Claude achieved the threshold for HTA-grade SLR, ChatGPT failed to reach this threshold even after 3 cycles. These failure modes and the iterative refinement burden of these general-purpose LLMs must factor into assessments of AI-assisted screening tools.
METHODS: ChatGPT (GPT-5.5) and Claude (Claude Sonnet 4.6) independently screened titles/abstracts of 992 cardiovascular disease literature review references using a PICOS-based 5-question algorithm for full-text screening inclusion. Performance was assessed against 98 human-reviewed inclusions, with 100% sensitivity as the target threshold. 3 refinement cycles were conducted and researcher time recorded.
RESULTS: At first-pass title/abstract screening, ChatGPT missed 84 of 98 references (sensitivity=0.14), 80 (95%) due to misclassifying registry studies and databases as non-human studies. Claude missed 12 (sensitivity=0.88), not recognizing national cardiac registries and administrative databases as eligible designs. ChatGPT produced identical results across repeat runs, while Claude produced slightly different results across sessions on the same prompt without modification. After 3 cycles, ChatGPT sensitivity reached 0.94, with 575 references incorrectly advanced for review (false positives) and 6 references still missed; 3 involved younger adults misclassified as non-adult and 3 involved national health database studies not recognized as eligible. Claude maintained sensitivity=1.00 throughout, with false positives reducing from 406 to 312; however, 312 persisted despite targeted exclusion criteria. Failure pattern identification and criteria redesign required ~6 combined researcher hours.
CONCLUSIONS: ChatGPT excluded most relevant studies at first pass and shifted output unpredictably with prompt changes. Claude over-included despite targeted exclusion instructions. Neither model resolved these issues without refinement cycles. While Claude achieved the threshold for HTA-grade SLR, ChatGPT failed to reach this threshold even after 3 cycles. These failure modes and the iterative refinement burden of these general-purpose LLMs must factor into assessments of AI-assisted screening tools.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
SA44
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Literature Review & Synthesis
Disease
Cardiovascular Disorders (including MI, Stroke, Circulatory)