COMPARISON OF GPT-5 MODELS FOR ARTIFICIAL INTELLIGENCE (AI)-ASSISTED DATA EXTRACTION FROM RANDOMIZED CLINICAL TRIALS (RCTS) IN SYSTEMATIC LITERATURE REVIEWS (SLRS): ACCURACY, REPRODUCIBILITY, AND TIME-EFFICIENCY

Author(s)

Mariana Farraia, PhD1, Aiswarya Shree, MSc2, Archana Rajadhyax, M.Pharm3, Anuja Pandey, MD4, SOHAN NITIN DESHPANDE, MSc4, Carolina Casañas i Comabella, MSc, PhD5.
1Thermo Fischer Scientific, Lisbon, Portugal, 2Thermo Fisher Scientific, Bhubaneswar, India, 3Thermo Fisher Scientific, Mumbai, India, 4Thermo Fisher Scientific, London, United Kingdom, 5Thermo Fisher Scientific, Oxford, United Kingdom.
OBJECTIVES: Large-language models (LLMs) show promising capabilities to support data extraction in evidence synthesis (ES). However, their rapid development poses challenges regarding reproducibility, replicability, and accuracy. Therefore, ongoing evaluation in real-world workflows is required. This study evaluated the reproducibility of three GPT models by comparing their accuracy and time-efficiency for AI-assisted data extraction from RCTs.
METHODS: A structured prompt was developed using meta-prompting and iterative human-led testing and refinement. The prompt was used to extract data from ten RCTs using GPT-5.3 Instant, GPT-5.5 Thinking-Standard, and GPT-5.5 Thinking-Extended. Time required for prompt development and data extraction was recorded. AI-extracted variables were compared with a human-conducted SLR, which was used as the gold standard. Each AI-extracted data point was categorized as correct, incorrect, incomplete, or missing. Accuracy was defined as the proportion of correctly extracted variables. Incorrect, incomplete, or missing were considered errors.
RESULTS: Prompt development and optimization required ~2h. GPT-5.5 Thinking-Standard was the fastest model (mean 2m46s extraction time per publication [p/p]), followed by GPT-5.5 Thinking-Extended (4m44s p/p) and GPT-5.3 Instant (5m14s p/p), versus over 1h p/p for human extraction. GPT-5.5 Thinking models achieved the highest accuracy, matching or exceeding our previous evaluation (average: 84%, GPT-4 on June 2024). Thinking models achieved higher performance for study- and patient-level characteristics, reaching 100% for some publications. Hallucinated extractions were uncommon. Most errors reflected incomplete extraction, or over-extraction.
CONCLUSIONS: All GPT models achieved high accuracy with remarkably faster execution compared to humans, but still required full human validation to address errors. Reproducibility and replicability remain a challenge when using AI in SLRs. Human oversight is required to design efficient prompts, provide templates, and review/validate outputs. Further evaluation in larger datasets and different study designs is needed to evaluate efficiency in ES workflows while maintaining quality through human oversight.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR42

Topic

Methodological & Statistical Research, Organizational Practices, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×