COMPARISON OF GPT-5 MODELS FOR ARTIFICIAL INTELLIGENCE (AI)-ASSISTED DATA EXTRACTION FROM RANDOMIZED CLINICAL TRIALS (RCTS) IN SYSTEMATIC LITERATURE REVIEWS (SLRS): ACCURACY, REPRODUCIBILITY, AND TIME-EFFICIENCY
Author(s)
Mariana Farraia, PhD1, Aiswarya Shree, MSc2, Archana Rajadhyax, M.Pharm3, Anuja Pandey, MD4, SOHAN NITIN DESHPANDE, MSc4, Carolina Casañas i Comabella, MSc, PhD5.
1Thermo Fischer Scientific, Lisbon, Portugal, 2Thermo Fisher Scientific, Bhubaneswar, India, 3Thermo Fisher Scientific, Mumbai, India, 4Thermo Fisher Scientific, London, United Kingdom, 5Thermo Fisher Scientific, Oxford, United Kingdom.
1Thermo Fischer Scientific, Lisbon, Portugal, 2Thermo Fisher Scientific, Bhubaneswar, India, 3Thermo Fisher Scientific, Mumbai, India, 4Thermo Fisher Scientific, London, United Kingdom, 5Thermo Fisher Scientific, Oxford, United Kingdom.
OBJECTIVES: Large-language models (LLMs) show promising capabilities to support data extraction in evidence synthesis (ES). However, their rapid development poses challenges regarding reproducibility, replicability, and accuracy. Therefore, ongoing evaluation in real-world workflows is required. This study evaluated the reproducibility of three GPT models by comparing their accuracy and time-efficiency for AI-assisted data extraction from RCTs.
METHODS: A structured prompt was developed using meta-prompting and iterative human-led testing and refinement. The prompt was used to extract data from ten RCTs using GPT-5.3 Instant, GPT-5.5 Thinking-Standard, and GPT-5.5 Thinking-Extended. Time required for prompt development and data extraction was recorded. AI-extracted variables were compared with a human-conducted SLR, which was used as the gold standard. Each AI-extracted data point was categorized as correct, incorrect, incomplete, or missing. Accuracy was defined as the proportion of correctly extracted variables. Incorrect, incomplete, or missing were considered errors.
RESULTS: Prompt development and optimization required ~2h. GPT-5.5 Thinking-Standard was the fastest model (mean 2m46s extraction time per publication [p/p]), followed by GPT-5.5 Thinking-Extended (4m44s p/p) and GPT-5.3 Instant (5m14s p/p), versus over 1h p/p for human extraction. GPT-5.5 Thinking models achieved the highest accuracy, matching or exceeding our previous evaluation (average: 84%, GPT-4 on June 2024). Thinking models achieved higher performance for study- and patient-level characteristics, reaching 100% for some publications. Hallucinated extractions were uncommon. Most errors reflected incomplete extraction, or over-extraction.
CONCLUSIONS: All GPT models achieved high accuracy with remarkably faster execution compared to humans, but still required full human validation to address errors. Reproducibility and replicability remain a challenge when using AI in SLRs. Human oversight is required to design efficient prompts, provide templates, and review/validate outputs. Further evaluation in larger datasets and different study designs is needed to evaluate efficiency in ES workflows while maintaining quality through human oversight.
METHODS: A structured prompt was developed using meta-prompting and iterative human-led testing and refinement. The prompt was used to extract data from ten RCTs using GPT-5.3 Instant, GPT-5.5 Thinking-Standard, and GPT-5.5 Thinking-Extended. Time required for prompt development and data extraction was recorded. AI-extracted variables were compared with a human-conducted SLR, which was used as the gold standard. Each AI-extracted data point was categorized as correct, incorrect, incomplete, or missing. Accuracy was defined as the proportion of correctly extracted variables. Incorrect, incomplete, or missing were considered errors.
RESULTS: Prompt development and optimization required ~2h. GPT-5.5 Thinking-Standard was the fastest model (mean 2m46s extraction time per publication [p/p]), followed by GPT-5.5 Thinking-Extended (4m44s p/p) and GPT-5.3 Instant (5m14s p/p), versus over 1h p/p for human extraction. GPT-5.5 Thinking models achieved the highest accuracy, matching or exceeding our previous evaluation (average: 84%, GPT-4 on June 2024). Thinking models achieved higher performance for study- and patient-level characteristics, reaching 100% for some publications. Hallucinated extractions were uncommon. Most errors reflected incomplete extraction, or over-extraction.
CONCLUSIONS: All GPT models achieved high accuracy with remarkably faster execution compared to humans, but still required full human validation to address errors. Reproducibility and replicability remain a challenge when using AI in SLRs. Human oversight is required to design efficient prompts, provide templates, and review/validate outputs. Further evaluation in larger datasets and different study designs is needed to evaluate efficiency in ES workflows while maintaining quality through human oversight.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR42
Topic
Methodological & Statistical Research, Organizational Practices, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas