ACCURACY AND EFFICIENCY OF A HUMAN-IN-THE-LOOP PROMPT-FEEDBACK WORKFLOW FOR ARTIFICIAL INTELLIGENCE-ENABLED EVIDENCE SYNTHESIS: A BILIARY TRACT CANCER SYSTEMATIC LITERATURE REVIEW CASE STUDY

Author(s)

Tushar Pyne, Ph.D.1, Nilanjan Sinha, Ph.D.1, Ayan Chakraborty, M.Sc.1, Tirna Bhattacharya, M.Sc.1, Samreen Kour, M.Tech.1, Aachal Shinde, M.Tech.1, Mrunal Kulkarni, M Pharm2, Ankita ., M.Pharm.1, Abhishikta Mukhopadhyay, M.Sc.1, Varun Ektare, MPH2.
1Indence Research Private Limited, North 24 Paraganas, India, 2Indence Research Private Limited, Thane West, India.
OBJECTIVES: Evidence for artificial intelligence (AI)-assisted systematic literature review (SLR) screening and extraction is evolving, particularly in oncology reviews with heterogeneous population, intervention, comparator, outcomes, and study design (PICOS) criteria and real-world evidence (RWE). We evaluated a human-in-the-loop (HITL) prompt-feedback ChatGPT (GPT-5.5, OpenAI) workflow using reviewer-corrected examples and indication-specific instructions for title-and-abstract (Ti-Ab) and full-text (FT) screening, and data extraction in a biliary tract cancer (BTC) SLR.
METHODS: This retrospective case study compared large language model (LLM) outputs with human reviewer quality control (QC). Indication- and study-design-specific instructions encoded eligibility criteria, design, population and outcome logic, controlled terms, evidence-based comments, and rules for complex RWE cases. The LLM screened and extracted records; reviewers corrected outputs during QC; and corrections informed prompts. Human reviewers assessed random samples of 50% of Ti-Ab outputs (n=5,000) and FT outputs (n=1,500), plus extraction from 10 records.
RESULTS: For Ti-Ab screening, specificity was 91.2%, precision 98.6%, recall 98.6%, and F1 score 98.6%; 5,000 records were processed in 2.5 hours. For FT screening, specificity was 79.1%, precision 94.0%, recall 98.8%, and F1 score 96.3%; 1,500 articles were processed in 16 hours. Cell-level extraction accuracy and agreement were 94.3%, with a 5.7% error rate, 11.7% missing-data rate, and 0.7% hallucination rate. Agreement ranged from 89.6% for clinical outcomes to 100% for cost and biomarker outcomes. Missingness was highest for treatment patterns (23.8%), healthcare resource utilization (HCRU; 22.2%), and clinical outcomes (20.0%).
CONCLUSIONS: This HITL workflow achieved higher precision and F1 scores than previously reported oncology Ti-Ab screening methods, while maintaining near-perfect recall and strong extraction agreement. Processing 6,500 records in 18.5 hours demonstrates potential to reduce review time while preserving human oversight. Reviewer-corrected examples and indication-specific instructions may improve AI-assisted SLR efficiency without compromising methodological rigour.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR151

Topic

Methodological & Statistical Research, Real World Data & Information Systems, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas, Oncology

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×