ACCURACY AND EFFICIENCY OF A HUMAN-IN-THE-LOOP PROMPT-FEEDBACK WORKFLOW FOR ARTIFICIAL INTELLIGENCE-ENABLED EVIDENCE SYNTHESIS: A BILIARY TRACT CANCER SYSTEMATIC LITERATURE REVIEW CASE STUDY
Author(s)
Tushar Pyne, Ph.D.1, Nilanjan Sinha, Ph.D.1, Ayan Chakraborty, M.Sc.1, Tirna Bhattacharya, M.Sc.1, Samreen Kour, M.Tech.1, Aachal Shinde, M.Tech.1, Mrunal Kulkarni, M Pharm2, Ankita ., M.Pharm.1, Abhishikta Mukhopadhyay, M.Sc.1, Varun Ektare, MPH2.
1Indence Research Private Limited, North 24 Paraganas, India, 2Indence Research Private Limited, Thane West, India.
1Indence Research Private Limited, North 24 Paraganas, India, 2Indence Research Private Limited, Thane West, India.
OBJECTIVES: Evidence for artificial intelligence (AI)-assisted systematic literature review (SLR) screening and extraction is evolving, particularly in oncology reviews with heterogeneous population, intervention, comparator, outcomes, and study design (PICOS) criteria and real-world evidence (RWE). We evaluated a human-in-the-loop (HITL) prompt-feedback ChatGPT (GPT-5.5, OpenAI) workflow using reviewer-corrected examples and indication-specific instructions for title-and-abstract (Ti-Ab) and full-text (FT) screening, and data extraction in a biliary tract cancer (BTC) SLR.
METHODS: This retrospective case study compared large language model (LLM) outputs with human reviewer quality control (QC). Indication- and study-design-specific instructions encoded eligibility criteria, design, population and outcome logic, controlled terms, evidence-based comments, and rules for complex RWE cases. The LLM screened and extracted records; reviewers corrected outputs during QC; and corrections informed prompts. Human reviewers assessed random samples of 50% of Ti-Ab outputs (n=5,000) and FT outputs (n=1,500), plus extraction from 10 records.
RESULTS: For Ti-Ab screening, specificity was 91.2%, precision 98.6%, recall 98.6%, and F1 score 98.6%; 5,000 records were processed in 2.5 hours. For FT screening, specificity was 79.1%, precision 94.0%, recall 98.8%, and F1 score 96.3%; 1,500 articles were processed in 16 hours. Cell-level extraction accuracy and agreement were 94.3%, with a 5.7% error rate, 11.7% missing-data rate, and 0.7% hallucination rate. Agreement ranged from 89.6% for clinical outcomes to 100% for cost and biomarker outcomes. Missingness was highest for treatment patterns (23.8%), healthcare resource utilization (HCRU; 22.2%), and clinical outcomes (20.0%).
CONCLUSIONS: This HITL workflow achieved higher precision and F1 scores than previously reported oncology Ti-Ab screening methods, while maintaining near-perfect recall and strong extraction agreement. Processing 6,500 records in 18.5 hours demonstrates potential to reduce review time while preserving human oversight. Reviewer-corrected examples and indication-specific instructions may improve AI-assisted SLR efficiency without compromising methodological rigour.
METHODS: This retrospective case study compared large language model (LLM) outputs with human reviewer quality control (QC). Indication- and study-design-specific instructions encoded eligibility criteria, design, population and outcome logic, controlled terms, evidence-based comments, and rules for complex RWE cases. The LLM screened and extracted records; reviewers corrected outputs during QC; and corrections informed prompts. Human reviewers assessed random samples of 50% of Ti-Ab outputs (n=5,000) and FT outputs (n=1,500), plus extraction from 10 records.
RESULTS: For Ti-Ab screening, specificity was 91.2%, precision 98.6%, recall 98.6%, and F1 score 98.6%; 5,000 records were processed in 2.5 hours. For FT screening, specificity was 79.1%, precision 94.0%, recall 98.8%, and F1 score 96.3%; 1,500 articles were processed in 16 hours. Cell-level extraction accuracy and agreement were 94.3%, with a 5.7% error rate, 11.7% missing-data rate, and 0.7% hallucination rate. Agreement ranged from 89.6% for clinical outcomes to 100% for cost and biomarker outcomes. Missingness was highest for treatment patterns (23.8%), healthcare resource utilization (HCRU; 22.2%), and clinical outcomes (20.0%).
CONCLUSIONS: This HITL workflow achieved higher precision and F1 scores than previously reported oncology Ti-Ab screening methods, while maintaining near-perfect recall and strong extraction agreement. Processing 6,500 records in 18.5 hours demonstrates potential to reduce review time while preserving human oversight. Reviewer-corrected examples and indication-specific instructions may improve AI-assisted SLR efficiency without compromising methodological rigour.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR151
Topic
Methodological & Statistical Research, Real World Data & Information Systems, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas, Oncology