ACCURACY AND EFFICIENCY OF A HUMAN-IN-THE-LOOP PROMPT-FEEDBACK WORKFLOW FOR ARTIFICIAL INTELLIGENCE-ENABLED EVIDENCE SYNTHESIS: A MULTIPLE MYELOMA SYSTEMATIC LITERATURE REVIEW CASE STUDY
Author(s)
Aniket Das, Ph.D.1, Ankita Roy, M.Sc.1, Ayan Chakraborty, M.Sc.1, Ashmita Chatterjee, M.Sc.1, Jishna Das, M Pharm2, Shubhangi Chatterjee, M.Sc.1, Arunima Bose, M.Sc1, Varun Ektare, MPH3.
1Indence Research Private Limited, North 24 Paraganas, India, 2Indence Research Private Limited, West Bengal, India, 3Indence Research Private Limited, Thane West, India.
1Indence Research Private Limited, North 24 Paraganas, India, 2Indence Research Private Limited, West Bengal, India, 3Indence Research Private Limited, Thane West, India.
OBJECTIVES: Evidence for artificial intelligence (AI)-assisted systematic literature review (SLR) screening and extraction is evolving, particularly in oncology reviews with complex population, intervention, comparator, outcomes, and study design (PICOS) criteria and heterogeneous real-world evidence (RWE). We evaluated a human-in-the-loop (HITL) prompt-feedback ChatGPT (GPT-5.5, OpenAI) workflow using reviewer-corrected examples and indication-specific instructions for title-and-abstract (Ti-Ab) and full-text (FT) screening, and data extraction in a multiple myeloma (MM) SLR.
METHODS: This retrospective case study compared large language model (LLM) outputs with human reviewer quality control (QC). Indication- and study-design-specific instructions covered eligibility criteria, design, population and outcome logic, standard terminology, evidence-linked comments, and rules for complex RWE scenarios. The LLM screened and extracted records; reviewers corrected outputs during QC; and corrections informed later prompts. Human reviewers independently assessed random samples of 50% of Ti-Ab outputs (n=5,000) and FT outputs (n=1,500), plus extraction from 10 records.
RESULTS: For Ti-Ab screening, specificity was 98.4%, precision 98.6%, recall 98.6%, and F1 score 98.6%; 5,000 records were processed in 2.5 hours. For FT screening, specificity was 91.6%, precision 94.0%, recall 98.7%, and F1 score 96.3%; 1,500 articles were processed in 16 hours. Across extraction domains, overall agreement was 91.5%, with an 8.5% error rate, 5.0% missing-data rate, and 7.1% hallucination rate. Agreement ranged from 82.5% for clinical outcomes to 98.3% for economic models. Missingness was highest for clinical outcomes (17.8%), followed by clinical safety (6.3%).
CONCLUSIONS: This HITL prompt-feedback workflow achieved high screening precision and recall with strong extraction agreement. Processing 6,500 screening records in 18.5 hours demonstrated the potential to reduce review time while preserving human oversight. Reviewer-corrected examples, indication-specific instructions, and iterative prompt refinement may improve the reliability and efficiency of AI-assisted SLRs while maintaining methodological rigour.
METHODS: This retrospective case study compared large language model (LLM) outputs with human reviewer quality control (QC). Indication- and study-design-specific instructions covered eligibility criteria, design, population and outcome logic, standard terminology, evidence-linked comments, and rules for complex RWE scenarios. The LLM screened and extracted records; reviewers corrected outputs during QC; and corrections informed later prompts. Human reviewers independently assessed random samples of 50% of Ti-Ab outputs (n=5,000) and FT outputs (n=1,500), plus extraction from 10 records.
RESULTS: For Ti-Ab screening, specificity was 98.4%, precision 98.6%, recall 98.6%, and F1 score 98.6%; 5,000 records were processed in 2.5 hours. For FT screening, specificity was 91.6%, precision 94.0%, recall 98.7%, and F1 score 96.3%; 1,500 articles were processed in 16 hours. Across extraction domains, overall agreement was 91.5%, with an 8.5% error rate, 5.0% missing-data rate, and 7.1% hallucination rate. Agreement ranged from 82.5% for clinical outcomes to 98.3% for economic models. Missingness was highest for clinical outcomes (17.8%), followed by clinical safety (6.3%).
CONCLUSIONS: This HITL prompt-feedback workflow achieved high screening precision and recall with strong extraction agreement. Processing 6,500 screening records in 18.5 hours demonstrated the potential to reduce review time while preserving human oversight. Reviewer-corrected examples, indication-specific instructions, and iterative prompt refinement may improve the reliability and efficiency of AI-assisted SLRs while maintaining methodological rigour.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR139
Topic
Methodological & Statistical Research, Real World Data & Information Systems, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas, Oncology