ACCURACY AND EFFICIENCY OF A HUMAN-IN-THE-LOOP PROMPT-FEEDBACK WORKFLOW FOR ARTIFICIAL INTELLIGENCE-ENABLED EVIDENCE SYNTHESIS: A MULTIPLE MYELOMA SYSTEMATIC LITERATURE REVIEW CASE STUDY

Author(s)

Aniket Das, Ph.D.1, Ankita Roy, M.Sc.1, Ayan Chakraborty, M.Sc.1, Ashmita Chatterjee, M.Sc.1, Jishna Das, M Pharm2, Shubhangi Chatterjee, M.Sc.1, Arunima Bose, M.Sc1, Varun Ektare, MPH3.
1Indence Research Private Limited, North 24 Paraganas, India, 2Indence Research Private Limited, West Bengal, India, 3Indence Research Private Limited, Thane West, India.
OBJECTIVES: Evidence for artificial intelligence (AI)-assisted systematic literature review (SLR) screening and extraction is evolving, particularly in oncology reviews with complex population, intervention, comparator, outcomes, and study design (PICOS) criteria and heterogeneous real-world evidence (RWE). We evaluated a human-in-the-loop (HITL) prompt-feedback ChatGPT (GPT-5.5, OpenAI) workflow using reviewer-corrected examples and indication-specific instructions for title-and-abstract (Ti-Ab) and full-text (FT) screening, and data extraction in a multiple myeloma (MM) SLR.
METHODS: This retrospective case study compared large language model (LLM) outputs with human reviewer quality control (QC). Indication- and study-design-specific instructions covered eligibility criteria, design, population and outcome logic, standard terminology, evidence-linked comments, and rules for complex RWE scenarios. The LLM screened and extracted records; reviewers corrected outputs during QC; and corrections informed later prompts. Human reviewers independently assessed random samples of 50% of Ti-Ab outputs (n=5,000) and FT outputs (n=1,500), plus extraction from 10 records.
RESULTS: For Ti-Ab screening, specificity was 98.4%, precision 98.6%, recall 98.6%, and F1 score 98.6%; 5,000 records were processed in 2.5 hours. For FT screening, specificity was 91.6%, precision 94.0%, recall 98.7%, and F1 score 96.3%; 1,500 articles were processed in 16 hours. Across extraction domains, overall agreement was 91.5%, with an 8.5% error rate, 5.0% missing-data rate, and 7.1% hallucination rate. Agreement ranged from 82.5% for clinical outcomes to 98.3% for economic models. Missingness was highest for clinical outcomes (17.8%), followed by clinical safety (6.3%).
CONCLUSIONS: This HITL prompt-feedback workflow achieved high screening precision and recall with strong extraction agreement. Processing 6,500 screening records in 18.5 hours demonstrated the potential to reduce review time while preserving human oversight. Reviewer-corrected examples, indication-specific instructions, and iterative prompt refinement may improve the reliability and efficiency of AI-assisted SLRs while maintaining methodological rigour.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR139

Topic

Methodological & Statistical Research, Real World Data & Information Systems, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas, Oncology

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×