AUTOMATED QUALITY ASSURANCE FOR LLM-ASSISTED EVIDENCE EXTRACTION IN SYSTEMATIC REVIEWS
Author(s)
Artur Nowak, MSc, Ewa Borowiack, MSc, Ewelina Sadowska, MPharm, Monika Opalek, PhD.
Evidence Prime, Krakow, Poland.
Evidence Prime, Krakow, Poland.
OBJECTIVES: Large language models can prepopulate evidence-extraction forms, but systematic reviews still require human validation of large numbers of suggested values. We developed an automated quality-assurance layer that routes each extracted field to one of three actions: accept the model-generated value, escalate it for human verification, or reject/hide it so the user enters the value manually.
METHODS: We evaluated the approach using a retrospective replay of structured evidence-extraction tasks, with ground-truth labels representing perfect human correction. The primary endpoint was mean human effort per extracted field, expressed in manual-entry-equivalent units, where fully manual entry was assigned a value of 1.0. Auto-accepted fields required no human effort. Rejected fields were treated as manual entry and assigned 1.0. Fields sent for verification were assigned 0.25 when the model-generated value was correct, representing rapid verification, and 1.25 when incorrect, representing verification plus correction. Routing policies were tuned on development data to minimize the human effort while targeting 95% accuracy and subsequently evaluated once on held-out test data.
RESULTS: The best-performing policy used field-specific thresholds based on raw model confidence. Among 272 held-out extraction fields, the policy auto-accepted 116 fields, yielding 42.6% automation coverage with 95.7% accuracy among auto-accepted values. It escalated 154 fields and rejected 2 fields. Overall accuracy was 98.2%. Mean human effort was 0.248 manual-entry-equivalent units per field, representing a 75.2% reduction versus fully manual extraction and a 33.2% reduction versus checking every field under the same observed model-error rates.
CONCLUSIONS: The proposed quality-assurance layer supports auditable automation: high-confidence values are accepted silently, uncertain values are escalated for verification, and likely erroneous suggestions can be suppressed before reaching the reviewer. Further work should assess generalizability across evidence fields, review types, and richer verification signals.
METHODS: We evaluated the approach using a retrospective replay of structured evidence-extraction tasks, with ground-truth labels representing perfect human correction. The primary endpoint was mean human effort per extracted field, expressed in manual-entry-equivalent units, where fully manual entry was assigned a value of 1.0. Auto-accepted fields required no human effort. Rejected fields were treated as manual entry and assigned 1.0. Fields sent for verification were assigned 0.25 when the model-generated value was correct, representing rapid verification, and 1.25 when incorrect, representing verification plus correction. Routing policies were tuned on development data to minimize the human effort while targeting 95% accuracy and subsequently evaluated once on held-out test data.
RESULTS: The best-performing policy used field-specific thresholds based on raw model confidence. Among 272 held-out extraction fields, the policy auto-accepted 116 fields, yielding 42.6% automation coverage with 95.7% accuracy among auto-accepted values. It escalated 154 fields and rejected 2 fields. Overall accuracy was 98.2%. Mean human effort was 0.248 manual-entry-equivalent units per field, representing a 75.2% reduction versus fully manual extraction and a 33.2% reduction versus checking every field under the same observed model-error rates.
CONCLUSIONS: The proposed quality-assurance layer supports auditable automation: high-confidence values are accepted silently, uncertain values are escalated for verification, and likely erroneous suggestions can be suppressed before reaching the reviewer. Further work should assess generalizability across evidence fields, review types, and richer verification signals.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR213
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas