AUTOMATED QUALITY ASSURANCE FOR LLM-ASSISTED EVIDENCE EXTRACTION IN SYSTEMATIC REVIEWS

Author(s)

Artur Nowak, MSc, Ewa Borowiack, MSc, Ewelina Sadowska, MPharm, Monika Opalek, PhD.
Evidence Prime, Krakow, Poland.
OBJECTIVES: Large language models can prepopulate evidence-extraction forms, but systematic reviews still require human validation of large numbers of suggested values. We developed an automated quality-assurance layer that routes each extracted field to one of three actions: accept the model-generated value, escalate it for human verification, or reject/hide it so the user enters the value manually.
METHODS: We evaluated the approach using a retrospective replay of structured evidence-extraction tasks, with ground-truth labels representing perfect human correction. The primary endpoint was mean human effort per extracted field, expressed in manual-entry-equivalent units, where fully manual entry was assigned a value of 1.0. Auto-accepted fields required no human effort. Rejected fields were treated as manual entry and assigned 1.0. Fields sent for verification were assigned 0.25 when the model-generated value was correct, representing rapid verification, and 1.25 when incorrect, representing verification plus correction. Routing policies were tuned on development data to minimize the human effort while targeting 95% accuracy and subsequently evaluated once on held-out test data.
RESULTS: The best-performing policy used field-specific thresholds based on raw model confidence. Among 272 held-out extraction fields, the policy auto-accepted 116 fields, yielding 42.6% automation coverage with 95.7% accuracy among auto-accepted values. It escalated 154 fields and rejected 2 fields. Overall accuracy was 98.2%. Mean human effort was 0.248 manual-entry-equivalent units per field, representing a 75.2% reduction versus fully manual extraction and a 33.2% reduction versus checking every field under the same observed model-error rates.
CONCLUSIONS: The proposed quality-assurance layer supports auditable automation: high-confidence values are accepted silently, uncertain values are escalated for verification, and likely erroneous suggestions can be suppressed before reaching the reviewer. Further work should assess generalizability across evidence fields, review types, and richer verification signals.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR213

Topic

Methodological & Statistical Research, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×