HUMAN-IN-THE-LOOP AI-ASSISTED FULL TEXT REVIEW FOR SAFETY EVIDENCE IDENTIFICATION

Author(s)

Artur Nowak, MSc, Ewa Borowiack, MSc, Ewelina Sadowska, MPharm, Monika Opalek, PhD.
Evidence Prime, Krakow, Poland.
OBJECTIVES: Full-text review is a resource-intensive component of safety evidence review. Large language models (LLMs) have demonstrated promising performance in screening tasks. However, translating model outputs into transparent, auditable, and human verifiable workflows remains a challenge.
As part of a broader research program evaluating AI-assisted full-text screening, initial studies focused on benchmarking model performance against established reference standards. The present study complements these evaluations by assessing how reviewers interact with, verify, and use AI-generated recommendations within a real world human-in-the-loop (HITL) workflow.
METHODS: Publications for five drugs representing diverse therapeutic areas (Metronidazole, Semaglutide, Amiodarone, Isotretinoin, Pembrolizumab) were selected. For each drug, approximately 100 randomly selected PubMed-indexed PDF publications were included.
Prior to evaluation, the Laser AI eligibility-based decision logic, including screening questions, was iteratively refined through expert guided testing and analysis of model outputs. The system applied the final decision logic, generating answers to screening questions, supporting reasoning, highlighted evidence passages, and links to relevant source text. Individual answers were aggregated into the final decision hint. An experienced reviewer evaluated each PDF within Laser AI and either accepted or rejected AI suggestions. Reviewer decisions served as the reference standard.
Outcomes included sensitivity, specificity, precision, accuracy, disagreement rates for individual screening questions, and workflow usability
RESULTS: Agreement between AI-generated recommendations and final reviewer decisions corresponded to a sensitivity of 98.77%, specificity of 98.22%, precision of 96.41%, and overall accuracy of 97.87%. At the individual screening-question level, disagreement rates remained low (0.0% to 2.0%).
CONCLUSIONS: While benchmark-based evaluations demonstrated the technical performance of AI-assisted full-text screening, this study provides evidence of its practical utility within a real-world reviewer workflow. By combining AI-generated recommendations, supporting reasoning, and direct access to source evidence, the framework enabled efficient validation of AI assessments while preserving expert control over final decisions.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR103

Topic

Epidemiology & Public Health, Methodological & Statistical Research, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×