HUMAN-IN-THE-LOOP AI-ASSISTED FULL TEXT REVIEW FOR SAFETY EVIDENCE IDENTIFICATION
Author(s)
Artur Nowak, MSc, Ewa Borowiack, MSc, Ewelina Sadowska, MPharm, Monika Opalek, PhD.
Evidence Prime, Krakow, Poland.
Evidence Prime, Krakow, Poland.
OBJECTIVES: Full-text review is a resource-intensive component of safety evidence review. Large language models (LLMs) have demonstrated promising performance in screening tasks. However, translating model outputs into transparent, auditable, and human verifiable workflows remains a challenge.
As part of a broader research program evaluating AI-assisted full-text screening, initial studies focused on benchmarking model performance against established reference standards. The present study complements these evaluations by assessing how reviewers interact with, verify, and use AI-generated recommendations within a real world human-in-the-loop (HITL) workflow.
METHODS: Publications for five drugs representing diverse therapeutic areas (Metronidazole, Semaglutide, Amiodarone, Isotretinoin, Pembrolizumab) were selected. For each drug, approximately 100 randomly selected PubMed-indexed PDF publications were included.
Prior to evaluation, the Laser AI eligibility-based decision logic, including screening questions, was iteratively refined through expert guided testing and analysis of model outputs. The system applied the final decision logic, generating answers to screening questions, supporting reasoning, highlighted evidence passages, and links to relevant source text. Individual answers were aggregated into the final decision hint. An experienced reviewer evaluated each PDF within Laser AI and either accepted or rejected AI suggestions. Reviewer decisions served as the reference standard.
Outcomes included sensitivity, specificity, precision, accuracy, disagreement rates for individual screening questions, and workflow usability
RESULTS: Agreement between AI-generated recommendations and final reviewer decisions corresponded to a sensitivity of 98.77%, specificity of 98.22%, precision of 96.41%, and overall accuracy of 97.87%. At the individual screening-question level, disagreement rates remained low (0.0% to 2.0%).
CONCLUSIONS: While benchmark-based evaluations demonstrated the technical performance of AI-assisted full-text screening, this study provides evidence of its practical utility within a real-world reviewer workflow. By combining AI-generated recommendations, supporting reasoning, and direct access to source evidence, the framework enabled efficient validation of AI assessments while preserving expert control over final decisions.
As part of a broader research program evaluating AI-assisted full-text screening, initial studies focused on benchmarking model performance against established reference standards. The present study complements these evaluations by assessing how reviewers interact with, verify, and use AI-generated recommendations within a real world human-in-the-loop (HITL) workflow.
METHODS: Publications for five drugs representing diverse therapeutic areas (Metronidazole, Semaglutide, Amiodarone, Isotretinoin, Pembrolizumab) were selected. For each drug, approximately 100 randomly selected PubMed-indexed PDF publications were included.
Prior to evaluation, the Laser AI eligibility-based decision logic, including screening questions, was iteratively refined through expert guided testing and analysis of model outputs. The system applied the final decision logic, generating answers to screening questions, supporting reasoning, highlighted evidence passages, and links to relevant source text. Individual answers were aggregated into the final decision hint. An experienced reviewer evaluated each PDF within Laser AI and either accepted or rejected AI suggestions. Reviewer decisions served as the reference standard.
Outcomes included sensitivity, specificity, precision, accuracy, disagreement rates for individual screening questions, and workflow usability
RESULTS: Agreement between AI-generated recommendations and final reviewer decisions corresponded to a sensitivity of 98.77%, specificity of 98.22%, precision of 96.41%, and overall accuracy of 97.87%. At the individual screening-question level, disagreement rates remained low (0.0% to 2.0%).
CONCLUSIONS: While benchmark-based evaluations demonstrated the technical performance of AI-assisted full-text screening, this study provides evidence of its practical utility within a real-world reviewer workflow. By combining AI-generated recommendations, supporting reasoning, and direct access to source evidence, the framework enabled efficient validation of AI assessments while preserving expert control over final decisions.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR103
Topic
Epidemiology & Public Health, Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas