EVALUATING FULLY AUTOMATED FULL-TEXT SCREENING FOR HTA SYSTEMATIC REVIEWS

Author(s)

Artur Nowak, MSc, Ewelina Sadowska, MPharm, Ewa Borowiack, MSc, Monika Opalek, PhD.
Evidence Prime, Krakow, Poland.
OBJECTIVES: Full-text (FT) screening is a critical but resource-intensive step in systematic reviews for health technology assessment (HTA). As part of a broader program evaluating AI-assisted full-text review, this study focused on technical validation of fully automated FT screening against reference standards. Laser AI uses question-based decision logic that translates eligibility criteria into screening questions and generates suggestions, reasoning, and links to relevant evidence within full-text publications. Performance was evaluated by comparing AI-generated inclusion and exclusion decisions with decisions reported in Cochrane reviews used as the reference standard.
METHODS: Five Cochrane reviews published in 2018 were used as a reference dataset. Full-text PDFs of publications classified as included or excluded in each review were collected. Eligibility criteria were copied verbatim from the published review reports and provided to Laser AI as the basis for automatically generated decision logic, without human review or refinement. Full-text publications were then processed using the generated workflow and compared with Cochrane inclusion/exclusion decisions. Seven reference inclusions with apparent discordance between the copied eligibility criteria and final review inclusion decisions were excluded from the reporting scenario. Performance was evaluated using recall (sensitivity), precision, and accuracy.
RESULTS: The reporting scenario included 204 publications, 17-114 per review; 138 were reference inclusions and 66 were reference exclusions. The automated workflow correctly identified 135/138 reference inclusions, yielding 97.8% sensitivity. It classified 150 publications as included, generating 15 false positives and 3 false negatives. Overall precision was 90.0%, range 33.3%-100.0%; accuracy was 91.2%, range 47.4%-100.0%.
CONCLUSIONS: Fully automated, eligibility-derived full-text screening achieved high sensitivity and precision in this five-review reference dataset after adjudicating criteria-label discordance. Remaining errors suggest that further improvements should focus on criteria interpretation and context preservation. Future evaluations should be conducted prospectively using Study Within A Review (SWAR) design, allowing review authors to steer AI-generated decision logic to match their objectives.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR41

Topic

Methodological & Statistical Research, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×