EVALUATING FULLY AUTOMATED FULL-TEXT SCREENING FOR HTA SYSTEMATIC REVIEWS
Author(s)
Artur Nowak, MSc, Ewelina Sadowska, MPharm, Ewa Borowiack, MSc, Monika Opalek, PhD.
Evidence Prime, Krakow, Poland.
Evidence Prime, Krakow, Poland.
OBJECTIVES: Full-text (FT) screening is a critical but resource-intensive step in systematic reviews for health technology assessment (HTA). As part of a broader program evaluating AI-assisted full-text review, this study focused on technical validation of fully automated FT screening against reference standards. Laser AI uses question-based decision logic that translates eligibility criteria into screening questions and generates suggestions, reasoning, and links to relevant evidence within full-text publications. Performance was evaluated by comparing AI-generated inclusion and exclusion decisions with decisions reported in Cochrane reviews used as the reference standard.
METHODS: Five Cochrane reviews published in 2018 were used as a reference dataset. Full-text PDFs of publications classified as included or excluded in each review were collected. Eligibility criteria were copied verbatim from the published review reports and provided to Laser AI as the basis for automatically generated decision logic, without human review or refinement. Full-text publications were then processed using the generated workflow and compared with Cochrane inclusion/exclusion decisions. Seven reference inclusions with apparent discordance between the copied eligibility criteria and final review inclusion decisions were excluded from the reporting scenario. Performance was evaluated using recall (sensitivity), precision, and accuracy.
RESULTS: The reporting scenario included 204 publications, 17-114 per review; 138 were reference inclusions and 66 were reference exclusions. The automated workflow correctly identified 135/138 reference inclusions, yielding 97.8% sensitivity. It classified 150 publications as included, generating 15 false positives and 3 false negatives. Overall precision was 90.0%, range 33.3%-100.0%; accuracy was 91.2%, range 47.4%-100.0%.
CONCLUSIONS: Fully automated, eligibility-derived full-text screening achieved high sensitivity and precision in this five-review reference dataset after adjudicating criteria-label discordance. Remaining errors suggest that further improvements should focus on criteria interpretation and context preservation. Future evaluations should be conducted prospectively using Study Within A Review (SWAR) design, allowing review authors to steer AI-generated decision logic to match their objectives.
METHODS: Five Cochrane reviews published in 2018 were used as a reference dataset. Full-text PDFs of publications classified as included or excluded in each review were collected. Eligibility criteria were copied verbatim from the published review reports and provided to Laser AI as the basis for automatically generated decision logic, without human review or refinement. Full-text publications were then processed using the generated workflow and compared with Cochrane inclusion/exclusion decisions. Seven reference inclusions with apparent discordance between the copied eligibility criteria and final review inclusion decisions were excluded from the reporting scenario. Performance was evaluated using recall (sensitivity), precision, and accuracy.
RESULTS: The reporting scenario included 204 publications, 17-114 per review; 138 were reference inclusions and 66 were reference exclusions. The automated workflow correctly identified 135/138 reference inclusions, yielding 97.8% sensitivity. It classified 150 publications as included, generating 15 false positives and 3 false negatives. Overall precision was 90.0%, range 33.3%-100.0%; accuracy was 91.2%, range 47.4%-100.0%.
CONCLUSIONS: Fully automated, eligibility-derived full-text screening achieved high sensitivity and precision in this five-review reference dataset after adjudicating criteria-label discordance. Remaining errors suggest that further improvements should focus on criteria interpretation and context preservation. Future evaluations should be conducted prospectively using Study Within A Review (SWAR) design, allowing review authors to steer AI-generated decision logic to match their objectives.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR41
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas