WHICH ARTICLES STILL NEED HUMAN EYES? A VALIDATED, SCORE-TIERED FRAMEWORK FOR AI-ASSISTED LITERATURE SCREENING

Author(s)

Angeline Babitha Dhas, BS1, Diwyashri Govindarajaperumal, B.Pharm2, Aditi Bajpai, PharmD1, Meghan Oates-Zalesky, MSc3.
1MadeAi, Cambridge, MA, USA, 2MadeAi, Nagercoil, India, 3Chief Marketing Officer, MadeAi, Cambridge, MA, USA.
OBJECTIVES: Background: AI-assisted screening approaches include ranking, generating tags, and binary classification of articles during title and abstract review, yet none provide a validated mechanism for identifying which specific articles require human review. This gap forces a costly choice: validate the entire corpus, eliminating efficiency gains, or trust AI outputs outright, risking audit and reproducibility failures in regulatory, submission-grade evidence.To develop and validate an AI scoring methodology that tells reviewers exactly which articles require human validation, enabling review depth to be matched to accuracy needs across submission-grade, scoping, and rapid/internal reviews.
METHODS: Three completed systematic reviews (7,557 articles) were screened using MadeAi that independently evaluates each article against every predefined inclusion/exclusion criterion (e.g., PICOS), classifying each as relevant, irrelevant, doubtful, or unavailable from the title/abstract. These criterion-level judgments determine the article's final relevance classification as relevant or irrelevant and a 0-10 relevance score. AI classifications were compared against human-adjudicated decisions to map screening errors by score. Three validation tiers were defined: Tier 1, submission-grade (score 7-10); Tier 2, scoping (score 9-10); Tier 3, rapid/internal (score 10 only).
RESULTS: Reviewing only the Tier 3 used just 9-20% of each corpus (mean 14%), reaching 89-98% accuracy (mean 94%). Expanding to Tier 2 lifted accuracy to 92-99% (mean 96%) at 15-37% reviewed. Tier 1 reached 99-100% accuracy at 37-74% reviewed (mean 60%), meeting submission-grade rigor without full manual validation. These gains may compound with reported 55% reductions in reviewer hours from AI-augmented workflows.
CONCLUSIONS: Reviewers do not need to re-check an entire AI-screened corpus to trust it—they need to know which fraction to check. A criterion-level, score-stratified validation framework answers this directly, delivering submission-grade accuracy at a reduced manual effort and closing a gap unresolved by ranking-, tagging-, or classification-only AI screening approaches.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR54

Topic

Health Technology Assessment, Methodological & Statistical Research, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×