TOWARD AI-AUGMENTED SCREENING FOR HTA-RELEVANT EVIDENCE SYNTHESIS: VALIDATION OF A GOVERNED SINGLE-REVIEWER WORKFLOW
Author(s)
Mahafroz Khatib, PhD1, Radha Savner, MPharm1, Keerthi Chowdhary Yelavarthi, Sr., PharmD2.
1Syneos Health Pvt Ltd, Bengaluru, India, 2Syneos Health Pvt. Ltd, Hyderabad, India.
1Syneos Health Pvt Ltd, Bengaluru, India, 2Syneos Health Pvt. Ltd, Hyderabad, India.
OBJECTIVES: The growing complexity of HTA submissions has increased the need for more efficient evidence-generation pathways without compromising methodological rigor. We evaluated an AI-assisted single-reviewer screening workflow with predefined governance controls and assessed its impact on screening efficiency.
METHODS: This pilot validation study used screening datasets from clinical efficacy, humanistic burden, and economic burden reviews, with prior human reviewer decisions serving as the reference standard. A proprietary, domain-adapted large language model workflow supported one reviewer during title/abstract and full-text screening using protocol-defined eligibility criteria and restricted to approved materials, including the screening protocol, gold-standard and borderline examples.The workflow evaluated each eligibility criterion independently before generating an overall recommendation, allowing protocol rules to be applied consistently across heterogeneous study designs and evidence sources. For each citation, the model recommended inclusion, exclusion, or manual review and provided the reasoning underlying its recommendation. Screening decisions followed predefined protocol rules, were documented with supporting rationale, confidence assessments, and included mandatory manual review of borderline cases, supporting key HTA expectations.
RESULTS: In full-text screening, agreement with the reference standard ranged from 92% to 96%. Sensitivity was 89%-93%, specificity 94%-100%, positive predictive value 93%-100%, negative predictive value (NPV) 87%-95%, and productivity gains 87%-92%. In title/abstract screening, agreement between the human and the AI-assisted approach ranged from 91% to 96%, with sensitivity of 85%-89%, specificity of 94%-100%, NPV of 92%-96%, and productivity gains of 81%-92% across evidence domains. Performance metrics were consistent across clinical efficacy, clinical, humanistic, and economic burden datasets, suggesting that the workflow generalized well across diverse evidence-generation contexts.
CONCLUSIONS: The AI-assisted workflow showed high agreement with human reviewers while maintaining low false-exclusion rates, remaining within commonly accepted screening-performance benchmarks. These findings indicate that AI-assisted single-reviewer screening can improve efficiency while preserving transparent documentation and reviewer oversight. Similar governance principles may be applicable across other evidence-synthesis activities.
METHODS: This pilot validation study used screening datasets from clinical efficacy, humanistic burden, and economic burden reviews, with prior human reviewer decisions serving as the reference standard. A proprietary, domain-adapted large language model workflow supported one reviewer during title/abstract and full-text screening using protocol-defined eligibility criteria and restricted to approved materials, including the screening protocol, gold-standard and borderline examples.The workflow evaluated each eligibility criterion independently before generating an overall recommendation, allowing protocol rules to be applied consistently across heterogeneous study designs and evidence sources. For each citation, the model recommended inclusion, exclusion, or manual review and provided the reasoning underlying its recommendation. Screening decisions followed predefined protocol rules, were documented with supporting rationale, confidence assessments, and included mandatory manual review of borderline cases, supporting key HTA expectations.
RESULTS: In full-text screening, agreement with the reference standard ranged from 92% to 96%. Sensitivity was 89%-93%, specificity 94%-100%, positive predictive value 93%-100%, negative predictive value (NPV) 87%-95%, and productivity gains 87%-92%. In title/abstract screening, agreement between the human and the AI-assisted approach ranged from 91% to 96%, with sensitivity of 85%-89%, specificity of 94%-100%, NPV of 92%-96%, and productivity gains of 81%-92% across evidence domains. Performance metrics were consistent across clinical efficacy, clinical, humanistic, and economic burden datasets, suggesting that the workflow generalized well across diverse evidence-generation contexts.
CONCLUSIONS: The AI-assisted workflow showed high agreement with human reviewers while maintaining low false-exclusion rates, remaining within commonly accepted screening-performance benchmarks. These findings indicate that AI-assisted single-reviewer screening can improve efficiency while preserving transparent documentation and reviewer oversight. Similar governance principles may be applicable across other evidence-synthesis activities.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR247
Topic
Health Technology Assessment, Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics