SPEED, SCRUTINY, AND TRUST: A FRAMEWORK FOR EVALUATING AI-SUPPORTED SLR TOOLS AGAINST HTA REQUIREMENTS
Author(s)
Minoo Mazaheri, PharmD, MPH, MSc, Veena Jia Wen Lim, MSc, MD, Laura Sawyer, BA, MSc, Alex Diamantopoulos, MSc.
Symmetron, London, United Kingdom.
Symmetron, London, United Kingdom.
OBJECTIVES: AI-supported tools are increasingly promoted to improve the efficiency of systematic literature review (SLR), but their use in health technology assessment (HTA) raises practical questions about transparency, reproducibility, auditability, and human oversight. Researchers need a structured approach to weigh these trade-offs when selecting SLR tools. This study developed and piloted an exploratory framework to evaluate AI-supported SLR tools against HTA requirements and identify considerations for tool selection and workflow integration.
METHODS: The framework was developed through a review of guidance on responsible AI in evidence synthesis, including the RAISE framework, and HTA agency guidance on AI use in reimbursement submissions, alongside practical requirements identified by researchers selecting SLR tools for HTA or HEOR. Recurring themes were synthesised into candidate domains, then refined into explicit criteria with predefined scoring anchors. The framework was then pilot-applied to selected tools using publicly available documentation to assess whether criteria were consistently interpreted, whether available information supported scoring, and whether efficiency claims could be weighed against documentation and audit requirements.
RESULTS: The framework comprised four domains: workflow coverage; collaboration and customisation; documentation, traceability and auditability; and responsible AI integration. The criteria captured HTA-facing requirements such as end-to-end review support, decision-level traceability, audit trails, human oversight and accessible evidence on AI validation. The pilot suggested the framework was appropriate for structured comparison, while highlighting variation in the completeness of publicly available documentation. The framework helped distinguish documented capabilities from claims that could not be independently assessed from public materials.
CONCLUSIONS: The framework provides a structured approach to evaluate AI-supported SLR tools before integration into HTA workflows. The framework’s value lies in making the efficiency-versus-trust trade-off explicit: tools may reduce operational burden only if their workflows remain sufficiently transparent, reproducible and auditable. Further work should validate the framework with its intended users and assess scoring consistency across independent assessors.
METHODS: The framework was developed through a review of guidance on responsible AI in evidence synthesis, including the RAISE framework, and HTA agency guidance on AI use in reimbursement submissions, alongside practical requirements identified by researchers selecting SLR tools for HTA or HEOR. Recurring themes were synthesised into candidate domains, then refined into explicit criteria with predefined scoring anchors. The framework was then pilot-applied to selected tools using publicly available documentation to assess whether criteria were consistently interpreted, whether available information supported scoring, and whether efficiency claims could be weighed against documentation and audit requirements.
RESULTS: The framework comprised four domains: workflow coverage; collaboration and customisation; documentation, traceability and auditability; and responsible AI integration. The criteria captured HTA-facing requirements such as end-to-end review support, decision-level traceability, audit trails, human oversight and accessible evidence on AI validation. The pilot suggested the framework was appropriate for structured comparison, while highlighting variation in the completeness of publicly available documentation. The framework helped distinguish documented capabilities from claims that could not be independently assessed from public materials.
CONCLUSIONS: The framework provides a structured approach to evaluate AI-supported SLR tools before integration into HTA workflows. The framework’s value lies in making the efficiency-versus-trust trade-off explicit: tools may reduce operational burden only if their workflows remain sufficiently transparent, reproducible and auditable. Further work should validate the framework with its intended users and assess scoring consistency across independent assessors.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR211
Topic
Health Technology Assessment, Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas