SETTING THE BAR WITHOUT A NUMBER: HTA AND GUIDELINE EXPECTATIONS FOR AI IN SLRS
Author(s)
Stephan N. Martin, MPH, Alyssa Simon, MPH, Grace E. Fox, PhD.
OPEN Health HEOR & Market Access, New York, NY, USA.
OPEN Health HEOR & Market Access, New York, NY, USA.
OBJECTIVES: With artificial intelligence (AI) increasingly used in systematic literature reviews (SLRs), we aimed to identify and synthesize guidance from health technology assessment (HTA) bodies and other organizations regarding assessment of AI performance across all SLR tasks.
METHODS: An AI web crawler (Claude Sonnet 4.6) was used to update a rapid literature review of methodological standards on AI usage in SLRs, to include any published in 2025-2026. Sources included reporting guidelines (e.g., PRISMA), methodological frameworks (i.e., RAISE guidelines), and position statements from evidence synthesis organizations (i.e., Cochrane, Campbell, CEE, JBI) and the WHO. Information was extracted on performance metrics and acceptable performance thresholds for AI used in any SLR task (e.g., title/abstract screening).
RESULTS: Over the period of interest, 4 major evidence synthesis organizations issued a joint position statement endorsing RAISE guidelines, which mandate human oversight, transparent reporting, and context-specific validation. Similarly, ELEVATE-GenAI and PRISMA-trAIce checklists both recommend that authors report appropriate performance metrics and validation methods of any AI tool used. A WHO policy discussion situates these recommendations within a governance architecture endorsing AI augmentation over AI automation in evidence synthesis, emphasizing that AI “complements rather than replaces human judgment,” while noting that AI will accelerate the paradigm shift toward living SLRs as standard practice.
CONCLUSIONS: HTA bodies have still not reported specific numerical thresholds for performance metrics. The consensus among NICE, CDA-AMC, and IQWiG remains that AI tools must demonstrate non-inferiority to dual-human processes and very high sensitivity, ensuring no evidence is missed, even at the cost of lower efficiency. In the absence of specific thresholds, we advise those using AI for HTA-grade SLRs to prioritize high sensitivity, transparent reporting of methodologies and validation, and human-in-the-loop processes to align with emerging regulatory standards while improving efficiency. In such an approach, AI augments rather than replaces humans’ role in conducting SLRs.
METHODS: An AI web crawler (Claude Sonnet 4.6) was used to update a rapid literature review of methodological standards on AI usage in SLRs, to include any published in 2025-2026. Sources included reporting guidelines (e.g., PRISMA), methodological frameworks (i.e., RAISE guidelines), and position statements from evidence synthesis organizations (i.e., Cochrane, Campbell, CEE, JBI) and the WHO. Information was extracted on performance metrics and acceptable performance thresholds for AI used in any SLR task (e.g., title/abstract screening).
RESULTS: Over the period of interest, 4 major evidence synthesis organizations issued a joint position statement endorsing RAISE guidelines, which mandate human oversight, transparent reporting, and context-specific validation. Similarly, ELEVATE-GenAI and PRISMA-trAIce checklists both recommend that authors report appropriate performance metrics and validation methods of any AI tool used. A WHO policy discussion situates these recommendations within a governance architecture endorsing AI augmentation over AI automation in evidence synthesis, emphasizing that AI “complements rather than replaces human judgment,” while noting that AI will accelerate the paradigm shift toward living SLRs as standard practice.
CONCLUSIONS: HTA bodies have still not reported specific numerical thresholds for performance metrics. The consensus among NICE, CDA-AMC, and IQWiG remains that AI tools must demonstrate non-inferiority to dual-human processes and very high sensitivity, ensuring no evidence is missed, even at the cost of lower efficiency. In the absence of specific thresholds, we advise those using AI for HTA-grade SLRs to prioritize high sensitivity, transparent reporting of methodologies and validation, and human-in-the-loop processes to align with emerging regulatory standards while improving efficiency. In such an approach, AI augments rather than replaces humans’ role in conducting SLRs.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
SA17
Topic
Study Approaches
Topic Subcategory
Literature Review & Synthesis
Disease
No Additional Disease & Conditions/Specialized Treatment Areas