ARE AI-ASSISTED EVIDENCE REVIEWS HTA-READY? A TARGETED LITERATURE REVIEW USING THE ELEVATE-GENAI QUALITY FRAMEWORK

Author(s)

Sheena Singh, MS1, Wenxi Tang, MS2, Andrew Easton, MS1, Lin Zhan, PhD2.
1Cytel, London, United Kingdom, 2BeOne Medicines, Ltd, San Carlos, CA, USA.
OBJECTIVES: Large language models (LLMs) are increasingly applied to evidence synthesis (ES) tasks informing health technology assessment (HTA) submissions. However, whether these tools meet HTA reporting standards remains uncharacterized. This study conducted a targeted literature review to characterize the use of artificial intelligence (AI) in oncology HEOR, and to evaluate HTA readiness using the ELEVATE-GenAI framework.[1]
[1]Fleurence RL, Dawoud D, Bian J, et al. ELEVATE-GenAI: reporting guidelines for the use of large language models in health economics and outcomes research: an ISPOR Working Group report. Value Health. 2025;28(11):1611-1625. doi:10.1016/j.jval.2025.06.018
METHODS: A structured search of MEDLINE, Embase, and grey literature (January 2023-December 2025) identified studies evaluating AI applied to ES tasks in oncology populations. The ELEVATE-GenAI checklist was applied to eligible full-text publications, evaluating 10 reporting domains. Each domain was rated clearly reported (CR; three points), ambiguous (two points), or not reported (NR; one point). A score of 100% indicated that all applicable domains were CR, reflecting reporting completeness, not study quality or AI performance.
RESULTS: Forty-one oncology ES studies were identified; 17 LLM-based studies were eligible for ELEVATE-GenAI assessment. Overall scores ranged from 50-83% (median 75%). Accuracy assessment was the best-reported domain (88% CR), with model characteristics (76% CR), comprehensiveness (71% CR), and factuality verification (76% CR) similarly documented. Four domains critical to HTA acceptance demonstrated systematic under-reporting: calibration and uncertainty quantification (59% NR), deployment/efficiency context (59% NR), robustness testing (53% NR), and security/data privacy (94% NR). Fairness/bias monitoring was not reported in any study.
CONCLUSIONS: Although AI tools demonstrate meaningful efficiency gains in oncology ES, critical reporting gaps persist in domains such as calibration and uncertainty quantification, robustness testing, security/data privacy and fairness, and bias monitoring. Standardization of AI reporting practices, supported by frameworks such as ELEVATE-GenAI, is needed to close these gaps and enable AI-assisted evidence to meet HTA submission standards.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR149

Topic

Health Technology Assessment, Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas, Oncology

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×