ARE AI-ASSISTED EVIDENCE REVIEWS HTA-READY? A TARGETED LITERATURE REVIEW USING THE ELEVATE-GENAI QUALITY FRAMEWORK
Author(s)
Sheena Singh, MS1, Wenxi Tang, MS2, Andrew Easton, MS1, Lin Zhan, PhD2.
1Cytel, London, United Kingdom, 2BeOne Medicines, Ltd, San Carlos, CA, USA.
1Cytel, London, United Kingdom, 2BeOne Medicines, Ltd, San Carlos, CA, USA.
OBJECTIVES: Large language models (LLMs) are increasingly applied to evidence synthesis (ES) tasks informing health technology assessment (HTA) submissions. However, whether these tools meet HTA reporting standards remains uncharacterized. This study conducted a targeted literature review to characterize the use of artificial intelligence (AI) in oncology HEOR, and to evaluate HTA readiness using the ELEVATE-GenAI framework.[1]
[1]Fleurence RL, Dawoud D, Bian J, et al. ELEVATE-GenAI: reporting guidelines for the use of large language models in health economics and outcomes research: an ISPOR Working Group report. Value Health. 2025;28(11):1611-1625. doi:10.1016/j.jval.2025.06.018
METHODS: A structured search of MEDLINE, Embase, and grey literature (January 2023-December 2025) identified studies evaluating AI applied to ES tasks in oncology populations. The ELEVATE-GenAI checklist was applied to eligible full-text publications, evaluating 10 reporting domains. Each domain was rated clearly reported (CR; three points), ambiguous (two points), or not reported (NR; one point). A score of 100% indicated that all applicable domains were CR, reflecting reporting completeness, not study quality or AI performance.
RESULTS: Forty-one oncology ES studies were identified; 17 LLM-based studies were eligible for ELEVATE-GenAI assessment. Overall scores ranged from 50-83% (median 75%). Accuracy assessment was the best-reported domain (88% CR), with model characteristics (76% CR), comprehensiveness (71% CR), and factuality verification (76% CR) similarly documented. Four domains critical to HTA acceptance demonstrated systematic under-reporting: calibration and uncertainty quantification (59% NR), deployment/efficiency context (59% NR), robustness testing (53% NR), and security/data privacy (94% NR). Fairness/bias monitoring was not reported in any study.
CONCLUSIONS: Although AI tools demonstrate meaningful efficiency gains in oncology ES, critical reporting gaps persist in domains such as calibration and uncertainty quantification, robustness testing, security/data privacy and fairness, and bias monitoring. Standardization of AI reporting practices, supported by frameworks such as ELEVATE-GenAI, is needed to close these gaps and enable AI-assisted evidence to meet HTA submission standards.
[1]Fleurence RL, Dawoud D, Bian J, et al. ELEVATE-GenAI: reporting guidelines for the use of large language models in health economics and outcomes research: an ISPOR Working Group report. Value Health. 2025;28(11):1611-1625. doi:10.1016/j.jval.2025.06.018
METHODS: A structured search of MEDLINE, Embase, and grey literature (January 2023-December 2025) identified studies evaluating AI applied to ES tasks in oncology populations. The ELEVATE-GenAI checklist was applied to eligible full-text publications, evaluating 10 reporting domains. Each domain was rated clearly reported (CR; three points), ambiguous (two points), or not reported (NR; one point). A score of 100% indicated that all applicable domains were CR, reflecting reporting completeness, not study quality or AI performance.
RESULTS: Forty-one oncology ES studies were identified; 17 LLM-based studies were eligible for ELEVATE-GenAI assessment. Overall scores ranged from 50-83% (median 75%). Accuracy assessment was the best-reported domain (88% CR), with model characteristics (76% CR), comprehensiveness (71% CR), and factuality verification (76% CR) similarly documented. Four domains critical to HTA acceptance demonstrated systematic under-reporting: calibration and uncertainty quantification (59% NR), deployment/efficiency context (59% NR), robustness testing (53% NR), and security/data privacy (94% NR). Fairness/bias monitoring was not reported in any study.
CONCLUSIONS: Although AI tools demonstrate meaningful efficiency gains in oncology ES, critical reporting gaps persist in domains such as calibration and uncertainty quantification, robustness testing, security/data privacy and fairness, and bias monitoring. Standardization of AI reporting practices, supported by frameworks such as ELEVATE-GenAI, is needed to close these gaps and enable AI-assisted evidence to meet HTA submission standards.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR149
Topic
Health Technology Assessment, Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas, Oncology