EVALUATION OF AN AI-AUGMENTED LITERATURE REVIEW PLATFORM IN TARGETED LITERATURE REVIEWS - A PRACTICAL INDUSTRY HEOR PERSPECTIVE
Author(s)
Jack Timmons, PharmD1, Uche Mordi, PharmD, MS1, JEANPIERRE COAQUIRA CASTRO, MPH1, SAEID SHAHRAZ, MD, PhD1, Yumi Asukai, MSc2.
1Gilead Sciences, Foster City, CA, USA, 2Gilead Sciences, London, United Kingdom.
1Gilead Sciences, Foster City, CA, USA, 2Gilead Sciences, London, United Kingdom.
OBJECTIVES: Full-scale adoption of Artificial Intelligence (AI) augmented literature review platforms in industry remains limited due to concerns around health technology assessment (HTA) acceptability, copyright constraints, and operational feasibility. However, these platforms may provide near-term value by supporting targeted literature reviews (TLRs). This study benchmarked performance between an AI platform and a large language model (LLM) against original manual outputs (manual) by replicating four TLRs.
METHODS: An AI platform and a LLM were used to replicate four completed TLRs across key domains: health economics (HE), indirect treatment comparison (ITC), real-world evidence (RWE), and clinical outcomes assessment (COA). The evaluation focused on AI-augmented search and title/abstract screening, with selective manual full-text screening. Performance was assessed in each TLR using recall, precision, accuracy and time savings. Recall was defined as the proportion of included studies over the pooled reference set of included studies from all three TLR approaches.
RESULTS: The AI platform demonstrated higher or equivalent recall across TLRs. In HE, recall was 67% versus 33% for LLM and manual search; in RWE, 95% versus 58% (manual) and 11% (LLM); in COA, 94% versus 100% (manual) and 28% (LLM); and in ITC, 100% versus 56% (manual) and 33% (LLM). AI-augmented search yielded high precision across each use case, with 100% for HE, 90% for RWE, 89% for COA, and 100% for ITC. Copyright restrictions limited performance, with open access articles representing as little as 20% of included studies in some cases.
CONCLUSIONS: Despite copyright-related limitations, the AI platform showed improved performance relative to manual search and LLMs. These findings support the potential for AI-augmented literature review platforms to deliver short-term value on lower impact applications such as TLRs, facilitate familiarity and adoption within HEOR teams, pending resolution on key barriers required for full scale adoption and end-to-end implementation in HTA evidence generation workflows.
METHODS: An AI platform and a LLM were used to replicate four completed TLRs across key domains: health economics (HE), indirect treatment comparison (ITC), real-world evidence (RWE), and clinical outcomes assessment (COA). The evaluation focused on AI-augmented search and title/abstract screening, with selective manual full-text screening. Performance was assessed in each TLR using recall, precision, accuracy and time savings. Recall was defined as the proportion of included studies over the pooled reference set of included studies from all three TLR approaches.
RESULTS: The AI platform demonstrated higher or equivalent recall across TLRs. In HE, recall was 67% versus 33% for LLM and manual search; in RWE, 95% versus 58% (manual) and 11% (LLM); in COA, 94% versus 100% (manual) and 28% (LLM); and in ITC, 100% versus 56% (manual) and 33% (LLM). AI-augmented search yielded high precision across each use case, with 100% for HE, 90% for RWE, 89% for COA, and 100% for ITC. Copyright restrictions limited performance, with open access articles representing as little as 20% of included studies in some cases.
CONCLUSIONS: Despite copyright-related limitations, the AI platform showed improved performance relative to manual search and LLMs. These findings support the potential for AI-augmented literature review platforms to deliver short-term value on lower impact applications such as TLRs, facilitate familiarity and adoption within HEOR teams, pending resolution on key barriers required for full scale adoption and end-to-end implementation in HTA evidence generation workflows.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR168
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas