DOING MORE WITH LESS: STRATEGIES TO REDUCE THE COST OF GENERATIVE AI IN HEOR

Author(s)

Oliver Pople, MSc, Sarah Cudworth, MRes.
Estima Scientific, London, United Kingdom.
OBJECTIVES: Advances in reasoning capabilities within modern Large Language Models (LLMs) are rapidly expanding their potential applications. However, implementing these models is becoming more expensive. Costs are driven by both higher per-token pricing of advanced models and greater token utilization associated with complex reasoning tasks. Within Health Economics and Outcomes Research (HEOR), there is growing interest in leveraging these technologies to automate and enhance project workflows. Therefore, reducing token consumption and associated costs is critical to the scalable adoption of generative AI.
METHODS: A targeted review was conducted to evaluate four approaches for reducing token consumption in HEOR workflows:
Prompt Caching saves processed tokens locally, enabling the LLM to reference existing context.
Smart Model Routing assigns tasks to models based on complexity - simpler tasks utilise lower-cost models.
Batch Processing leverages discounted pricing if LLM requests are submitted in bulk for deferred processing.
Result caching stores outputs from previously completed tasks, allowing reuse of previous output without generating new LLM calls.
RESULTS: The review identified multiple opportunities to reduce LLM-related costs through a combination of technical and operational optimization strategies. These approaches minimize redundant token processing, reduce repeated analysis of identical information, and improve alignment between task complexity and model capability. Collectively, the identified strategies address both key drivers of LLM expenditure by decreasing per-token costs of advanced reasoning models and reducing token utilization required for complex analytical tasks. The potential benefits are particularly relevant for HEOR workflows, where large document sets, iterative analyses, and repeated queries can substantially increase token consumption over time.
CONCLUSIONS: Optimization strategies can substantially improve the cost-efficiency of generative AI within HEOR workflows. As the use of generative AI expands across HEOR, proactive management of costs will become increasingly important to support the long-term scalability of automation.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR174

Topic

Methodological & Statistical Research, Organizational Practices, Study Approaches

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×