AI DELPHI: A FRAMEWORK FOR BENCHMARKING THE WRITING QUALITY OF LARGE LANGUAGE MODELS (LLMS) IN HEALTH ECONOMICS AND OUTCOMES RESEARCH (HEOR) SETTINGS, A CANADIAN PROOF-OF-CONCEPT
Author(s)
Gabriel Tremblay, MSc, DPhil1, Sharada Harricharan, PharmD2.
1Untraceable AI, Levis, QC, Canada, 2Toronto, ON, Canada.
1Untraceable AI, Levis, QC, Canada, 2Toronto, ON, Canada.
OBJECTIVES: To develop and validate AI Delphi, a reproducible multi-criteria framework for evaluating AI performance in HEOR, and compare ten LLMs across five providers using Canadian HTA documentation.
METHODS: AI Delphi uses four blinded AI judges (Claude Opus 4.8+RAG, GPT 5.5+RAG, Gemini 3.1 Pro+RAG, Claude Sonnet 4.6+RAG) scoring outputs across nine criteria: page/token budget (5%), writing quality and HEOR/payer style (25%), market access/payer insights, economic evaluation, clinical evaluation, comparative effectiveness, disease burden/unmet need, missing key information, and hallucination (10% each). To reduce bias, self-preferential scores were excluded. Convergence was assessed iteratively: scores <5% were averaged; divergent criteria re-evaluated with cross-judge information; unresolved discordance underwent senior HEOR adjudication. Ten LLMs generated a formatted abstract and four-page payer-centric summary of a Canadian HTA report for sacituzumab govitecan in HR+/HER2− metastatic breast cancer post-endocrine therapy and at least two metastatic systemic therapies. Models included four Anthropic, two OpenAI, two Google, Kimi K2.6, and DeepSeek V4 Pro. Each was tested with and without HEOR-specific hybrid RAG anchored by a human-authored, Canadian Grounding Codex.
RESULTS: Anthropic and OpenAI achieved the highest provider-level scores (both 75.5/100), followed by DeepSeek/Kimi combined (69.2), then Google (66.5). Top-performing configurations were Claude Opus 4.8+RAG (79.1) and GPT 5.5+RAG (78.5). RAG improved high-reasoning models but yielded inconsistent or negligible gains for others. Writing quality varied, with Google models producing bullet-point outputs incompatible with deliverable standards. Missing key information was most heterogeneous, while hallucination rates were low. Delphi convergence occurred in Round 1 for 30.8%, Round 2 for 61.1%, and required human adjudication in 8.1%.
CONCLUSIONS: AI Delphi provides a reproducible framework to evaluate LLM performance in HEOR. High-reasoning models with domain-specific RAG generated the strongest payer-centric outputs, however RAG benefits were model-dependent. Formatting gaps demonstrate that HEOR AI evaluation must assess both structure and content. Future work should expand testing across HTA bodies and indications.
METHODS: AI Delphi uses four blinded AI judges (Claude Opus 4.8+RAG, GPT 5.5+RAG, Gemini 3.1 Pro+RAG, Claude Sonnet 4.6+RAG) scoring outputs across nine criteria: page/token budget (5%), writing quality and HEOR/payer style (25%), market access/payer insights, economic evaluation, clinical evaluation, comparative effectiveness, disease burden/unmet need, missing key information, and hallucination (10% each). To reduce bias, self-preferential scores were excluded. Convergence was assessed iteratively: scores <5% were averaged; divergent criteria re-evaluated with cross-judge information; unresolved discordance underwent senior HEOR adjudication. Ten LLMs generated a formatted abstract and four-page payer-centric summary of a Canadian HTA report for sacituzumab govitecan in HR+/HER2− metastatic breast cancer post-endocrine therapy and at least two metastatic systemic therapies. Models included four Anthropic, two OpenAI, two Google, Kimi K2.6, and DeepSeek V4 Pro. Each was tested with and without HEOR-specific hybrid RAG anchored by a human-authored, Canadian Grounding Codex.
RESULTS: Anthropic and OpenAI achieved the highest provider-level scores (both 75.5/100), followed by DeepSeek/Kimi combined (69.2), then Google (66.5). Top-performing configurations were Claude Opus 4.8+RAG (79.1) and GPT 5.5+RAG (78.5). RAG improved high-reasoning models but yielded inconsistent or negligible gains for others. Writing quality varied, with Google models producing bullet-point outputs incompatible with deliverable standards. Missing key information was most heterogeneous, while hallucination rates were low. Delphi convergence occurred in Round 1 for 30.8%, Round 2 for 61.1%, and required human adjudication in 8.1%.
CONCLUSIONS: AI Delphi provides a reproducible framework to evaluate LLM performance in HEOR. High-reasoning models with domain-specific RAG generated the strongest payer-centric outputs, however RAG benefits were model-dependent. Formatting gaps demonstrate that HEOR AI evaluation must assess both structure and content. Future work should expand testing across HTA bodies and indications.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
P18
Topic
Methodological & Statistical Research, Organizational Practices, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
Oncology