HTA-BENCH: A HUMAN-VALIDATED BENCHMARK FOR AI GENERATION OF EUROPEAN HTA AND MARKET-ACCESS DOCUMENTS

Author(s)

Anton O. Wiehe, M.Sc.1, Pia A. Cuk, M.Sc.1, Susanne Schneller, Ph.D.2, Florian Woeste, M.Sc.1.
1Pharos Labs, Hamburg, Germany, 2SmartStep Consulting, Hamburg, Germany.
OBJECTIVES: AI is increasingly used to draft HTA deliverables such as AMNOG dossier modules, G-BA scientific advice, and comparator-therapy (zVT) predictions. Existing regulatory LLM benchmarks evaluate question answering, not document generation, and output quality depends on the whole system, model plus agentic harness, not the model alone. We introduce HTA-Bench, to our knowledge the first benchmark ranking AI systems on end-to-end generation of real HTA deliverables against expert ground truth via a human-calibrated LLM judge.
METHODS: The dataset comprises seven authentic German AMNOG tasks of increasing difficulty, each pairing real inputs (study reports, protocols, guidelines) with an expert gold output: G-BA scientific-advice requests, zVT predictions, IQWiG-statement responses, and full dossier modules (2 to 4); JCA and EU tasks are in development. To separate model from harness, two system classes, a domain-specialized medical writing agent and a general-purpose assistant, each run on two frontier LLMs, give four configurations from identical inputs. A blinded LLM judge scores outputs per section for factual accuracy and writing quality. It is calibrated to expert raters whose criteria are encoded, with agreement measured on held-out comparisons (per ELEVATE-GenAI).
RESULTS: We report feasibility and judge calibration; system scoring is underway. The dataset and end-to-end generation pipeline are operational across all task types. In the calibration pilot, judge to human agreement reached κ=0.74 (87% concordance), up from κ=0.51, supporting the judge as a scalable proxy for expert assessment. System-level scoring, separating base-model from agentic-harness contributions at identical inputs, is in progress.
CONCLUSIONS: HTA-Bench is an expert-grounded, human-validated benchmark for end-to-end generation of real HTA documents, with a judge that agrees closely with expert raters. It tests whether the agentic harness, not the model alone, drives output quality on authentic market-access deliverables, which question-answering benchmarks cannot assess. Results are calibration-stage; full system scoring, expanded tasks and models, and JCA coverage are ongoing.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

HTA86

Topic

Health Technology Assessment, Methodological & Statistical Research, Real World Data & Information Systems

Topic Subcategory

Value Frameworks & Dossier Format

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×