HTA-BENCH: A HUMAN-VALIDATED BENCHMARK FOR AI GENERATION OF EUROPEAN HTA AND MARKET-ACCESS DOCUMENTS
Author(s)
Anton O. Wiehe, M.Sc.1, Pia A. Cuk, M.Sc.1, Susanne Schneller, Ph.D.2, Florian Woeste, M.Sc.1.
1Pharos Labs, Hamburg, Germany, 2SmartStep Consulting, Hamburg, Germany.
1Pharos Labs, Hamburg, Germany, 2SmartStep Consulting, Hamburg, Germany.
OBJECTIVES: AI is increasingly used to draft HTA deliverables such as AMNOG dossier modules, G-BA scientific advice, and comparator-therapy (zVT) predictions. Existing regulatory LLM benchmarks evaluate question answering, not document generation, and output quality depends on the whole system, model plus agentic harness, not the model alone. We introduce HTA-Bench, to our knowledge the first benchmark ranking AI systems on end-to-end generation of real HTA deliverables against expert ground truth via a human-calibrated LLM judge.
METHODS: The dataset comprises seven authentic German AMNOG tasks of increasing difficulty, each pairing real inputs (study reports, protocols, guidelines) with an expert gold output: G-BA scientific-advice requests, zVT predictions, IQWiG-statement responses, and full dossier modules (2 to 4); JCA and EU tasks are in development. To separate model from harness, two system classes, a domain-specialized medical writing agent and a general-purpose assistant, each run on two frontier LLMs, give four configurations from identical inputs. A blinded LLM judge scores outputs per section for factual accuracy and writing quality. It is calibrated to expert raters whose criteria are encoded, with agreement measured on held-out comparisons (per ELEVATE-GenAI).
RESULTS: We report feasibility and judge calibration; system scoring is underway. The dataset and end-to-end generation pipeline are operational across all task types. In the calibration pilot, judge to human agreement reached κ=0.74 (87% concordance), up from κ=0.51, supporting the judge as a scalable proxy for expert assessment. System-level scoring, separating base-model from agentic-harness contributions at identical inputs, is in progress.
CONCLUSIONS: HTA-Bench is an expert-grounded, human-validated benchmark for end-to-end generation of real HTA documents, with a judge that agrees closely with expert raters. It tests whether the agentic harness, not the model alone, drives output quality on authentic market-access deliverables, which question-answering benchmarks cannot assess. Results are calibration-stage; full system scoring, expanded tasks and models, and JCA coverage are ongoing.
METHODS: The dataset comprises seven authentic German AMNOG tasks of increasing difficulty, each pairing real inputs (study reports, protocols, guidelines) with an expert gold output: G-BA scientific-advice requests, zVT predictions, IQWiG-statement responses, and full dossier modules (2 to 4); JCA and EU tasks are in development. To separate model from harness, two system classes, a domain-specialized medical writing agent and a general-purpose assistant, each run on two frontier LLMs, give four configurations from identical inputs. A blinded LLM judge scores outputs per section for factual accuracy and writing quality. It is calibrated to expert raters whose criteria are encoded, with agreement measured on held-out comparisons (per ELEVATE-GenAI).
RESULTS: We report feasibility and judge calibration; system scoring is underway. The dataset and end-to-end generation pipeline are operational across all task types. In the calibration pilot, judge to human agreement reached κ=0.74 (87% concordance), up from κ=0.51, supporting the judge as a scalable proxy for expert assessment. System-level scoring, separating base-model from agentic-harness contributions at identical inputs, is in progress.
CONCLUSIONS: HTA-Bench is an expert-grounded, human-validated benchmark for end-to-end generation of real HTA documents, with a judge that agrees closely with expert raters. It tests whether the agentic harness, not the model alone, drives output quality on authentic market-access deliverables, which question-answering benchmarks cannot assess. Results are calibration-stage; full system scoring, expanded tasks and models, and JCA coverage are ongoing.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
HTA86
Topic
Health Technology Assessment, Methodological & Statistical Research, Real World Data & Information Systems
Topic Subcategory
Value Frameworks & Dossier Format