DEEP CHARACTERIZATION OF CLINICAL TRIAL POPULATIONS: BENCHMARKING A HIGH-GRANULARITY GENERATIVE ARTIFICIAL INTELLIGENCE FRAMEWORK FOR EU JOINT CLINICAL ASSESSMENTS
Author(s)
Manuel Cossio1, Ramiro Eugenio Gilardino, MSc, MD2.
1Head of AI Solutions, Cytel, Dubendorf, Switzerland, 2Insights & Impact, Zurich, Switzerland.
1Head of AI Solutions, Cytel, Dubendorf, Switzerland, 2Insights & Impact, Zurich, Switzerland.
OBJECTIVES: The Joint Clinical Assessment process under the European Union Health Technology Assessment Regulation (EU HTAR) relies on a statutory survey-driven scoping mechanism where member-state requirements are consolidated by the Assessor and Co-Assessor into a unified PICO matrix. Navigating these complex multi-country demands requires Health Technology Developers (HTDs) to perform meticulous extractions of granular clinical trial subpopulation data. While generic AI solutions extract high-level variables, they systematically fail to capture granular data layers. This study benchmarks a specialized Generative AI framework engineered to overcome these extraction limitations.
METHODS: A zero-shot, role-enforced prompt was engineered based on an expert HTA analyst persona, mandating a strict 19-row validation schema targeting granular variables omitted by generic tools—including exact diagnostic testing methodologies for biomarker confirmation, detailed age/race subsets, and explicit ECOG distributions—with zero conversational text. The framework evaluated six frontier Large Language Models: Anthropic Claude 3.5 Sonnet, Meta Llama 3, OpenAI GPT-4o, Google Gemini 1.5 Pro, DeepSeek-V3, and Mistral Large, using the primary phase III LIBRETTO-531 trial manuscript as a proxy benchmark for complex, biomarker-stratified data environments.
RESULTS: No hallucinations were observed across models. Accuracy varied significantly based on completeness. Claude 3.5 Sonnet and Llama 3 achieved perfect accuracy scores (1.00), successfully extracting all 19 schema rows. Crucially, they captured the exact nested variables conventional tools miss: specific laboratory test boundaries (PCR/NGS), control-arm safety vs. ITT denominator splits, and multi-variable subgroup outcomes (events/median PFS across age, sex, race, and sub-mutations). GPT-4o achieved moderate completeness (0.76). Gemini, DeepSeek, and Mistral failed on exhaustiveness (0.59) due to persistent omissions of granular context.
CONCLUSIONS: Generic AI solutions are insufficient for navigating consolidated EU JCA pipelines. Highly constrained framework architectures utilizing frontier models provide the semantic precision required for exhaustive, audit-ready subpopulation mapping, potentially reducing the operational burden of multi-country PICO alignment for HTDs.
METHODS: A zero-shot, role-enforced prompt was engineered based on an expert HTA analyst persona, mandating a strict 19-row validation schema targeting granular variables omitted by generic tools—including exact diagnostic testing methodologies for biomarker confirmation, detailed age/race subsets, and explicit ECOG distributions—with zero conversational text. The framework evaluated six frontier Large Language Models: Anthropic Claude 3.5 Sonnet, Meta Llama 3, OpenAI GPT-4o, Google Gemini 1.5 Pro, DeepSeek-V3, and Mistral Large, using the primary phase III LIBRETTO-531 trial manuscript as a proxy benchmark for complex, biomarker-stratified data environments.
RESULTS: No hallucinations were observed across models. Accuracy varied significantly based on completeness. Claude 3.5 Sonnet and Llama 3 achieved perfect accuracy scores (1.00), successfully extracting all 19 schema rows. Crucially, they captured the exact nested variables conventional tools miss: specific laboratory test boundaries (PCR/NGS), control-arm safety vs. ITT denominator splits, and multi-variable subgroup outcomes (events/median PFS across age, sex, race, and sub-mutations). GPT-4o achieved moderate completeness (0.76). Gemini, DeepSeek, and Mistral failed on exhaustiveness (0.59) due to persistent omissions of granular context.
CONCLUSIONS: Generic AI solutions are insufficient for navigating consolidated EU JCA pipelines. Highly constrained framework architectures utilizing frontier models provide the semantic precision required for exhaustive, audit-ready subpopulation mapping, potentially reducing the operational burden of multi-country PICO alignment for HTDs.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
HTA186
Topic
Health Technology Assessment, Real World Data & Information Systems
Topic Subcategory
Decision & Deliberative Processes
Disease
Oncology