PICO PROVENANCE EVALUATING AI TRACEABILITY FOR EU JOINT CLINICAL ASSESSMENTS
Author(s)
Rhythm Arora, Masters in Pharmacy1, Hanna-Liisa Vilu, Master's degree, Molecular and Cellular Biochemist2, Victoria L Molenkamp, MSc Health Psychology3.
1Associate Consultant, Evidera Ltd., Part of PPD, Thermo Fisher Scientific, Hyderabad, India, 2Associate Director, Evidera Ltd., Part of PPD, Thermo Fisher Scientific, Tallinn, Estonia, 3Sr Market Access Writer, Evidera Ltd., Part of PPD, Thermo Fisher Scientific, London, United Kingdom.
1Associate Consultant, Evidera Ltd., Part of PPD, Thermo Fisher Scientific, Hyderabad, India, 2Associate Director, Evidera Ltd., Part of PPD, Thermo Fisher Scientific, Tallinn, Estonia, 3Sr Market Access Writer, Evidera Ltd., Part of PPD, Thermo Fisher Scientific, London, United Kingdom.
OBJECTIVES: Health technology developers increasingly use generative artificial intelligence (AI) tools to anticipate Population, Intervention, Comparator, and Outcome (PICO) scenarios ahead of EU Joint Clinical Assessment (JCA). While prior research has benchmarked AI-generated PICO counts and relevance against manual scoping, the traceability of AI-generated claims for predicted PICOs remains a gap in existing validation frameworks. This study evaluated AI-generated PICO outputs to determine whether the supporting evidence for each predicted PICO was verifiable, accurate, and supported by the cited source.
METHODS: The custom GPT was built on OpenAI GPT-5.5, with instructions operationalizing the PICO framework and a likelihood-rating system for predicting comparator inclusion logic, supplemented by example PICO outputs and EU HTA/JCA methodology reference files. Combining a pre-loaded national clinical guideline with open web browsing for payer and sponsor sources, the tool generated a single-country (Sweden) PICO grid for an investigational biologic across two indications (Crohn's Disease and Ulcerative Colitis). Each cited source (n=18) was independently retrieved and compared against the tool's stated claim using a three-point rubric: 0 (unverifiable or fabricated), 1 (verifiable but inaccurate or overstated), 2 (verifiable and accurate).
RESULTS: Overall traceability reached 89% (16/18 cited sources verifiable and accurate), rising to 100% for claims drawn from the uploaded reference guideline across both indications. Accuracy of browsed-source citations diverged sharply by indication: 57% for Crohn's disease (containing the only fabrication and both overstated-specificity cases) versus 100% for Ulcerative colitis, despite identical tool configuration and source types.
CONCLUSIONS: This analysis shows AI-generated PICO predictions can achieve high traceability when grounded in pre-specified source materials. However, inaccuracies and fabricated claims in externally sourced evidence show that source fidelity is not guaranteed by default and depends on tool configuration, highlighting the need for systematic manual source validation as part of AI-enabled JCA preparation workflows.
METHODS: The custom GPT was built on OpenAI GPT-5.5, with instructions operationalizing the PICO framework and a likelihood-rating system for predicting comparator inclusion logic, supplemented by example PICO outputs and EU HTA/JCA methodology reference files. Combining a pre-loaded national clinical guideline with open web browsing for payer and sponsor sources, the tool generated a single-country (Sweden) PICO grid for an investigational biologic across two indications (Crohn's Disease and Ulcerative Colitis). Each cited source (n=18) was independently retrieved and compared against the tool's stated claim using a three-point rubric: 0 (unverifiable or fabricated), 1 (verifiable but inaccurate or overstated), 2 (verifiable and accurate).
RESULTS: Overall traceability reached 89% (16/18 cited sources verifiable and accurate), rising to 100% for claims drawn from the uploaded reference guideline across both indications. Accuracy of browsed-source citations diverged sharply by indication: 57% for Crohn's disease (containing the only fabrication and both overstated-specificity cases) versus 100% for Ulcerative colitis, despite identical tool configuration and source types.
CONCLUSIONS: This analysis shows AI-generated PICO predictions can achieve high traceability when grounded in pre-specified source materials. However, inaccuracies and fabricated claims in externally sourced evidence show that source fidelity is not guaranteed by default and depends on tool configuration, highlighting the need for systematic manual source validation as part of AI-enabled JCA preparation workflows.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
HTA153
Topic
Health Technology Assessment, Methodological & Statistical Research, Study Approaches
Topic Subcategory
Decision & Deliberative Processes
Disease
No Additional Disease & Conditions/Specialized Treatment Areas