FROM PICO TO DOSSIER: EVALUATING A SOURCE-GROUNDED EVIDENCE-VALIDATION LAYER FOR EVIDENCE TRACEABILITY AND CONTRADICTION DETECTION IN HTA-READY CLAIMS
Author(s)
Ankit Dhaundiyal, MS1, Jagriti Prasad, MPH2, Tejesh S, MPH2, Mona Thangamma, MSc Health Economics2, Megha Tharad, PhD3.
1Evalueserve GmbH, Rheinbach, Germany, 2Evalueserve Pvt Ltd, Bengaluru, India, 3Evalueserve SEZ Pvt Ltd, Gurugram, India.
1Evalueserve GmbH, Rheinbach, Germany, 2Evalueserve Pvt Ltd, Bengaluru, India, 3Evalueserve SEZ Pvt Ltd, Gurugram, India.
OBJECTIVES: To evaluate whether a source-grounded AI evidence-validation workflow could classify HTA-ready claim-source pairs, assess source alignment, and detect unsupported or contradictory claims within a PICO-to-dossier traceability framework.
METHODS: An exploratory controlled validation was conducted across 120 curated claim-source pairs from 10 oncology product-indication cases, using public clinical and regulatory evidence. Claims covered population, intervention/comparator, trial design, efficacy, safety, HRQoL/PRO, subgroup, and uncertainty/economic inputs. Each pair included a dossier-style claim and corresponding source passage. Reference labels - fully supported, partially supported, unsupported, or contradicted - were assigned before AI classification and withheld during clean evaluation runs. The dataset was balanced across four classes, with 30 claims per class, to support per-class evaluation. Task-specialised modules parsed claims, aligned them with source passages, classified support status, detected contradictions, and generated rationales using only supplied evidence. The primary endpoint was four-class classification accuracy. Secondary endpoints included problematic-claim sensitivity/precision, supported-claim specificity, contradiction-flag sensitivity/specificity/precision, and macro F1.
RESULTS: The workflow correctly classified 107/120 claim-source pairs, giving 89.2% accuracy (Wilson 95% CI: 82.3%-93.6%). Problematic-claim detection achieved 100.0% sensitivity (90/90), 100.0% supported-claim specificity (30/30), and 100.0% precision (90/90). Contradiction detection identified all contradicted claims (30/30; 100.0% sensitivity), with 86.7% specificity (78/90) and 71.4% precision (30/42). Macro F1 was 89.1%. Most errors reflected conservative over-flagging, where unsupported or partially supported claims were escalated as contradicted.
CONCLUSIONS: These findings suggest proof-of-concept value for a source-grounded evidence-validation layer to support HTA claim substantiation and dossier quality control. Its practical role is to help expert teams prioritise problematic claims for review, rather than replace reviewer judgement. Interpretation should remain cautious because the dataset was controlled, balanced, oncology-focused, and based on public evidence, with no manual dual-review or baseline NLP comparator. Larger prospective studies in real-world dossier workflows, including non-oncology settings and workflow-efficiency endpoints, are warranted before dossier-critical use.
METHODS: An exploratory controlled validation was conducted across 120 curated claim-source pairs from 10 oncology product-indication cases, using public clinical and regulatory evidence. Claims covered population, intervention/comparator, trial design, efficacy, safety, HRQoL/PRO, subgroup, and uncertainty/economic inputs. Each pair included a dossier-style claim and corresponding source passage. Reference labels - fully supported, partially supported, unsupported, or contradicted - were assigned before AI classification and withheld during clean evaluation runs. The dataset was balanced across four classes, with 30 claims per class, to support per-class evaluation. Task-specialised modules parsed claims, aligned them with source passages, classified support status, detected contradictions, and generated rationales using only supplied evidence. The primary endpoint was four-class classification accuracy. Secondary endpoints included problematic-claim sensitivity/precision, supported-claim specificity, contradiction-flag sensitivity/specificity/precision, and macro F1.
RESULTS: The workflow correctly classified 107/120 claim-source pairs, giving 89.2% accuracy (Wilson 95% CI: 82.3%-93.6%). Problematic-claim detection achieved 100.0% sensitivity (90/90), 100.0% supported-claim specificity (30/30), and 100.0% precision (90/90). Contradiction detection identified all contradicted claims (30/30; 100.0% sensitivity), with 86.7% specificity (78/90) and 71.4% precision (30/42). Macro F1 was 89.1%. Most errors reflected conservative over-flagging, where unsupported or partially supported claims were escalated as contradicted.
CONCLUSIONS: These findings suggest proof-of-concept value for a source-grounded evidence-validation layer to support HTA claim substantiation and dossier quality control. Its practical role is to help expert teams prioritise problematic claims for review, rather than replace reviewer judgement. Interpretation should remain cautious because the dataset was controlled, balanced, oncology-focused, and based on public evidence, with no manual dual-review or baseline NLP comparator. Larger prospective studies in real-world dossier workflows, including non-oncology settings and workflow-efficiency endpoints, are warranted before dossier-critical use.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR275
Topic
Health Technology Assessment, Methodological & Statistical Research, Real World Data & Information Systems
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas