GOVERNANCE-FIRST MACHINE LEARNING FOR EXPLAINABLE HTA EVIDENCE-RISK INTELLIGENCE IN ONCOLOGY: A PILOT STUDY IN PEMBROLIZUMAB FOR NSCLC
Author(s)
Athanasios Chalikias, PhD1, Charalampos Gourzis, MSc2.
1Likovrisi, Greece, 2Aristarchus AI, Athens, Greece.
1Likovrisi, Greece, 2Aristarchus AI, Athens, Greece.
OBJECTIVES: To evaluate MA-OS, a governance-first machine-learning framework, for transforming oncology health technology assessment (HTA) evidence into explainable, auditable evidence-risk intelligence under uncertainty. The objective was not reimbursement prediction, but reproduction of HTA-style methodological reasoning and identification of restriction-risk drivers.
METHODS: HTA reports, regulatory artifacts, and clinical evidence for pembrolizumab in NSCLC were transformed into structured evidence objects using provenance-linked extraction, normalization, and source-validation controls. AI/ML components included report-native feature engineering, TF-IDF evidence representation, supervised multiclass logistic-regression classification, isotonic calibration, probability-distribution diagnostics, and driver attribution. Features captured PICO alignment, comparator validity, endpoint maturity, subgroup robustness, restriction language, economic uncertainty, and cross-agency divergence. Outputs were governed through Run ID, Snapshot ID, Estimand Registry, claim-boundary metadata, citation anchors, and human-review controls. Robustness analyses included calibration assessment, no-leakage stress testing, source-integrity checks, directional concordance assessment, and driver-stability evaluation.
RESULTS: The internal case-study set included 98 labelled decisions across five HTA agencies (AEMPS, AIFA, G-BA, HAS, NICE), with complete audit coverage across evidence objects, citation anchors, claim boundaries, Run IDs, and Snapshot IDs. In the validation set (73 training, 25 holdout), accuracy was 0.92 and balanced accuracy 0.67, with calibrated probability estimates (Brier score 0.138, ECE 0.109). A separate no-leakage evidence-only stress test showed reduced performance (accuracy 0.523; balanced accuracy 0.346), supporting absence of shortcut learning. Driver analysis identified comparator validity, PD-L1 subgroup definition, evidence maturity, line-of-therapy positioning, and economic uncertainty as dominant restriction-risk contributors. Aggregate retrospective directional concordance across historical HTA decisions was retained in the robustness package. The analysis also preserved the distinction between evidence-structure concordance and autonomous prediction, strengthening interpretability for HTA, HEOR, and market-access review audiences.
CONCLUSIONS: MA-OS demonstrates a governance-first AI approach for converting HTA evidence into structured evidence-risk intelligence, emphasizing reproducibility, uncertainty awareness, leakage control, and early identification of restriction drivers before submission for transparent, reproducible, human-governed pre-submission decision-support workflows across agencies.
METHODS: HTA reports, regulatory artifacts, and clinical evidence for pembrolizumab in NSCLC were transformed into structured evidence objects using provenance-linked extraction, normalization, and source-validation controls. AI/ML components included report-native feature engineering, TF-IDF evidence representation, supervised multiclass logistic-regression classification, isotonic calibration, probability-distribution diagnostics, and driver attribution. Features captured PICO alignment, comparator validity, endpoint maturity, subgroup robustness, restriction language, economic uncertainty, and cross-agency divergence. Outputs were governed through Run ID, Snapshot ID, Estimand Registry, claim-boundary metadata, citation anchors, and human-review controls. Robustness analyses included calibration assessment, no-leakage stress testing, source-integrity checks, directional concordance assessment, and driver-stability evaluation.
RESULTS: The internal case-study set included 98 labelled decisions across five HTA agencies (AEMPS, AIFA, G-BA, HAS, NICE), with complete audit coverage across evidence objects, citation anchors, claim boundaries, Run IDs, and Snapshot IDs. In the validation set (73 training, 25 holdout), accuracy was 0.92 and balanced accuracy 0.67, with calibrated probability estimates (Brier score 0.138, ECE 0.109). A separate no-leakage evidence-only stress test showed reduced performance (accuracy 0.523; balanced accuracy 0.346), supporting absence of shortcut learning. Driver analysis identified comparator validity, PD-L1 subgroup definition, evidence maturity, line-of-therapy positioning, and economic uncertainty as dominant restriction-risk contributors. Aggregate retrospective directional concordance across historical HTA decisions was retained in the robustness package. The analysis also preserved the distinction between evidence-structure concordance and autonomous prediction, strengthening interpretability for HTA, HEOR, and market-access review audiences.
CONCLUSIONS: MA-OS demonstrates a governance-first AI approach for converting HTA evidence into structured evidence-risk intelligence, emphasizing reproducibility, uncertainty awareness, leakage control, and early identification of restriction drivers before submission for transparent, reproducible, human-governed pre-submission decision-support workflows across agencies.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
HTA260
Topic
Health Technology Assessment, Methodological & Statistical Research, Real World Data & Information Systems
Topic Subcategory
Decision & Deliberative Processes, Value Frameworks & Dossier Format
Disease
Oncology