INTEGRATING REAL-WORLD PHARMACOGENOMIC EVIDENCE WITH MACHINE LEARNING FOR DRUG RESPONSE PREDICTION: IMPLICATIONS FOR EQUITABLE HEALTH TECHNOLOGY ASSESSMENT
Author(s)
Jashuva T, PharmD1, Manoj Kumar Mudigubba, MPH, PharmD, PhD2.
1Department of Pharmacy Practice, Raghavendra Institute of Pharmaceutical Education & Research (RIPER), Anantapur, India, 2Department of Pharmacy Practice, Raghavendra Institute of Pharmaceutical Education and research, Anantapur, India.
1Department of Pharmacy Practice, Raghavendra Institute of Pharmaceutical Education & Research (RIPER), Anantapur, India, 2Department of Pharmacy Practice, Raghavendra Institute of Pharmaceutical Education and research, Anantapur, India.
OBJECTIVES: Pharmacogenomic clinical decision support tools predominantly rely on curated variant databases, limiting generalizability across diverse populations and introducing uncertainty into health technology assessment (HTA). This study evaluated the impact of integrating real-world genotype-phenotype data with curated PGx evidence on machine learning (ML)-based drug response prediction, characterized ancestry-stratified performance disparities, and assessed its potential to strengthen decision-relevant evidence for AI-enabled HTA.
METHODS: Real-world genotype-phenotype data from the eMERGE-PGx cohort were integrated with PharmGKB Level 1A/1B variant-drug associations and ClinVar pharmacogene annotations across five CPIC guideline-supported gene-drug pairs: CYP2C19-clopidogrel, CYP2C9/VKORC1-warfarin, SLCO1B1-statins, TPMT-thiopurines, and DPYD-fluoropyrimidines (aligned with EMA recommendations for pre-treatment testing). Drug-response phenotypes were defined using standardized eMERGE criteria. XGBoost, random forest, and logistic regression models were developed under two conditions using ancestry-stratified 10-fold cross-validation: (1) RWE-augmented and (2) curated-only. Primary outcomes were discrimination (AUC), calibration, and ancestry-stratified subgroup performance. Exploratory scenario analysis examined potential implications for prescribing decisions and HTA evidence quality.
RESULTS: RWE-augmented XGBoost achieved the highest discrimination (AUC 0.71-0.79), a 7-10% improvement over curated-only models. Feature importance remained biologically concordant with established CPIC relationships, supporting model interpretability. Despite improved overall performance, models trained predominantly in European ancestry populations exhibited a 6-11% reduction in discrimination among South Asian and African ancestry subgroups - a systematic disparity with direct implications for equitable HTA generalizability. Exploratory scenario analysis suggested improved discrimination may support prescribing decisions with potential downstream implications for healthcare utilization.
CONCLUSIONS: Integrating real-world evidence with curated pharmacogenomic knowledge improves ML-based drug response prediction while maintaining biological interpretability, but exposes critical ancestry-related evidence gaps that HTA bodies must address to ensure equitable implementation. These findings provide decision-relevant evidence to inform future evaluation standards for AI-enabled pharmacogenomic technologies within HTA.
METHODS: Real-world genotype-phenotype data from the eMERGE-PGx cohort were integrated with PharmGKB Level 1A/1B variant-drug associations and ClinVar pharmacogene annotations across five CPIC guideline-supported gene-drug pairs: CYP2C19-clopidogrel, CYP2C9/VKORC1-warfarin, SLCO1B1-statins, TPMT-thiopurines, and DPYD-fluoropyrimidines (aligned with EMA recommendations for pre-treatment testing). Drug-response phenotypes were defined using standardized eMERGE criteria. XGBoost, random forest, and logistic regression models were developed under two conditions using ancestry-stratified 10-fold cross-validation: (1) RWE-augmented and (2) curated-only. Primary outcomes were discrimination (AUC), calibration, and ancestry-stratified subgroup performance. Exploratory scenario analysis examined potential implications for prescribing decisions and HTA evidence quality.
RESULTS: RWE-augmented XGBoost achieved the highest discrimination (AUC 0.71-0.79), a 7-10% improvement over curated-only models. Feature importance remained biologically concordant with established CPIC relationships, supporting model interpretability. Despite improved overall performance, models trained predominantly in European ancestry populations exhibited a 6-11% reduction in discrimination among South Asian and African ancestry subgroups - a systematic disparity with direct implications for equitable HTA generalizability. Exploratory scenario analysis suggested improved discrimination may support prescribing decisions with potential downstream implications for healthcare utilization.
CONCLUSIONS: Integrating real-world evidence with curated pharmacogenomic knowledge improves ML-based drug response prediction while maintaining biological interpretability, but exposes critical ancestry-related evidence gaps that HTA bodies must address to ensure equitable implementation. These findings provide decision-relevant evidence to inform future evaluation standards for AI-enabled pharmacogenomic technologies within HTA.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
RWD178
Topic
Health Technology Assessment, Methodological & Statistical Research, Real World Data & Information Systems
Disease
Cardiovascular Disorders (including MI, Stroke, Circulatory), Genetic, Regenerative & Curative Therapies, Oncology, Personalized & Precision Medicine