Machine Learning for Subpopulation Analysis in Datasets without Control Arm
Author(s)
Wang Y1, Wei G1, Wang Y1, Behnke M2, Reiner E3, Chaudhuri K4, Reeve R5, McKemey A6, Oliva C7, Reynolds M5
1IQVIA, Plymouth Meeting, PA, USA, 2IQVIA, Overland Park, KS, USA, 3IQVIA, Pound Ridge, NY, USA, 4IQVIA, Ambler, PA, USA, 5IQVIA, Durham, NC, USA, 6IQVIA, New York, NY, USA, 7IQVIA, Reading, UK
OBJECTIVES : Develop machine learning (ML) and predictive analysis models to identify patient attributes and biomarkers predictive of a particular outcome – desirable or undesirable. This can be used in one arm clinical trials and real world data towards various outcomes including disease-free survival and adverse events. METHODS : We used decision trees to rank the top biomarkers splitting criteria and validated them using hypothesis tests. We used these to develop ML models predicting the mortality of sepsis patients. 10-fold cross-validation was performed using the best parameters from the grid search to predict the study outcome. The finalized decision tree has at least 30 samples on each node and returns the biomarkers splitting criterion of each node until maximum depth reaches five. We used Receiver Operating Characteristic (ROC) to assess model performance. permutation tests were applied for splitting criterion validation plus survival status. We then applied scanning algorithms to search all values of each variable. After data splitting found each threshold, log-rank tests of Kaplan-Meier survival curves and likelihood ratio tests from Cox proportional hazard models were applied, testing for survival variances. RESULTS : The decision tree method returned three sub-groups with differentiated outcomes, and the information gain increased by an average of 0.2 for each one. Assessed cutoffs based on the splitting criteria from the decision tree, Kaplan-Meier and Cox Regression models showed statistical significance of the outcome of the optimal sub-group vs. the rest. CONCLUSIONS : ML models and survival analysis used to extract optimal sub-groups enhances identification of the differential outcome between sub-groups in clinical trial data without a control arm. We envision that these algorithms have the potential to do subpopulation analysis of clinical trial data as well as real world data when having a control arm is not practical.
Conference/Value in Health Info
2020-11, ISPOR Europe 2020, Milan, Italy
Value in Health, Volume 23, Issue S2 (December 2020)
Code
PNS203
Topic
Clinical Outcomes, Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics, Clinical Outcomes Assessment
Disease
No Specific Disease