Machine Learning for Subpopulation Analysis in Datasets without Control Arm

Author(s)

Wang Y1, Wei G1, Wang Y1, Behnke M2, Reiner E3, Chaudhuri K4, Reeve R5, McKemey A6, Oliva C7, Reynolds M5
1IQVIA, Plymouth Meeting, PA, USA, 2IQVIA, Overland Park, KS, USA, 3IQVIA, Pound Ridge, NY, USA, 4IQVIA, Ambler, PA, USA, 5IQVIA, Durham, NC, USA, 6IQVIA, New York, NY, USA, 7IQVIA, Reading, UK

OBJECTIVES : Develop machine learning (ML) and predictive analysis models to identify patient attributes and biomarkers predictive of a particular outcome – desirable or undesirable. This can be used in one arm clinical trials and real world data towards various outcomes including disease-free survival and adverse events.

METHODS : We used decision trees to rank the top biomarkers splitting criteria and validated them using hypothesis tests. We used these to develop ML models predicting the mortality of sepsis patients. 10-fold cross-validation was performed using the best parameters from the grid search to predict the study outcome. The finalized decision tree has at least 30 samples on each node and returns the biomarkers splitting criterion of each node until maximum depth reaches five. We used Receiver Operating Characteristic (ROC) to assess model performance. permutation tests were applied for splitting criterion validation plus survival status. We then applied scanning algorithms to search all values of each variable. After data splitting found each threshold, log-rank tests of Kaplan-Meier survival curves and likelihood ratio tests from Cox proportional hazard models were applied, testing for survival variances.

RESULTS : The decision tree method returned three sub-groups with differentiated outcomes, and the information gain increased by an average of 0.2 for each one. Assessed cutoffs based on the splitting criteria from the decision tree, Kaplan-Meier and Cox Regression models showed statistical significance of the outcome of the optimal sub-group vs. the rest.

CONCLUSIONS : ML models and survival analysis used to extract optimal sub-groups enhances identification of the differential outcome between sub-groups in clinical trial data without a control arm. We envision that these algorithms have the potential to do subpopulation analysis of clinical trial data as well as real world data when having a control arm is not practical.

Conference/Value in Health Info

2020-11, ISPOR Europe 2020, Milan, Italy

Value in Health, Volume 23, Issue S2 (December 2020)

Code

PNS203

Topic

Clinical Outcomes, Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics, Clinical Outcomes Assessment

Disease

No Specific Disease

Explore Related HEOR by Topic


Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×