COMPARISON OF MACHINE LEARNING, STATISTICAL AND HYBRID METHODS TO IDENTIFY PREDICTORS OF POSITIVE TREATMENT OUTCOMES IN COMORBID CONDITIONS USING EMR DATA
Author(s)
Lipkovich I1, Griner BP2, Niemira J3, Jin C2
1Quintiles, Morrisville, MA, USA, 2Quintiles, New York, NY, USA, 3Quintiles, Cambridge, MA, USA
OBJECTIVES: Electronic Healthcare databases and use of machine learning algorithms has created opportunities for rapid learning. However, the indiscriminate application of machine learning algorithms to non-experimental healthcare databases may result in incorrect inferences about possible treatment benefits. This paper compares results from large healthcare databases produced by statistical, machine learning, and hybrid methods to emphasize the importance of controlling for known biases in healthcare databases and machine learning algorithms. METHODS: MS patient cases were selected from an EMR database that met the following criteria: Must be on an MS therapy for at least one year and diagnosed with at least one co morbidity prior to and during the treatment period. Co morbidity improvements are measured by changes in specific lab values measured prior to and during treatment. To facilitate method comparison, a binary variable was constructed to measure improvements in co morbidities experienced during MS therapy. Treatment groups were defined by specific MS therapies and compared to control groups treated with alternative therapies during the observation period. Propensity scores were used with all methods. Statistical and machine learning algorithms were compared to a hybrid algorithm, SIDES, originally designed for subpopulation analysis in RCT’s while controlling for multiplicity bias (Lipkovich et al. 2011). RESULTS: Initial analyses identified differences in predictors of co mobidity improvements. The presentation will cover specific comparisons between different methods, highlight similarities and differences in findings and provide rationales for divergent results. CONCLUSION: Machine learning methods (such as SIDES) designed for use in RCTs can be adapted for use with large healthcare databases to accelerate learning and discovery while also including protections against known sources of bias in healthcare data (treatment selection) and machine learning methods (multiplicity) that can lead to incorrect inferences.
Conference/Value in Health Info
2015-05, ISPOR 2015, Philadelphia, PA, USA
Value in Health, Vol. 18, No. 3 (May 2015)
Code
PHP183
Topic
Health Policy & Regulatory
Disease
Multiple Diseases