USING MACHINE LEARNING TECHNIQUES TO CLASSIFY OECD COUNTRIES ACCORDING TO HEALTH EXPENDITURES
Author(s)
Cinaroglu S
Hacettepe University, Ankara, Turkey
OBJECTIVES: Machine learning techniques are used for analysis of large complex datasets. Classification is an important part of machine learning applications, it defines groups within population. There are many different methods which are compare results to determine the best classification. In this study we aim to use machine learning techniques to classify OECD countries according to their health expenditures. METHODS: Different algorithms can be use in machine learning techniques; C4.5 which is an extension version of ID.3 algorithm and CART algorithm are one of these most commonly use algorithms. Random Forest which constructs a lot of number of trees is one of another useful technique for solving both classification and regression problems. In this study we compare classification performances of different decision trees (C4.5, CART) and Random Forest which was generated by using 50 trees. We perform this prediction model for predicting OECD countries health expenditures for the year 2011. We use number of independent variables for this prediction. These are; life expectancy at birth, number of physicians, number of hospitals, hospital aggregates, alcohol consumption, GDP per capita, perceived health status and immunization. We use AUC results and ROC curve graph for performance comparison. RESULTS: As a result of this study it was seen that classification performances of machine learning techniques were good (AUC≥0.90) and Random Forest [50] classification performance results much higher [AUC=0.98] than CART (0.95) and C4.5 (0.90). Decision tree graphs shows that GDP per capita was a variable which has more information gain for predicting health expenditures. CONCLUSIONS: To conclude according to our knowledge this is the first study applied machine learning classification methods to health expenditure data. Future studies will compare classification performances of Random Forest using different types of health expenditure datasets, different predictor variables while increasing the number of trees in the forest.
Conference/Value in Health Info
2015-11, ISPOR Europe 2015, Milan, Italy
Value in Health, Vol. 18, No. 7 (November 2015)
Code
PRM216
Topic
Methodological & Statistical Research
Topic Subcategory
Confounding, Selection Bias Correction, Causal Inference
Disease
Multiple Diseases