MACHINE LEARNING PREDICTION OF ASTHMA NON-SCHEDULED VISITS IN THE BRAZILIAN PRIVATE HEALTHCARE SETTING
Author(s)
Silveira A1, dos Santos F1, Alves DSB2, Zilocchi R3, Soares CS1
1GlaxoSmithKline, Rio de Janeiro, Brazil, 2Universidade Federal do Estado do Rio de Janeiro – UNIRIO, Rio de Janeiro, Brazil, 3Orizon, São Paulo, Brazil
Presentation Documents
OBJECTIVES: Asthma related non-scheduled visits are non-easily measured in burden studies using claims databases. Machine learning can provide an alternative method for predicting those events and improve healthcare burden assessment. Our objective is to find the best classifying algorithm and compare how it performs against manual classification. METHODS: Asthma patients (ICD-10 J45) aged ≥12 years old were identified from Jan 2010 to June 2016 in Orizon database (medical invoice claims Brazilian dataset). SADT data, which registers any type of medical care that was not a hospitalization, was used. The “Type of Care” field where healthcare provider informs if the SADT was an emergency was considered gold-standard. For specialist’s classification, eligible SADT entries were analyzed by two epidemiologists and one physician through semi-manual identification of emergency procedures in the expense description. Asthma-related events were defined as pre-scheduled physician visits and emergency visits (EV), with the following primary ICD-10 codes: J45, J46, J11, J12.8, J12.9 or J13-18. Additional codes (J20, J96.0, R05 and R06) were included when associated with systemic corticosteroid use. Both labels were compared to different classifiers machine learning algorithms: Naive Bayes, Decision Tree, Random Forest, SVM (linear, polynomial, gaussian) and Neural Network. Performance was measured using accuracy, confusion matrices and ROC curves. RESULTS: Specialists algorithm had an accuracy of 91.2% (true-positive: 39.9% and true-negative: 51.3%). All classifiers had less accuracy than specialist’s algorithm, where the highest were observed for Decision Tree (89.8%), followed by Neural Network (89.5%) and Random Forest (89.2%), and the lowest for Polynomial SVM (58.4%), followed by Naïve Bayes (70.3%) and Linear SVM (76.2%). Accuracy for Gaussian SVM was 83.0%. CONCLUSIONS: Although specialists algorithm had the highest accuracy, it demanded longer time for completion (±6 weeks). Use of machine learning classifiers can support physicians and payers to predict higher burden patients. Funding: GSK (212382).
Conference/Value in Health Info
2019-11, ISPOR Europe 2019, Copenhagen, Denmark
Code
PRS56
Topic
Epidemiology & Public Health, Methodological & Statistical Research, Real World Data & Information Systems
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics, Disease Classification & Coding, Health & Insurance Records Systems
Disease
Respiratory-Related Disorders