USING TEXT NOTES FROM CALL CENTER DATA TO PREDICT HOSPITALIZATION
Author(s)
Lorenzana A, Tyagi M, Wang QC, Chawla R, Nigam S
Independence Blue Cross, Philadelphia, PA, USA
Presentation Documents
OBJECTIVES: With the increasing amount of data available for health care analytics, new models and advances in data science have become critical to extracting value from the non-numerical data sources in machine learning algorithms. These models can use unstructured text data to extract information and glean insights that can be used to reduce the cost of health care. This study aims to identify hospitalization trends found in textual data. METHODS: Text mining is a sensitive process with varying results based on the text preprocessing. The data compose two parts: 2.6 million text-based call center comments transcribed by call center agents with members occurring between October 2014 and March 2015 and hospitalization occurring between April 2015 and September 2015. Text comments are often fewer than 300 characters in length and typically contain various abbreviations and spelling errors. We dismiss any spelling errors by removing tokens that occur fewer than 1000 times in the collection of 2.6 million comments. Tokenization consists of separating a string of text into numerous fields by removing whitespace and punctuation. In order to retain contextual information we duplicate the number of tokens by creating new tokens out of sequences of tokens. These tokens and hospitalization outcomes then undergo a 60/40 split, are selected using a Chi-squared test and passed to a Naïve Bayes Classifier to predict the hospitalization outcomes. RESULTS: The Naïve Bayes classifier can predict hospitalization outcomes with up to 82% accuracy, with .65 area under the ROC curve. Further, the Chi-squared test can be used to find words associated with hospitalization, which is useful for making qualitative insights into the population of members. CONCLUSIONS: This study showed that text features yield both explanatory and predictive information about hospitalization. Text data can aid in understanding members that are later hospitalized and be used to make predictions.
Conference/Value in Health Info
2016-05, ISPOR 2016, Washington DC, USA
Value in Health, Vol. 19, No. 3 (May 2016)
Code
PRM89
Topic
Methodological & Statistical Research
Topic Subcategory
Modeling and simulation
Disease
Multiple Diseases