PREDICTING HYPOPARATHYROIDISM DIAGNOSIS IN MEDICARE CLAIMS - EXPLORING THE UTILITY OF SEQUENCE ANALYSIS
Author(s)
Li S1, Yu B1, Wang X1, Yu G1, Singh D1, Miyasato G2, Yajima M1
1Boston University, Boston, MA, USA, 2Trinity Partners, LLC, Waltham, MA, USA
OBJECTIVES : Claims-based prediction models commonly use dichotomous indicator variables (e.g., presence/absence of disease) or count variables (e.g., frequency of claims) as predictors. The research objective was to compare the predictive ability of this standard model with other models that account for the sequence of events. The models were applied in the context of predicting hypoparathyroidism diagnosis among Medicare beneficiaries. METHODS : 24-month claim histories for patients with and without hypoparathyroidism were sourced from a 5% random sample of Medicare fee-for-service beneficiaries (2010-2016). We created a baseline model using a random forest approach, including only demographics and frequency variables based on Clinical Classification Software (CCS) groupings of ICD codes. Comparator models (hidden Markov model [HMM], random forest and logistic regression) were constructed using claim sequence data. To reduce the dimensionality of the sequence data, a sequential pattern mining method (PrefixSpan) was used to identify the candidate subsequences that most differentiated those with hypoparathyroidism. We assessed model performance by comparing prediction accuracy, sensitivity and specificity. RESULTS : 3,111 and 8,572 patients with and without hypoparathyroidism, respectively, were analyzed. The baseline model included 30 CCS groupings (determined via feature selection) and resulted in an accuracy, specificity and sensitivity of 0.82, 0.96 and 0.44, respectively. The sequence-based random forest and logistic regression models had similar accuracy and specificity (0.80 and 0.94, 0.81 and 0.93, respectively) but lower sensitivity (0.4 and 0.47, respectively). The HMM model had higher sensitivity (0.54) but lower accuracy and specificity (0.61 and 0.67, respectively). CONCLUSIONS : The inclusion of sequence information did not improve our predictions, possibly because the sequences were very repetitive or, in some cases, very short or the CCS groupings did not provide enough granularity to distinguish sequences. Our research highlights the challenges involved with using sequence data but we believe refinement of our approaches can lead to better predictions.
Conference/Value in Health Info
2019-05, ISPOR 2019, New Orleans, LA, USA
Value in Health, Volume 22, Issue S1 (2019 May)
Code
PDB100
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
Diabetes/Endocrine/Metabolic Disorders