IDENTIFYING DIAGNOSES FOR PRESCRIBED MEDICATION IN JAPAN HEALTH INSURANCE CLAIM DATA- AN APPLICATION OF NATURAL LANGUAGE PROCESSING TECHNIQUE ON REAL WORLD HEALTHCARE DATABASE
Author(s)
Chu C
IQVIA JAPAN, Tokyo, Japan
Presentation Documents
OBJECTIVES: Healthcare claims database provides rich information for treatment evaluation and forecasting. A blindside remains as multiple diagnoses often linked to multiple medications on a visit. Pinpointing diagnoses for a medication could unveil treatment usage of a drug. Conversely, knowing the prescribed medications for a disease could support clinical decisions. This was a pilot study aimed to identify prescribing purpose of medication using latent Dirichlet allocation (LDA), topic modeling from natural language processing, which often used by search engines to sort articles into topics. Topic modeling allows us to identify latent topics, find synonym and polysemy, and data reduction. METHODS: Data from July 2015 was extracted from IQVIA Claim database. Diagnoses and medications were linked by claim IDs. The corpus of medications contained diagnoses gathered from patient claims. 1/3 of medications were used for training. To determine the optimum hyperparameters, perplexities were calculated for a series of topic numbers. A ten-fold cross-validation (CV) was performed to check overfitting. Associated diagnostic topics were checked against usage guide. An illustration of proton pump inhibitor (PPI) drugs was used to demonstrate prescribing trends and common comorbidities. RESULTS: The training data contained 4,000 medications. The optimal topic number with the lowest perplexity (217.7) was 200. The ten-fold CV showed little difference between training and hold-out sets (perplexities 198.9 vs 217.6). Sampled medications had diagnoses conformed to the usage guide. PPI drugs were used differently, where two PPI products were prescribed to hypertension for prevention of NSAIDS-associated ulcer while another two were prescribed when esophageal/stomach ulcer occurred. CONCLUSIONS: Topic modeling correctly identified diseases for prescribed medications while excluding unapproved uses of health conditions. Natural language processing enabled us to reduce 20,000 of diagnostic codes into hundreds of meaningful categories, identify comorbidities, and observe medication usage. Mining on large scale healthcare database helped us detect seasonal and regional prescribing trends.
Conference/Value in Health Info
2018-09, ISPOR Asia Pacific 2018, Tokyo, Japan
Value in Health, Vol. 21, S2 (September 2018)
Code
PRM27
Topic
Methodological & Statistical Research
Topic Subcategory
Modeling and simulation
Disease
Multiple Diseases