IDENTIFYING CHRONIC KIDNEY DISEASE STAGES USING PATIENT PRESCRIPTION INFORMATION

Author(s)

Cai Y*1;Han Y2;Jiao X3, Mu G1 1IMS Health, Plymouth Meeting, PA, USA, 2IPSEN Biopharmaceuticals, Inc., Basking Ridge, NJ, USA, 3IMS Health, Alexandria, VA, USA

OBJECTIVES: It’s important for healthcare policy makers, payers and drug manufacturers to identify Chronic Kidney Disease (CDK) patient stages in order to evaluate the prevalence, economic burden or market opportunities. The office based medical claims data, which contains ICD-9 diagnosis and treatment information of CDK stages, has very limited coverage. To identify more CDK patients, we built and compared varies of statistical models and machine learning algorithms to project CKD stages using prescription database. METHODS: The model data contained year 2011 patient level CKD stage indications, longitudinal drug therapies, days of supplies, titration rates, Demographic characteristics, payment type, and physician specialties, etc. The data were randomly divided into a training set (66.7%) and a validation set (33.3%). The classification models used were Logistic regression, linear/quadratic Discriminant, Classification and Regression Tree (CART), C4.5 decision tree, Logit Boost classification tree, Bayes learning networks, Support Vector Machine and Neural networks. Bagging and boosting techniques were also tested to improve the precision.  RESULTS: Logistic regression showed the best classification accuracy out of all models. The overall correct classification rate was 66.8%, Kappa Statistics 0.35, F-measure 0.63 and receiver operating characteristic (ROC) area 0.75. Among all CDK stages, CDK stage 1-3 were predicted most accurately with True Positive Rate (TPR) 93% and 65% precision. ESRD was moderately identified with TPR 46% and precision 76%. CDK 4 was most difficult to identify, with TPR 9% and precision 40%. The bagging and boosting improved decision trees showed comparable results.   CONCLUSIONS: The modern machine learning algorithms were proved to be more accurate in many cases but for the CKD dataset the classical logistic model worked out better. The logistic model could fairly classify severe CKD stages (4-5) patients from lower stages (1-3) using prescription information but it had hard time to separate CKD stage 4 from 5.

Conference/Value in Health Info

2013-05, ISPOR 2013, New Orleans, LA, USA

Value in Health, Vol. 16, No. 3 (May 2013)

Code

PRM71

Topic

Real World Data & Information Systems

Topic Subcategory

Reproducibility & Replicability

Disease

Urinary/Kidney Disorders

Explore Related HEOR by Topic


Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×