Machine Learning for Imputing Missing Pharmacy Costs in Claims Data

Author(s)

Vojjala SK1, Barron J2, Kumar A2, Grabner M2, Eshete B2, Tan H3, Willey V2
1Carelon Research, Middletown, DE, USA, 2Carelon Research, Wilmington, DE, USA, 3Carelon Research, San Diego, CA, USA

Presentation Documents

OBJECTIVES: Cost data may be missing from administrative claims data for several reasons, and multiple methods of cost imputation have been used over time. Machine learning techniques are being utilized with greater regularity across all of healthcare, and their use in cost imputation may allow for broader utility of previously incomplete data. To that end, we developed and compared several machine learning algorithms to impute missing pharmacy claims costs in a large US commercially insured population.

METHODS: Pharmacy claims with non-missing cost data from 1/1/2013 to 12/31/2021 within the HealthCore Integrated Research Database were randomly divided into training (60%), validation (20%), and test (20%) datasets. We considered various predictors of allowed pharmacy costs including national drug code (NDC), quantity dispensed, pharmacy state, other medication properties (e.g., specialty, wholesale prices), health plan characteristics, and patient demographics. Linear regression, categorical boosting, light gradient boosting, and XGBoost algorithms were developed and evaluated with training and validation datasets. The test dataset was used to assess algorithm performance using root mean square error (RMSE) residual analysis.

RESULTS: A total of approximately 120 million claims and 43,000 NDCs were analyzed. The RMSE were 0.71, 0.43, 0.41, 0.31 for linear regression, categorical boosting, light gradient boosting, and XGBoost, respectively. The XGBoost algorithm had the best fit. The top 5 predictors were NDC, quantity dispensed, wholesale unit price, account type (local/national), and insurance plan type. The mean cost differences between the imputed and actual costs were less than $10 across over 90% of low-cost (<$100) NDCs and less than 25% across over 90% of medium- (≥$100) and high-cost (≥$1,000) NDCs.

CONCLUSIONS: Use of machine learning techniques provided acceptable estimates for missing pharmacy costs using real-world data. This approach will allow for use of data that previously could not be utilized for healthcare cost analyses.

Conference/Value in Health Info

2023-05, ISPOR 2023, Boston, MA, USA

Value in Health, Volume 26, Issue 6, S2 (June 2023)

Acceptance Code

P21

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics, Missing Data

Disease

biologics-biosimilars, Drugs, Generics, no-additional-disease-conditions-specialized-treatment-areas

Explore Related HEOR by Topic


Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×