Beyond Accuracy: Fairness Implications in High-Cost User Prediction Models Among Incident Lung Cancer Survivors
Author(s)
Pathak M1, Gupta M2, Zhou B3, Dehghan A4, Sambamoorthi N3, Sambamoorthi U1
1University of North Texas Health Sciences Center, Denton, TX, USA, 2Southern Methodist University, Dallas, TX, USA, 3University of North Texas Health Sciences Center, Fort Worth, TX, USA, 4University of North Texas Health Science Center, Denton, TX, USA
OBJECTIVES: Cost containment efforts often focus on high-cost and high-need patients. However, studies in predicting high-cost users are limited. This study identifies leading predictors of high-cost users among older adults with incident lung cancer and evaluate fairness of predicting high-cost in subgroups of gender, race and ethnicity, poverty, and metro status.
METHODS: We used SEER-Medicare database with diverse subgroups (60% women, 6.6% NHA, 4.4% Hispanic, and 5.4% AA/PI/AN) and adopted a retrospective cohort analysis of older adults (age > 66 years) with incident lung cancer diagnosed between 2010-2017. High-cost users were identified as having higher than 80th percentile($109,021.7) in Medicare payments. We used interpretable machine learning model (XGBoost with 5-fold cross-validation and SHAP). Group and counterfactual fairness were measured with demographic parity (Equalization of Odds Ratio(EOR), Disparate imPact Difference(DPD), and Equal OPportunity Difference(EOPD)). Ratios between 0.8 and 1.2 and differences between -0.1 and +0.1 were considered fair.
RESULTS: 18.5% women,18.8% NHW, 25.5% NHB, 24.4% Hispanic, 26.4% AA/PI/AN, 26% other race,12.9% non-metro and 23.6% with low income had high-cost. The model fit was excellent with AUC 0.90; recall 0.72; precision 0.77. AUC was higher among men, NHB, Hispanic, and AA/PI/AN compared to women and NHW. Multi-morbidity, cancer treatment, localized and distant cancer stage were the top 6 predictors of high-cost. EOR suggested unfairness of prediction among NHB(2.63), Hispanic(1.52), AA/PI/AN(1.49), and gender(1.79) and EOPD was higher than the threshold for Hispanics and AA/PI/AN. DPDs were not different among sensitive groups compared to their normative counterparts.
CONCLUSIONS: Accuracy of the overall model was very high. However, fairness in prediction varied by type of fairness metrics. Although the data consisted sensitive attributes, small number of patients in these groups may have led to inconsistent findings in fairness. Future research with large number of patients with sensitive attributes are needed to evaluate fairness in predicting high-cost users.
Conference/Value in Health Info
Value in Health, Volume 27, Issue 6, S1 (June 2024)
Acceptance Code
P11
Topic
Economic Evaluation, Health Policy & Regulatory, Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics, Health Disparities & Equity
Disease
Oncology