PREDICTING NASH PATIENTS USING INNOVATIVE MACHINE LEARNING TECHNIQUES

Author(s)

Docherty M1, Huang J1, Regnier SA2, Capkun G2, Balp MM2, Ye Q1, Janssens N2, Lopez P2, Pedrosa M2, Schattenberg JM3
1ZS, Princeton, NJ, USA, 2Novartis Pharma AG, Basel, Switzerland, 3University Medical Center Mainz, Mainz, Germany

OBJECTIVES : Non-alcoholic steatohepatitis (NASH) is an inflammatory form of non-alcoholic fatty liver disease (NAFLD), which can progress to cirrhosis and hepatocellular carcinoma. The confirmatory diagnosis requires invasive liver biopsy and NASH is commonly underdiagnosed in clinical practice. The objective of this study was to develop and validate a machine learning model for early prediction of NASH patients using non-invasive clinical parameters available in electronic medical records (EMR).

METHODS : A data-driven approach using machine learning techniques and two real-world databases were used to develop predictive models for identification of NASH patients. First, exploratory analysis, feature extraction, model training, and parameter tuning were conducted on the NAFLD Adult Database from the National Institute of Diabetes, Digestive and Kidney Diseases (NIDDK). This dataset has ~450 confirmed NASH and ~250 confirmed non-NASH, NAFLD patients. The best-performing model from NIDDK was tested on Optum Humedica EMR database using a cohort of 1,016 patients with NASH confirmed by liver biopsy. The model performance was evaluated and selected based on the area under the curve (AUC). Additional performance measures such as sensitivity, specificity, and overall accuracy were also analyzed to understand model performance.

RESULTS : A gradient boosting model (XGBoost) was the best performing model (AUC: 0.82 in NIDDK, 0.76 in Optum EMR). This model retains 14 variables after recursive feature elimination. A smaller model with 5 variables (in order: HbA1c, AST, ALT, triglycerides, total protein) trades off slightly lower performance (AUC: 0.80 in NIDDK, 0.74 in Optum EMR) for reduced input data requirements.

CONCLUSIONS : This model may be utilized within existing EMRs as an effective and scalable pre-screening support tool for referring patients at risk of NASH to specialists. Further validation in other real-world databases is planned to evaluate the value of the tool in clinical practice and for clinical trial recruitment.

Conference/Value in Health Info

2019-11, ISPOR Europe 2019, Copenhagen, Denmark

Code

PDB125

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

Diabetes/Endocrine/Metabolic Disorders

Explore Related HEOR by Topic


Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×