LARGE LANGUAGE MODEL-EXTRACTED CLINICAL FEATURES ENHANCED STRUCTURED ELECTRONIC MEDICAL RECORDS DATA FOR PREDICTING TREATMENT RESPONSE DEPTH IN MULTIPLE MYELOMA
Author(s)
Meenakshi Dubey, MS1, Lee Yi FOO, BSc1, Kok Joon CHONG, MSc1, Yuba Raj PUN, BSc1, Kee Yuan NGIAM, FRCS2, Allison Tso, FRC Path3, Hwee-Lin Wee, PhD4, Melissa Ooi, PhD2.
1Saw Swee Hock School of Public Health, National University of Singapore, Singapore, Singapore, 2National University Hospital Singapore, Singapore, Singapore, 3Tan Tock Seng Hospital, Singapore, Singapore, 4National University of Singapore, Singapore, Singapore.
1Saw Swee Hock School of Public Health, National University of Singapore, Singapore, Singapore, 2National University Hospital Singapore, Singapore, Singapore, 3Tan Tock Seng Hospital, Singapore, Singapore, 4National University of Singapore, Singapore, Singapore.
OBJECTIVES: Treatment response depth is an important prognostic factor for multiple myeloma (MM). We aimed to predict treatment response depth in a multi-ethnic Singaporean cohort of patients with MM using structured electronic health record (EHR) data augmented with large language model (LLM)-extracted clinical features.
METHODS: We used data from 354 MM patients at the National University Hospital Singapore (2017-2021) and extracted variables required for IMWG response staging from clinical notes using a locally trained LLM (LLAMA 3.1 backbone). We evaluated three prediction frameworks: 5-class response (CR, VGPR, PR, MR, PD), one-vs-rest, and binary (CR+VGPR vs rest), using 3-fold cross-validation and temporal validation (train: 2017-2019, test: 2020-2021). We defined two feature sets: Model A using structured EHR data (demographics, laboratory values, staging, medications) and Model B adding LLM-extracted features (FISH cytogenetics, ECOG performance status, frailty, CRAB symptoms). Performance was measured by macro-averaged area under the receiver operating characteristic curve (AUC). We then compared imputation strategies for handling: median, K-nearest neighbors, and multiple imputation by chained equations [MICE].
RESULTS: The following variables were extracted by LLM but missingness were high as these are currently not routinely assessed: minimal residual disease (MRD) status (48.9), ECOG performance% status (67.0%), ISS staging (derived 44.0%, among active cases), FISH cytogenetics (58.3%. among active cases) and frailty assessment (83.9%). Baseline 5-class AUC improved from 0.61 (cross-validation, structured variables alone) to 0.66 when LLM-extracted clinical features were added. MICE imputation provided further gains to 0.67. Binary prediction (CR+VGPR vs rest) achieved the highest AUC 0.74 (cross-validation) and 0.70 (temporal validation), representing a cumulative improvement of +12%. CR+VGPR was shown previously to predict median progression-free survival.
CONCLUSIONS: LLM-extracted clinical features from unstructured EHR notes improve prediction of treatment response depth in multiple myeloma beyond structured data alone. Response-depth prediction supports downstream applications including trial-eligibility screening, treatment-selection support, and real-world health technology assessment.
METHODS: We used data from 354 MM patients at the National University Hospital Singapore (2017-2021) and extracted variables required for IMWG response staging from clinical notes using a locally trained LLM (LLAMA 3.1 backbone). We evaluated three prediction frameworks: 5-class response (CR, VGPR, PR, MR, PD), one-vs-rest, and binary (CR+VGPR vs rest), using 3-fold cross-validation and temporal validation (train: 2017-2019, test: 2020-2021). We defined two feature sets: Model A using structured EHR data (demographics, laboratory values, staging, medications) and Model B adding LLM-extracted features (FISH cytogenetics, ECOG performance status, frailty, CRAB symptoms). Performance was measured by macro-averaged area under the receiver operating characteristic curve (AUC). We then compared imputation strategies for handling: median, K-nearest neighbors, and multiple imputation by chained equations [MICE].
RESULTS: The following variables were extracted by LLM but missingness were high as these are currently not routinely assessed: minimal residual disease (MRD) status (48.9), ECOG performance% status (67.0%), ISS staging (derived 44.0%, among active cases), FISH cytogenetics (58.3%. among active cases) and frailty assessment (83.9%). Baseline 5-class AUC improved from 0.61 (cross-validation, structured variables alone) to 0.66 when LLM-extracted clinical features were added. MICE imputation provided further gains to 0.67. Binary prediction (CR+VGPR vs rest) achieved the highest AUC 0.74 (cross-validation) and 0.70 (temporal validation), representing a cumulative improvement of +12%. CR+VGPR was shown previously to predict median progression-free survival.
CONCLUSIONS: LLM-extracted clinical features from unstructured EHR notes improve prediction of treatment response depth in multiple myeloma beyond structured data alone. Response-depth prediction supports downstream applications including trial-eligibility screening, treatment-selection support, and real-world health technology assessment.
Conference/Value in Health Info
2026-09, ISPOR Asia Pacific 2026, Bangkok, Thailand
Value in Health, Volume 55, Issue S1
Code
MSR24
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas, SDC: Oncology