EXTENDING THE VALIDATION OF ACCURACY FOR LARGE LANGUAGE MODEL (LLM)-/MACHINE LEARNING (ML)-EXTRACTED INFORMATION AND DATA (VALID) FRAMEWORK GLOBALLY: QUALITY ASSESSMENT OF LLM-DERIVED PROSTATE CANCER REAL-WORLD DATA IN THE UK
Author(s)
Victor Lhoste, PhD1, Arun Sujenthiran, MD, FCRS1, Kathi Seidl-Rathkopf, PhD2, Natalia Viani, PhD1, Marina Kushnir, MS1, Nikola Dolezalova, PhD1, Anna Schwarz, MS1, Amin Hashemian, MS1, Conor Walsh, MS1, Alex Enrique, BA1, Golnessa Mojtahedi, MS3, Cornelius Thaiss, MD2, Patrycja Pluta, DVM2, Qianyi Zhang, MS3.
1Flatiron Health UK, London, United Kingdom, 2Flatiron Health Germany, Berlin, Germany, 3Flatiron Health, New York, NY, USA.
1Flatiron Health UK, London, United Kingdom, 2Flatiron Health Germany, Berlin, Germany, 3Flatiron Health, New York, NY, USA.
OBJECTIVES: The VALID framework assesses LLM-derived real-world data (RWD) quality across three dimensions: variable-level metrics (VLM), verification checks, and replication analyses. This study applied VALID to an LLM-derived UK prostate cancer (PC) dataset.
METHODS: LLMs selected patients with PC from the UK-based, electronic health record-derived, deidentified Flatiron Health Research Database and extracted clinically meaningful characteristics, including initial/metastatic diagnosis, castration-resistant status, treatment, and outcome. LLM-derived data were benchmarked against human-abstracted data. For VLM, validation sets of 587 patients were doubly abstracted. Verification checks assessed the prevalence of conflicting or clinically implausible data points. Replication benchmarked patient characteristics and real-world overall survival (rwOS) against human-abstracted data, and national registry data. All data were processed in a UK GDPR-compliant manner.
RESULTS: The LLM-derived dataset included 13,901 patients with PC. For VLM, F1 scores were within 2.6 (initial diagnosis date) and 2.3 (metastatic diagnosis date) percentage points of abstractor performance (30-day tolerance). Verification checks confirmed radical prostatectomy was rare among metastatic de novo cases (0.3%) vs. recurrent cases (15.2%). Replication showed similar rwOS from metastatic diagnosis between LLM-derived and abstracted datasets for patients with greater than 1 line containing an androgen receptor pathway inhibitor (ARPI) or taxane in the metastatic castration-resistant setting (median rwOS [months, 95% CI]: ARPI: 44.8 [41.5-47.8] vs. 42.3 [38.7-47.8]; taxane: 37.1 [33.3-42.0] vs. 37.0 [33.1-44.6]). LLM-derived versus national cancer registry data showed consistent 1-year survival by stage (stage III: 98.9% vs. 97.7%; stage IV: 85.4% vs. 84.3%) and 2021 proportions metastatic at diagnosis (18.6% vs. 16.4%) and receiving radical prostatectomy (15.6% vs. 14.6%).
CONCLUSIONS: Applying the VALID framework to a UK PC dataset indicated that LLMs can extract data suitable for generating accurate and reliable real-world evidence.
METHODS: LLMs selected patients with PC from the UK-based, electronic health record-derived, deidentified Flatiron Health Research Database and extracted clinically meaningful characteristics, including initial/metastatic diagnosis, castration-resistant status, treatment, and outcome. LLM-derived data were benchmarked against human-abstracted data. For VLM, validation sets of 587 patients were doubly abstracted. Verification checks assessed the prevalence of conflicting or clinically implausible data points. Replication benchmarked patient characteristics and real-world overall survival (rwOS) against human-abstracted data, and national registry data. All data were processed in a UK GDPR-compliant manner.
RESULTS: The LLM-derived dataset included 13,901 patients with PC. For VLM, F1 scores were within 2.6 (initial diagnosis date) and 2.3 (metastatic diagnosis date) percentage points of abstractor performance (30-day tolerance). Verification checks confirmed radical prostatectomy was rare among metastatic de novo cases (0.3%) vs. recurrent cases (15.2%). Replication showed similar rwOS from metastatic diagnosis between LLM-derived and abstracted datasets for patients with greater than 1 line containing an androgen receptor pathway inhibitor (ARPI) or taxane in the metastatic castration-resistant setting (median rwOS [months, 95% CI]: ARPI: 44.8 [41.5-47.8] vs. 42.3 [38.7-47.8]; taxane: 37.1 [33.3-42.0] vs. 37.0 [33.1-44.6]). LLM-derived versus national cancer registry data showed consistent 1-year survival by stage (stage III: 98.9% vs. 97.7%; stage IV: 85.4% vs. 84.3%) and 2021 proportions metastatic at diagnosis (18.6% vs. 16.4%) and receiving radical prostatectomy (15.6% vs. 14.6%).
CONCLUSIONS: Applying the VALID framework to a UK PC dataset indicated that LLMs can extract data suitable for generating accurate and reliable real-world evidence.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR263
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
Oncology