EXTENDING THE VALIDATION OF ACCURACY FOR LARGE LANGUAGE MODEL (LLM)-/MACHINE LEARNING (ML)-EXTRACTED INFORMATION AND DATA (VALID) FRAMEWORK GLOBALLY: QUALITY ASSESSMENT OF LLM-DERIVED PROSTATE CANCER REAL-WORLD DATA IN THE UK

Author(s)

Victor Lhoste, PhD1, Arun Sujenthiran, MD, FCRS1, Kathi Seidl-Rathkopf, PhD2, Natalia Viani, PhD1, Marina Kushnir, MS1, Nikola Dolezalova, PhD1, Anna Schwarz, MS1, Amin Hashemian, MS1, Conor Walsh, MS1, Alex Enrique, BA1, Golnessa Mojtahedi, MS3, Cornelius Thaiss, MD2, Patrycja Pluta, DVM2, Qianyi Zhang, MS3.
1Flatiron Health UK, London, United Kingdom, 2Flatiron Health Germany, Berlin, Germany, 3Flatiron Health, New York, NY, USA.
OBJECTIVES: The VALID framework assesses LLM-derived real-world data (RWD) quality across three dimensions: variable-level metrics (VLM), verification checks, and replication analyses. This study applied VALID to an LLM-derived UK prostate cancer (PC) dataset.
METHODS: LLMs selected patients with PC from the UK-based, electronic health record-derived, deidentified Flatiron Health Research Database and extracted clinically meaningful characteristics, including initial/metastatic diagnosis, castration-resistant status, treatment, and outcome. LLM-derived data were benchmarked against human-abstracted data. For VLM, validation sets of 587 patients were doubly abstracted. Verification checks assessed the prevalence of conflicting or clinically implausible data points. Replication benchmarked patient characteristics and real-world overall survival (rwOS) against human-abstracted data, and national registry data. All data were processed in a UK GDPR-compliant manner.
RESULTS: The LLM-derived dataset included 13,901 patients with PC. For VLM, F1 scores were within 2.6 (initial diagnosis date) and 2.3 (metastatic diagnosis date) percentage points of abstractor performance (30-day tolerance). Verification checks confirmed radical prostatectomy was rare among metastatic de novo cases (0.3%) vs. recurrent cases (15.2%). Replication showed similar rwOS from metastatic diagnosis between LLM-derived and abstracted datasets for patients with greater than 1 line containing an androgen receptor pathway inhibitor (ARPI) or taxane in the metastatic castration-resistant setting (median rwOS [months, 95% CI]: ARPI: 44.8 [41.5-47.8] vs. 42.3 [38.7-47.8]; taxane: 37.1 [33.3-42.0] vs. 37.0 [33.1-44.6]). LLM-derived versus national cancer registry data showed consistent 1-year survival by stage (stage III: 98.9% vs. 97.7%; stage IV: 85.4% vs. 84.3%) and 2021 proportions metastatic at diagnosis (18.6% vs. 16.4%) and receiving radical prostatectomy (15.6% vs. 14.6%).
CONCLUSIONS: Applying the VALID framework to a UK PC dataset indicated that LLMs can extract data suitable for generating accurate and reliable real-world evidence.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR263

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

Oncology

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×