LOW-COST FEATURE EXTRACTION FROM CLINICAL NOTES USING HUMMINGBIRD: ECOG PERFORMANCE STATUS OF LUNG CANCER PATIENTS
Author(s)
Ishtiyaque Ahmad, PhD1, Vivek Rudrapatna, MD, PhD2, Trinabh Gupta, PhD1.
1DataUnite, Cupertino, CA, USA, 2University of California San Francisco, San Francisco, CA, USA.
1DataUnite, Cupertino, CA, USA, 2University of California San Francisco, San Francisco, CA, USA.
OBJECTIVES: Clinical information extraction from unstructured notes remains challenging. Manual annotation is not scalable; rule-based methods have limited generalizability, and large language models (LLMs) are constrained by privacy and cost. We evaluated Hummingbird, a lightweight CPU-based NLP framework, for extracting Eastern Cooperative Oncology Group (ECOG) performance status from oncology notes, enabling large-scale characterization of functional status and trial-relevant populations in real-world lung cancer care, while comparing cost with GPT-5.4-mini.
METHODS: Adult patients (≥18 years) with a first primary lung cancer diagnosis (ICD-10-CM C34.x; excluding C78.0x) between January 1, 2020 and January 31, 2026 were identifiedfrom the UCSF deidentified EHR. Hummingbird was deployed on CPU-only infrastructure and evaluated on 200 manually annotated notes. ECOG extraction was formulated as a multi-class classification task (ECOG 0-4 and no documented ECOG). Extraction costs were estimated projecting GPT-5.4-mini API token usage and measuring Hummingbird CPU utilization cost for AWS.
RESULTS: The cohort included 5,923 patients contributing 39,765 lung oncology notes containing the pattern “ECOG”. Hummingbird achieved 98.5% accuracy and a weighted F1-score of 98.6% for ECOG classification against manual review. Overall, 51% of patients had at least one documented ECOG assessment; among them, 79.9% had a baseline ECOG of 0-1 and 20.1% had ECOG ≥2. A total of 2,420 patients had at least one follow-up ECOG assessment, with a median of 7 assessments (IQR [3-16]) over a median follow-up of 8.6 months. Among patients with serial assessments, 44.4% experienced worsening of ECOG score. The estimated extraction cost was $119.3 using GPT-5.4-mini versus $2.16 using Hummingbird, a 55-fold reduction in cost.
CONCLUSIONS: Hummingbird extracted ECOG performance status from notes with high accuracy while enabling large-scale characterization of functional status in a real-world lung cancer population. With 55x lower cost than frontier-model APIs, Hummingbird provides a scalable, privacy-preserving approach for generating clinically meaningful real-world evidence.
METHODS: Adult patients (≥18 years) with a first primary lung cancer diagnosis (ICD-10-CM C34.x; excluding C78.0x) between January 1, 2020 and January 31, 2026 were identifiedfrom the UCSF deidentified EHR. Hummingbird was deployed on CPU-only infrastructure and evaluated on 200 manually annotated notes. ECOG extraction was formulated as a multi-class classification task (ECOG 0-4 and no documented ECOG). Extraction costs were estimated projecting GPT-5.4-mini API token usage and measuring Hummingbird CPU utilization cost for AWS.
RESULTS: The cohort included 5,923 patients contributing 39,765 lung oncology notes containing the pattern “ECOG”. Hummingbird achieved 98.5% accuracy and a weighted F1-score of 98.6% for ECOG classification against manual review. Overall, 51% of patients had at least one documented ECOG assessment; among them, 79.9% had a baseline ECOG of 0-1 and 20.1% had ECOG ≥2. A total of 2,420 patients had at least one follow-up ECOG assessment, with a median of 7 assessments (IQR [3-16]) over a median follow-up of 8.6 months. Among patients with serial assessments, 44.4% experienced worsening of ECOG score. The estimated extraction cost was $119.3 using GPT-5.4-mini versus $2.16 using Hummingbird, a 55-fold reduction in cost.
CONCLUSIONS: Hummingbird extracted ECOG performance status from notes with high accuracy while enabling large-scale characterization of functional status in a real-world lung cancer population. With 55x lower cost than frontier-model APIs, Hummingbird provides a scalable, privacy-preserving approach for generating clinically meaningful real-world evidence.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR144
Topic
Methodological & Statistical Research, Real World Data & Information Systems
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas, Oncology