EXTRACTION OF COMPLEX IBD FEATURES FROM CLINICAL NOTES USING HUMMINGBIRD: A COST-EFFECTIVE, PRIVACY-PRESERVING ALTERNATIVE TO LARGE LANGUAGE MODELS

Author(s)

Ishtiyaque Ahmad, PhD1, Yuntao Zou, MD, PhD2, Shadera Azzam, MSGH, RDN2, Vivek Rudrapatna, MD, PhD2.
1DataUnite, Cupertino, CA, USA, 2University of California San Francisco, San Francisco, CA, USA.
OBJECTIVES: Accurate extraction of complex clinical features from clinical notes is critical for real-world evidence generation, as many clinically relevant details are not captured in structured EHRs. While expert manual annotation remains the gold standard, it is labor-intensive and difficult to scale. Large language models (LLMs) offer a promising alternative but introduce privacy concerns due to sharing of notes with external service providers and can be prohibitively expensive for large-scale extraction. This study evaluates the accuracy and cost-effectiveness of Hummingbird, a lightweight natural language processing tool that runs on commodity CPU-hardware without requiring GPUs. Unlike LLMs, Hummingbird processes clinical notes entirely within the hospital environment, preserving patient privacy and reducing costs. Hummingbird was evaluated for extracting inflammatory bowel disease (IBD) complications from clinical notes within UCSF.
METHODS: Four clinically complex IBD complications were extracted: active fistula, symptomatic intestinal stricture, history of bowel resection, and intestinal ostomy. These tasks required interpretation of surgical history, procedure reversibility, and temporal context. A gastroenterologist developed detailed feature definitions and extraction criteria, which were provided to both Hummingbird and ChatGPT-5 for feature extraction. Performance was evaluated against gold-standard annotations from 77 manually reviewed clinical notes. Accuracy, precision, recall, and F1-score were assessed, and estimated costs were calculated from ChatGPT-5 token usage and AWS compute pricing for Hummingbird CPU usage.
RESULTS: Hummingbird achieved F1-scores of 0.86 (active fistula), 0.78 (symptomatic intestinal stricture), 0.91 ( history of bowel resection), and 0.87 (intestinal ostomy), versus ChatGPT-5 F1-scores of 0.74, 0.27, 0.84, and 0.92 respectively. Estimated extraction cost for 77 notes were approximately $1.10 for ChatGPT-5, compared to under $0.03 for Hummingbird, a 36x cost reduction.
CONCLUSIONS: Hummingbird achieved performance comparable to ChatGPT-5 on IBD complications phenotyping from clinical notes, while preserving privacy and reducing cost by 36-times. This will reduce the burden of manual chart review and enable large-cohort epidemiologic research.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR70

Topic

Health Technology Assessment, Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

Gastrointestinal Disorders

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×