IMPROVING IDENTIFICATION OF PATIENTS, CAREGIVERS, AND HEALTHCARE PROFESSIONALS IN SOCIAL MEDIA DATA USING FINE-TUNED PRETRAINED LANGUAGE MODELS IN FRENCH AND ENGLISH

Author(s)

Safaa Mahdir, MEng1, Manissa Talmatkadi, MS2, Amine Chouaki, MEng1, Joelle Malaab, MS, MPH3, Paméla Voillot, MS4, Nathalie Texier, PharmD5, Stéphane SCHÜCK, MD, MPH1.
1ULTIMA-I, Paris, France, 2Kap Code, Paris, France, 3Director of projects and scientific publications, Kap Code, Paris, France, 4Kap Code, PARIS, France, 5kappa Santé, paris, France.
OBJECTIVES: Social media and health forums are increasingly used in health economics and outcomes research (HEOR) to capture patient experiences and real-world perspectives. However, their value depends on accurately identifying whether authors are patients, caregivers, healthcare professionals (HCPs), or unrelated responders. Algorithms already exist for this classification, but performance can still be improved, particularly for minority profiles and across languages. This study aimed to improve natural language processing (NLP)-based classification of author profiles in French- and English-language health-related social media data.
METHODS: For French, generic and biomedical-domain pretrained Transformer models, including CamemBERT, DrBERT, and CamemBERT-bio-base, were fine-tuned and compared using a gold-standard dataset of approximately 4,000 messages built from existing labeled data and large language model (LLM)-assisted annotation. Model performance was compared with traditional machine-learning and deep-learning baselines previously used for patient/caregiver and HCP identification (XGBoost and LSTM-CNN). For English, a comparable dataset was developed using the same methodology. Several pretrained models (BioBERT-large, all-mpnet-base-v2, BiomedBERT, PubMedBERT) were fine-tuned and benchmarked against a translation-based baseline. Performance was evaluated using macro-averaged F1-score on held-out test sets. The username was included as an input feature, and a weighted loss addressed class imbalance.
RESULTS: For French, CamemBERT-bio-base achieved the highest performance (macro F1=85%), outperforming XGBoost and LSTM-CNN by 33 percentage points for patient/caregiver identification and 27 percentage points for HCP identification. For English, the translation-based baseline achieved a macro F1-score of 54%, while native fine-tuned models reached up to 73% (all-mpnet-base-v2 with HCP class weighting), confirming that native-language fine-tuning outperforms translation. Inclusion of usernames improved performance macro F1 by 12 percentage points.
CONCLUSIONS: Fine-tuning pretrained language models substantially improved classification of author profiles in French- and English-language health-related social media data. These findings may enhance the validity of social media-based studies used in health outcomes research and real-world evidence through more accurate identification of patients, caregivers, and HCPs.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

RWD75

Topic

Patient-Centered Research, Real World Data & Information Systems

Topic Subcategory

Reproducibility & Replicability

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×