IMPROVING IDENTIFICATION OF PATIENTS, CAREGIVERS, AND HEALTHCARE PROFESSIONALS IN SOCIAL MEDIA DATA USING FINE-TUNED PRETRAINED LANGUAGE MODELS IN FRENCH AND ENGLISH
Author(s)
Safaa Mahdir, MEng1, Manissa Talmatkadi, MS2, Amine Chouaki, MEng1, Joelle Malaab, MS, MPH3, Paméla Voillot, MS4, Nathalie Texier, PharmD5, Stéphane SCHÜCK, MD, MPH1.
1ULTIMA-I, Paris, France, 2Kap Code, Paris, France, 3Director of projects and scientific publications, Kap Code, Paris, France, 4Kap Code, PARIS, France, 5kappa Santé, paris, France.
1ULTIMA-I, Paris, France, 2Kap Code, Paris, France, 3Director of projects and scientific publications, Kap Code, Paris, France, 4Kap Code, PARIS, France, 5kappa Santé, paris, France.
OBJECTIVES: Social media and health forums are increasingly used in health economics and outcomes research (HEOR) to capture patient experiences and real-world perspectives. However, their value depends on accurately identifying whether authors are patients, caregivers, healthcare professionals (HCPs), or unrelated responders. Algorithms already exist for this classification, but performance can still be improved, particularly for minority profiles and across languages. This study aimed to improve natural language processing (NLP)-based classification of author profiles in French- and English-language health-related social media data.
METHODS: For French, generic and biomedical-domain pretrained Transformer models, including CamemBERT, DrBERT, and CamemBERT-bio-base, were fine-tuned and compared using a gold-standard dataset of approximately 4,000 messages built from existing labeled data and large language model (LLM)-assisted annotation. Model performance was compared with traditional machine-learning and deep-learning baselines previously used for patient/caregiver and HCP identification (XGBoost and LSTM-CNN). For English, a comparable dataset was developed using the same methodology. Several pretrained models (BioBERT-large, all-mpnet-base-v2, BiomedBERT, PubMedBERT) were fine-tuned and benchmarked against a translation-based baseline. Performance was evaluated using macro-averaged F1-score on held-out test sets. The username was included as an input feature, and a weighted loss addressed class imbalance.
RESULTS: For French, CamemBERT-bio-base achieved the highest performance (macro F1=85%), outperforming XGBoost and LSTM-CNN by 33 percentage points for patient/caregiver identification and 27 percentage points for HCP identification. For English, the translation-based baseline achieved a macro F1-score of 54%, while native fine-tuned models reached up to 73% (all-mpnet-base-v2 with HCP class weighting), confirming that native-language fine-tuning outperforms translation. Inclusion of usernames improved performance macro F1 by 12 percentage points.
CONCLUSIONS: Fine-tuning pretrained language models substantially improved classification of author profiles in French- and English-language health-related social media data. These findings may enhance the validity of social media-based studies used in health outcomes research and real-world evidence through more accurate identification of patients, caregivers, and HCPs.
METHODS: For French, generic and biomedical-domain pretrained Transformer models, including CamemBERT, DrBERT, and CamemBERT-bio-base, were fine-tuned and compared using a gold-standard dataset of approximately 4,000 messages built from existing labeled data and large language model (LLM)-assisted annotation. Model performance was compared with traditional machine-learning and deep-learning baselines previously used for patient/caregiver and HCP identification (XGBoost and LSTM-CNN). For English, a comparable dataset was developed using the same methodology. Several pretrained models (BioBERT-large, all-mpnet-base-v2, BiomedBERT, PubMedBERT) were fine-tuned and benchmarked against a translation-based baseline. Performance was evaluated using macro-averaged F1-score on held-out test sets. The username was included as an input feature, and a weighted loss addressed class imbalance.
RESULTS: For French, CamemBERT-bio-base achieved the highest performance (macro F1=85%), outperforming XGBoost and LSTM-CNN by 33 percentage points for patient/caregiver identification and 27 percentage points for HCP identification. For English, the translation-based baseline achieved a macro F1-score of 54%, while native fine-tuned models reached up to 73% (all-mpnet-base-v2 with HCP class weighting), confirming that native-language fine-tuning outperforms translation. Inclusion of usernames improved performance macro F1 by 12 percentage points.
CONCLUSIONS: Fine-tuning pretrained language models substantially improved classification of author profiles in French- and English-language health-related social media data. These findings may enhance the validity of social media-based studies used in health outcomes research and real-world evidence through more accurate identification of patients, caregivers, and HCPs.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
RWD75
Topic
Patient-Centered Research, Real World Data & Information Systems
Topic Subcategory
Reproducibility & Replicability
Disease
No Additional Disease & Conditions/Specialized Treatment Areas