CAPTURING REAL-WORLD PATIENT EXPERIENCE: CROSS-INDICATION VALIDATION OF AN LLM SOCIAL MEDIA LISTENING PIPELINE WITHOUT RETRAINING
Author(s)
Ana Amaral, MSc1, Natalia Hakimi Hawken, MSc, PhD2, Pranav Khurana, MSc3, Paul Kamudoni, MSc, PhD4.
1Merck Healthcare, Frankfurt am Main, Germany, 2PRWE, Merck Healthcare, Amsterdam, Netherlands, 3Merck Healthcare, Madrid, Spain, 4Merck Healthcare, Darmstadt, Germany.
1Merck Healthcare, Frankfurt am Main, Germany, 2PRWE, Merck Healthcare, Amsterdam, Netherlands, 3Merck Healthcare, Madrid, Spain, 4Merck Healthcare, Darmstadt, Germany.
OBJECTIVES: Social media offers a scalable source of patient experience evidence, but conventional listening studies are slow, costly, single-use, and require methodological redevelopment for each disease area. We developed an end-to-end large language model (LLM) pipeline spanning search-term generation, indication-specific schema instantiation, curation, entity extraction, controlled-vocabulary labeling, and reporting can produce human-comparable annotations and generalize across indications without retraining.
METHODS: The pipeline runs on a fixed architecture with indication as an input parameter: a master signs-and-symptoms schema is adapted per indication by the LLM and reviewed by domain experts before deployment, after which curation, extraction, and labeling proceed without retraining. Deployed across 23 indications, it ingests millions of public posts into a reusable knowledge base. We validated three indications spanning prevalence and discourse complexity: acute ischemic stroke, myasthenia gravis, and NF1-PN. Unmatched entities were routed to an “Other” stratum reviewed for schema gaps. Two reviewers adjudicated stratified samples against source text (curation n=700; extraction/labeling n=425).
RESULTS: Performance was consistent across indications despite a 100-fold corpus-size difference. Curation recall was 98% with 84% precision; the gap largely reflected correctable scope cases (unconfirmed diagnoses, non-human subjects, third-party reports, related stroke subtypes). Extraction faithfulness was 90%/85% (lenient/strict) and completeness 95%/92%; lapses mainly involved dropped modifiers. Among mapped entities, labeling precision was 90% (88-94%), errors predominantly adjacent granularity mismatches. Within the “Other” stratum, 75% were correct out-of-scope and 23% surfaced schema gaps; only 1.8% were true mislabels warranting an existing label.
CONCLUSIONS: This validated LLM pipeline is a fit-for-purpose, scalable solution for real-world patient experience research. Achieving near-human accuracy across three distinct indications without retraining, it converts one-off listening studies into a reusable, re-interrogable evidence asset. Residual errors are correctable through iterative schema refinement, confirming the pipeline's readiness for broad deployment. Social media representativeness remains a limitation.
METHODS: The pipeline runs on a fixed architecture with indication as an input parameter: a master signs-and-symptoms schema is adapted per indication by the LLM and reviewed by domain experts before deployment, after which curation, extraction, and labeling proceed without retraining. Deployed across 23 indications, it ingests millions of public posts into a reusable knowledge base. We validated three indications spanning prevalence and discourse complexity: acute ischemic stroke, myasthenia gravis, and NF1-PN. Unmatched entities were routed to an “Other” stratum reviewed for schema gaps. Two reviewers adjudicated stratified samples against source text (curation n=700; extraction/labeling n=425).
RESULTS: Performance was consistent across indications despite a 100-fold corpus-size difference. Curation recall was 98% with 84% precision; the gap largely reflected correctable scope cases (unconfirmed diagnoses, non-human subjects, third-party reports, related stroke subtypes). Extraction faithfulness was 90%/85% (lenient/strict) and completeness 95%/92%; lapses mainly involved dropped modifiers. Among mapped entities, labeling precision was 90% (88-94%), errors predominantly adjacent granularity mismatches. Within the “Other” stratum, 75% were correct out-of-scope and 23% surfaced schema gaps; only 1.8% were true mislabels warranting an existing label.
CONCLUSIONS: This validated LLM pipeline is a fit-for-purpose, scalable solution for real-world patient experience research. Achieving near-human accuracy across three distinct indications without retraining, it converts one-off listening studies into a reusable, re-interrogable evidence asset. Residual errors are correctable through iterative schema refinement, confirming the pipeline's readiness for broad deployment. Social media representativeness remains a limitation.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
PCR156
Topic
Methodological & Statistical Research, Patient-Centered Research
Topic Subcategory
Patient Engagement
Disease
No Additional Disease & Conditions/Specialized Treatment Areas