VALIDATING LLM-ASSISTED CLASSIFICATION OF PATIENT VOICE FROM SOCIAL MEDIA: A DUAL-ARTIFACT WORKFLOW WITH HELD-OUT REPRODUCIBILITY, DEMONSTRATED IN FERTILITY CARE

Author(s)

Rushabh Mehta, Post Graduate Program1, Varsha Patial, Dr., BAMS, MBA2, Vaidehi Ketkar3, Rahul Srivastava, BTech, MBA4.
1SIROai, Mumbai, India, 2Aksigen IVF, Mumbai, India, 3AKSIGEN IVF, India, 4SIRO Clinpharm Pvt. Ltd., Mumbai, India.
OBJECTIVES: To validate a workflow for LLM-assisted classification of patient voice from social media using separately iterated human-facing codebook and LLM-facing prompt instruments, and to demonstrate its reproducibility on a large-scale fertility patient-voice corpus from Reddit and YouTube using a held-out confirmation sample unseen during development.
METHODS: 22,619 multilingual posts and comments referencing fertility care were extracted post filtration. A four-domain codebook (Information Gaps, Decision Conflict, Emotional Distress, Navigation Issues) and a separate LLM-facing classification prompt were developed and maintained as two independently versioned artifacts. Two human raters classified a development sample (n=111); Claude Opus 4.7 classified the same items via API. Disagreements were diagnosed case-by-case: AI-human disagreements were either codebook gaps the AI’s literal-mindedness had surfaced (addressed in the codebook and prompt) or AI execution gaps requiring instructions meaningful only to the model (anti-keyword tests, removal rules; added to the prompt only). A held-out sample (n=80) unseen during instrument iteration was classified by both human raters and AI, using Cohen’s κ to compute agreement.
RESULTS: In the held-out sample, all six dimensions cleared κ ≥ 0.60. AI-human means: Unclassifiable 0.83, Information Gaps 0.79, Decision Conflict 0.86, Emotional Distress 0.85, Navigation Issues 0.97, Primary domain 0.86; human-human means ≥ 0.92 across dimensions. Navigation Issues, the most error prone domain in development (κ=0.45), reached κ=0.97 in the held out partition after prompt only remediation.
CONCLUSIONS: Maintaining the codebook and the prompt as separately iterated instruments and routing each disagreement to its source, produced reproducible AI-human agreement across six dimensions. The AI’s literal application of rules served as a diagnostic probe that surfaced codebook ambiguities humans tolerated silently, while genuine execution gaps were remediated in the prompt. The workflow enables scalable, auditable classification of patient voice from social media and is portable to all therapies where patient experience is expressed in unstructured text.

Conference/Value in Health Info

2026-09, ISPOR Asia Pacific 2026, Bangkok, Thailand

Value in Health, Volume 55, Issue S1

Code

MSR27

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas, SDC: Reproductive & Sexual Health

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×