VALIDATING LLM-ASSISTED CLASSIFICATION OF PATIENT VOICE FROM SOCIAL MEDIA: A DUAL-ARTIFACT WORKFLOW WITH HELD-OUT REPRODUCIBILITY, DEMONSTRATED IN FERTILITY CARE
Author(s)
Rushabh Mehta, Post Graduate Program1, Varsha Patial, Dr., BAMS, MBA2, Vaidehi Ketkar3, Rahul Srivastava, BTech, MBA4.
1SIROai, Mumbai, India, 2Aksigen IVF, Mumbai, India, 3AKSIGEN IVF, India, 4SIRO Clinpharm Pvt. Ltd., Mumbai, India.
1SIROai, Mumbai, India, 2Aksigen IVF, Mumbai, India, 3AKSIGEN IVF, India, 4SIRO Clinpharm Pvt. Ltd., Mumbai, India.
OBJECTIVES: To validate a workflow for LLM-assisted classification of patient voice from social media using separately iterated human-facing codebook and LLM-facing prompt instruments, and to demonstrate its reproducibility on a large-scale fertility patient-voice corpus from Reddit and YouTube using a held-out confirmation sample unseen during development.
METHODS: 22,619 multilingual posts and comments referencing fertility care were extracted post filtration. A four-domain codebook (Information Gaps, Decision Conflict, Emotional Distress, Navigation Issues) and a separate LLM-facing classification prompt were developed and maintained as two independently versioned artifacts. Two human raters classified a development sample (n=111); Claude Opus 4.7 classified the same items via API. Disagreements were diagnosed case-by-case: AI-human disagreements were either codebook gaps the AI’s literal-mindedness had surfaced (addressed in the codebook and prompt) or AI execution gaps requiring instructions meaningful only to the model (anti-keyword tests, removal rules; added to the prompt only). A held-out sample (n=80) unseen during instrument iteration was classified by both human raters and AI, using Cohen’s κ to compute agreement.
RESULTS: In the held-out sample, all six dimensions cleared κ ≥ 0.60. AI-human means: Unclassifiable 0.83, Information Gaps 0.79, Decision Conflict 0.86, Emotional Distress 0.85, Navigation Issues 0.97, Primary domain 0.86; human-human means ≥ 0.92 across dimensions. Navigation Issues, the most error prone domain in development (κ=0.45), reached κ=0.97 in the held out partition after prompt only remediation.
CONCLUSIONS: Maintaining the codebook and the prompt as separately iterated instruments and routing each disagreement to its source, produced reproducible AI-human agreement across six dimensions. The AI’s literal application of rules served as a diagnostic probe that surfaced codebook ambiguities humans tolerated silently, while genuine execution gaps were remediated in the prompt. The workflow enables scalable, auditable classification of patient voice from social media and is portable to all therapies where patient experience is expressed in unstructured text.
METHODS: 22,619 multilingual posts and comments referencing fertility care were extracted post filtration. A four-domain codebook (Information Gaps, Decision Conflict, Emotional Distress, Navigation Issues) and a separate LLM-facing classification prompt were developed and maintained as two independently versioned artifacts. Two human raters classified a development sample (n=111); Claude Opus 4.7 classified the same items via API. Disagreements were diagnosed case-by-case: AI-human disagreements were either codebook gaps the AI’s literal-mindedness had surfaced (addressed in the codebook and prompt) or AI execution gaps requiring instructions meaningful only to the model (anti-keyword tests, removal rules; added to the prompt only). A held-out sample (n=80) unseen during instrument iteration was classified by both human raters and AI, using Cohen’s κ to compute agreement.
RESULTS: In the held-out sample, all six dimensions cleared κ ≥ 0.60. AI-human means: Unclassifiable 0.83, Information Gaps 0.79, Decision Conflict 0.86, Emotional Distress 0.85, Navigation Issues 0.97, Primary domain 0.86; human-human means ≥ 0.92 across dimensions. Navigation Issues, the most error prone domain in development (κ=0.45), reached κ=0.97 in the held out partition after prompt only remediation.
CONCLUSIONS: Maintaining the codebook and the prompt as separately iterated instruments and routing each disagreement to its source, produced reproducible AI-human agreement across six dimensions. The AI’s literal application of rules served as a diagnostic probe that surfaced codebook ambiguities humans tolerated silently, while genuine execution gaps were remediated in the prompt. The workflow enables scalable, auditable classification of patient voice from social media and is portable to all therapies where patient experience is expressed in unstructured text.
Conference/Value in Health Info
2026-09, ISPOR Asia Pacific 2026, Bangkok, Thailand
Value in Health, Volume 55, Issue S1
Code
MSR27
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas, SDC: Reproductive & Sexual Health