A LARGE LANGUAGE MODEL PIPELINE FOR RAPID QUALITATIVE RE-ANALYSIS OF WITHIN-TRIAL PATIENT INTERVIEWS: A HEALTH-RELATED QUALITY-OF-LIFE CASE STUDY

Author(s)

Piper Fromy, PhD1, Mariam Bayram, PhD2.
1SeeingTheta, Souzay-Champigny, France, 2Souzay-Champigny, France.
OBJECTIVES: Qualitative interview data collected during clinical trials are often analysed for a predefined purpose. Secondary research questions and patient-centered themes may arise following results. We evaluated whether a large language model (LLM) pipeline could enable rapid re-analysis of existing trial interview transcripts to characterise specific health-related quality-of-life (HRQoL) themes the original analysis did not target, and to recover patient-level data enabling within-patient comparisons originally unaccounted for.
METHODS: The process started with per-patient structured extraction by an LLM custom agent, contextualised to the trial design, followed by quantitative aggregation in R. Double-transcript patients enabled within-patient contrasts between placebo and active-drug periods, the central advance over the period-stratified original analysis. Outputs were one pipe-delimited row per patient for tabulated aggregation. One prompt specified explicit field numbers and content, conditional formatting, and active self-checks. Extraction ran within a secure, access-controlled enterprise environment, each patient processed in one isolated agent session to preserve accuracy and avoid cross-patient contamination.
RESULTS: Structured HRQoL coding was obtained rapidly relative to manual coding. LLM extraction was sensitive to transcript heterogeneity, with field misalignment the dominant challenge, mitigated but not eliminated by self-checks. Misalignment followed empty data fields: quality control confirmed the explicit prompt eliminated false or hallucinated data, yielding empty fields rather than fabrication. Single-transcript outputs showed less misalignment, likely due to shorter expected output. Quantitative aggregation in R succeeded, the main challenge being parsing and coding of textual data.
CONCLUSIONS: A structured LLM pipeline can facilitate rapid within-patient secondary HRQoL analyses of existing trial interview data within a predefined concept. Reliability depends less on model capability than on explicit prompt engineering and transcript-aware handling. This pipeline can quickly uncover new HRQoL themes and their interplay, informing future study design.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

CO47

Topic

Clinical Outcomes, Patient-Centered Research

Topic Subcategory

Clinical Outcomes Assessment

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×