STANDARDIZING PERSONALIZED OUTCOME ASSESSMENT WITH GOAL ATTAINMENT SCALING: RELIABILITY OF A STRUCTURED GOAL APPRAISAL FRAMEWORK AND EVALUATION OF AI-ASSISTED REVIEW

Author(s)

Chere Chapman, MBA, MPH1, Rebecca Metcalfe, PhD2, Autum Mason, MSc1, Jacob Coyne, BSc1, Gunes Sevinc, BSc, MSc, PhD3.
1Ardea Outcomes, Halifax, NS, Canada, 2Ardea Outcomes, Toronto, ON, Canada, 3Director of Patient-Centered Outcomes, Ardea Outcomes, Vancouver, BC, Canada.
OBJECTIVES: Goal Attainment Scaling (GAS) is a personalized outcome assessment well-suited for heterogeneous patient populations; however, variability in adherence to goal-formulation and scaling principles can introduce measurement error. This study aimed to: (1) evaluate the interrater reliability (IRR) of a novel, structured 19-item Goal Appraisal Framework, and (2) assess the performance of a Large Language Model (LLM) in using this framework to support clinician review and improve consistency.
METHODS: GAS experts (n=4) developed 40 goal attainment scales that included common structural errors (e.g., vague wording, overlapping levels) to simulate realistic challenges in clinical trials. A structured 19-item Goal Appraisal Framework, designed to assess the psychometric quality of GAS scales using explicit pass/fail criteria, was applied. Phase 1 evaluated IRR between two independent human assessors on all 40 scales (760 items total) using Cohen’s κ and percent agreement. Phase 2 evaluated alignment between an LLM (GPT-5.4, OpenAI) and adjudicated human ratings via an observability platform (Braintrust AI). Finally, a post-hoc error analysis of human-AI disagreements was conducted to identify sources of discordance and opportunities for framework refinement.
RESULTS: In Phase 1, IRR for pass/fail decisions was substantial, with a Cohen’s κ of 0.68 (95% CI:0.61-0.75, z=8.8, p<0.001) and overall percent agreement of 92.2%. In Phase 2, agreement between adjudicated human assessments and the LLM averaged 84.52% across 30 independent runs (SD=10.19), indicating substantial concordance. Post-hoc review revealed ambiguity in appraisals of specificity and measurability.
CONCLUSIONS: These findings provide preliminary evidence that AI-assisted psychometric review may offer a scalable approach to supporting consistency in GAS implementation. Human-AI disagreements primarily highlighted opportunities to refine appraisal criteria within the framework. Further evaluation using real-world GAS data is needed before AI-assisted goal appraisal can be deployed in clinical trials.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

PCR153

Topic

Clinical Outcomes, Methodological & Statistical Research, Patient-Centered Research

Topic Subcategory

Instrument Development, Validation, & Translation, Patient-reported Outcomes & Quality of Life Outcomes

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×