STANDARDIZING PERSONALIZED OUTCOME ASSESSMENT WITH GOAL ATTAINMENT SCALING: RELIABILITY OF A STRUCTURED GOAL APPRAISAL FRAMEWORK AND EVALUATION OF AI-ASSISTED REVIEW
Author(s)
Chere Chapman, MBA, MPH1, Rebecca Metcalfe, PhD2, Autum Mason, MSc1, Jacob Coyne, BSc1, Gunes Sevinc, BSc, MSc, PhD3.
1Ardea Outcomes, Halifax, NS, Canada, 2Ardea Outcomes, Toronto, ON, Canada, 3Director of Patient-Centered Outcomes, Ardea Outcomes, Vancouver, BC, Canada.
1Ardea Outcomes, Halifax, NS, Canada, 2Ardea Outcomes, Toronto, ON, Canada, 3Director of Patient-Centered Outcomes, Ardea Outcomes, Vancouver, BC, Canada.
OBJECTIVES: Goal Attainment Scaling (GAS) is a personalized outcome assessment well-suited for heterogeneous patient populations; however, variability in adherence to goal-formulation and scaling principles can introduce measurement error. This study aimed to: (1) evaluate the interrater reliability (IRR) of a novel, structured 19-item Goal Appraisal Framework, and (2) assess the performance of a Large Language Model (LLM) in using this framework to support clinician review and improve consistency.
METHODS: GAS experts (n=4) developed 40 goal attainment scales that included common structural errors (e.g., vague wording, overlapping levels) to simulate realistic challenges in clinical trials. A structured 19-item Goal Appraisal Framework, designed to assess the psychometric quality of GAS scales using explicit pass/fail criteria, was applied. Phase 1 evaluated IRR between two independent human assessors on all 40 scales (760 items total) using Cohen’s κ and percent agreement. Phase 2 evaluated alignment between an LLM (GPT-5.4, OpenAI) and adjudicated human ratings via an observability platform (Braintrust AI). Finally, a post-hoc error analysis of human-AI disagreements was conducted to identify sources of discordance and opportunities for framework refinement.
RESULTS: In Phase 1, IRR for pass/fail decisions was substantial, with a Cohen’s κ of 0.68 (95% CI:0.61-0.75, z=8.8, p<0.001) and overall percent agreement of 92.2%. In Phase 2, agreement between adjudicated human assessments and the LLM averaged 84.52% across 30 independent runs (SD=10.19), indicating substantial concordance. Post-hoc review revealed ambiguity in appraisals of specificity and measurability.
CONCLUSIONS: These findings provide preliminary evidence that AI-assisted psychometric review may offer a scalable approach to supporting consistency in GAS implementation. Human-AI disagreements primarily highlighted opportunities to refine appraisal criteria within the framework. Further evaluation using real-world GAS data is needed before AI-assisted goal appraisal can be deployed in clinical trials.
METHODS: GAS experts (n=4) developed 40 goal attainment scales that included common structural errors (e.g., vague wording, overlapping levels) to simulate realistic challenges in clinical trials. A structured 19-item Goal Appraisal Framework, designed to assess the psychometric quality of GAS scales using explicit pass/fail criteria, was applied. Phase 1 evaluated IRR between two independent human assessors on all 40 scales (760 items total) using Cohen’s κ and percent agreement. Phase 2 evaluated alignment between an LLM (GPT-5.4, OpenAI) and adjudicated human ratings via an observability platform (Braintrust AI). Finally, a post-hoc error analysis of human-AI disagreements was conducted to identify sources of discordance and opportunities for framework refinement.
RESULTS: In Phase 1, IRR for pass/fail decisions was substantial, with a Cohen’s κ of 0.68 (95% CI:0.61-0.75, z=8.8, p<0.001) and overall percent agreement of 92.2%. In Phase 2, agreement between adjudicated human assessments and the LLM averaged 84.52% across 30 independent runs (SD=10.19), indicating substantial concordance. Post-hoc review revealed ambiguity in appraisals of specificity and measurability.
CONCLUSIONS: These findings provide preliminary evidence that AI-assisted psychometric review may offer a scalable approach to supporting consistency in GAS implementation. Human-AI disagreements primarily highlighted opportunities to refine appraisal criteria within the framework. Further evaluation using real-world GAS data is needed before AI-assisted goal appraisal can be deployed in clinical trials.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
PCR153
Topic
Clinical Outcomes, Methodological & Statistical Research, Patient-Centered Research
Topic Subcategory
Instrument Development, Validation, & Translation, Patient-reported Outcomes & Quality of Life Outcomes
Disease
No Additional Disease & Conditions/Specialized Treatment Areas