Abstract
Objectives
This study evaluated the viability of large language models, specifically GPT-4, in predicting patient health preference-consistent choices using a discrete choice experiment framework.
Methods
Synthetic data were generated from real discrete choice experiment responses by patients with a history of cancer. Analytical data included 50 synthetic patients, each answering 48 2-alternative treatment choice questions varying in expected survival, chance of long-term survival, health limitations, and out-of-pocket cost. GPT-4’s predictive performance was assessed across 4 experiments. In experiments 1 and 2, GPT-4 predicted 20 hold-out questions (ie, new choice questions) leveraging 28 fixed (experiment 1) or randomly selected (experiment 2) sample questions. Experiment 3 varied the number of sample questions to examine prediction accuracy and prediction confidence. Experiment 4 evaluated how characteristics of the hold-out questions influenced prediction accuracy.
Results
GPT-4 achieved an average prediction accuracy of 70.5% (95% confidence interval [CI]: 68.3%-72.7%) in experiment 1 and 69.9% (95% CI: 66.9%-72.9%) in experiment 2, with greater variability when sample questions were randomized. Experiment 3 revealed a learning curve, in which accuracy improved from 53% with 5 sample questions to 64% with 10, after which performance plateaued. Experiment 4 showed higher prediction accuracy for questions with more salient attribute differences.
Conclusions
GPT-4 demonstrated the ability to infer patient preferences from limited samples, achieving accuracy levels comparable to surrogate decision makers. Its performance remained consistent across randomized input sequences and improved as the number of sample questions increased, eventually reaching a plateau where additional training yielded diminishing returns.
Authors
Tina Cheng Juan Marcos Gonzalez Matthew M. Engelhard Shelby D. Reed Semra Ozdemir