EQUIVALENCE OF MEASUREMENT PROPERTIES OF THE DERMATOLOGY LIFE QUALITY INDEX (DLQI) ACROSS 13 EUROPEAN LANGUAGES USING BAYESIAN MULTIGROUP CONFIRMATORY FACTOR ANALYSIS (BAYESIAN MGCFA)
Author(s)
Jeffrey R. Johns, PhD.
Cardiff University, Cardiff, United Kingdom.
Cardiff University, Cardiff, United Kingdom.
OBJECTIVES: To examine measurement invariance of the Dermatology Life Quality Index (DLQI) across 13 European language versions using Bayesian multi-group confirmatory factor analysis (MG-CFA). Bayesian measurement-invariance testing follows the same hierarchy as frequentist approaches (configural, metric, scalar invariance) but evaluates models using Markov Chain Monte-Carlo (MCMC) estimation, posterior distributions, predictive checks and Bayesian model-comparison criteria rather than maximum-likelihood fit indices.
METHODS: DLQI data from 3,408 patients were analysed across all language groups using R-blavaan. Model convergence was assessed using Effective Sample Size (ESS), Gelman-Rubin convergence statistics (R-hat), traceplots and autocorrelation diagnostics. Posterior predictive checks were used to assess model fit, and Bayesian results were compared with previously reported frequentist MG-CFA findings.
RESULTS: All models demonstrated excellent convergence and adequate fit, with posterior predictive checks indicating that replicated data closely reproduced observed response-category frequencies across all DLQI items. No item showed evidence of systematic misfit despite substantial differences in response distributions between language groups. Factor loadings were comparable across groups, supporting metric invariance. However, item thresholds differed systematically, indicating failure of scalar invariance. Bayesian and frequentist estimates were highly concordant, with only minor shrinkage of loadings under Bayesian estimation, consistent with expected regularisation effects. No systematic inflation or attenuation of parameter estimates was observed, and items with lower frequentist loadings did not show disproportionate changes under Bayesian estimation. The close agreement between approaches suggests that conclusions were robust to the estimation framework.
CONCLUSIONS: Respondents across language groups appear to interpret DLQI items similarly, supporting metric invariance. However, differences in item thresholds indicate that response endorsement patterns vary across languages even at equivalent levels of latent dermatology-related quality-of-life impairment. Consequently, direct comparison of observed scores or mean values across language groups should be undertaken with caution. These findings replicate previous frequentist MG-CFA results, demonstrating consistency across Bayesian and frequentist frameworks.
METHODS: DLQI data from 3,408 patients were analysed across all language groups using R-blavaan. Model convergence was assessed using Effective Sample Size (ESS), Gelman-Rubin convergence statistics (R-hat), traceplots and autocorrelation diagnostics. Posterior predictive checks were used to assess model fit, and Bayesian results were compared with previously reported frequentist MG-CFA findings.
RESULTS: All models demonstrated excellent convergence and adequate fit, with posterior predictive checks indicating that replicated data closely reproduced observed response-category frequencies across all DLQI items. No item showed evidence of systematic misfit despite substantial differences in response distributions between language groups. Factor loadings were comparable across groups, supporting metric invariance. However, item thresholds differed systematically, indicating failure of scalar invariance. Bayesian and frequentist estimates were highly concordant, with only minor shrinkage of loadings under Bayesian estimation, consistent with expected regularisation effects. No systematic inflation or attenuation of parameter estimates was observed, and items with lower frequentist loadings did not show disproportionate changes under Bayesian estimation. The close agreement between approaches suggests that conclusions were robust to the estimation framework.
CONCLUSIONS: Respondents across language groups appear to interpret DLQI items similarly, supporting metric invariance. However, differences in item thresholds indicate that response endorsement patterns vary across languages even at equivalent levels of latent dermatology-related quality-of-life impairment. Consequently, direct comparison of observed scores or mean values across language groups should be undertaken with caution. These findings replicate previous frequentist MG-CFA results, demonstrating consistency across Bayesian and frequentist frameworks.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR112
Topic
Clinical Outcomes, Methodological & Statistical Research
Topic Subcategory
PRO & Related Methods
Disease
No Additional Disease & Conditions/Specialized Treatment Areas