STANDARDIZATION OF MENTAL HEALTH ASSESSMENT – USING ITEM RESPONSE THEORY (IRT) TO CROSS-CALIBRATE TWO SELF-REPORTED MENTAL HEALTH TOOLS- THE PATIENT HEALTH QUESTIONNAIRE (PHQ-9) AND THE SF-36V2 MENTAL HEALTH (MH) SCALE
Author(s)
Bjorner JB, White MK, Yarlas AS
Optum, Lincoln, RI, USA
Presentation Documents
OBJECTIVES: Mental health can be measured by numerous instruments, but scores are usually not directly comparable. The heterogeneity of scale specific metrics seriously impairs comparability across study results and the communication among researchers and clinicians. We aimed to develop and evaluate methods for cross-calibration of two popular mental health tools: the PHQ-9 and the SF-36v2 MH scale. METHODS: We analyzed data from the USA and the UK including a general population sample (US: 216, UK: 355) and a sample with suspected depression (US: 169, UK: 153). Multigroup confirmatory bifactor models tested whether the two instruments measured the same construct. Differential item function (DIF) between general population and depression samples was tested using logistic regression DIF tests. We estimated IRT item parameters using a multigroup generalized partial credit model and developed cross-calibration algorithms using the summed score cross-calibration approach. The measurement properties of the instruments were evaluated by test information functions. RESULTS: In the bifactor model, all items loaded strongly on a common factor, supporting that the two scales measure the same general mental health construct. We found no indication of DIF, supporting that the same item parameters apply to the general population and the depression samples. The cross-calibration algorithm revealed a fairly linear relation between PHQ-9 score and MH score in the PHQ-9 score range of 10-20 (moderate to severe depression), but a non-linear relation at more extreme scores. The PHQ-9 provided most information for persons with scores in the interval from the general population average down to two standard deviations below average, but the MH scale provided more information at the lower and upper extremes. CONCLUSIONS: We successfully developed a procedure for cross-calibrating the PHQ-9 and MH scales. These results can be used to compare scores between the two instruments.
Conference/Value in Health Info
2014-05, ISPOR 2014, Palais des Congres de Montreal
Value in Health, Vol. 17, No. 3 (May 2014)
Code
PRM110
Topic
Methodological & Statistical Research
Topic Subcategory
Confounding, Selection Bias Correction, Causal Inference, PRO & Related Methods
Disease
Mental Health