SAMPLING OF COMMONLY USED POPULATION CHARACTERISTICS- IS A NORMAL APPROXIMATION VALID?
Author(s)
Davis JA, Mittard V, Saunders R
Coreva Scientific, Freiburg im Breisgau, Germany
Presentation Documents
OBJECTIVES: Health economic models use a basecase that is generally representative of a subpopulation rather than the whole population. During sensitivity analysis, extrapolation of the model to other subpopulations or the whole population is estimated via sampling. Sampling is performed using summary statistics (e.g. mean and standard deviation) to inform generation of a distribution from which to draw values at random. Key population characteristics for healthcare include age, height, weight, and body mass index (BMI); all of which are commonly assumed to approximate to a normal distribution. Here the plausibility of this common assumption is tested. METHODS: Full data (N=451,075) were obtained from the 2010 Behavioral Risk Factor Surveillance System (BRFSS), a national, US, health-related, telephone survey. Data collected include age, gender, height and weight, with BMI being a calculated variable. Summary statistics and distributions were produced from the whole population. A sample of 2,500 records were extracted for in-depth analysis. Of these, 2,365 had complete data for age, gender, height, and weight. Analyses performed in R and Microsoft Excel® included subsampling, normality and Cullen-Frey tests. RESULTS: None of the data assessed were normally distributed. Cullen-Frey plots indicate that the best distributions to approximate the data are Beta, Log-normal, Beta, Log-normal for age, weight, height and BMI, respectively. Taking 1,000 subsamples of 300 patients, 67% of samples had a mean age falling outside of the 99% confidence interval for the population. For BMI the percentage was 62%. The ability of progressively smaller subsamples to represent the population was progressively worse. CONCLUSIONS: Many population characteristics of interest to healthcare do not follow a normal distribution. In the BRFSS dataset, the most descriptive distributions are the log-normal for BMI and the Beta distribution with negative skew for age. Age distribution skew may represent the aging population in the US setting.
Conference/Value in Health Info
2016-10, ISPOR Europe 2016, Vienna, Austria
Value in Health, Vol. 19, No. 7 (November 2016)
Code
PRM99
Topic
Methodological & Statistical Research
Topic Subcategory
Modeling and simulation
Disease
Multiple Diseases