Assessment of the Relationship between Collision Rate and Sample Size Using a Large US Mortality Dataset
Author(s)
Huynh S1, Liu T2, Leshin J2, Haskell T3
1Kantar, New York, NY, USA, 2Datavant, San Francisco, CA, USA, 3Kantar, Havertown, PA, USA
Presentation Documents
OBJECTIVES: To estimate expected collision rate in a large US mortality dataset and specifically examine the relationship between collision rate and sample size. We propose the hypothesize that expected collision rate scales linearly with sample size. METHODS: The hypothesized relationship between collision rate and sample size was validated on Datavant’s Mortality dataset, which contains over 100 million unique Datavant Tokens (Social Security Number + First Name) on a patient-level basis and is based on data from US government sources. A series of random samples were drawn from the dataset and the number of unique individuals (N), unique Tokens (K), and collisions (C) and the collision rate (C/N) were computed at each sample size. Expected vs. observed collision rate were compared. This same analysis was repeated to examine the collision of combinations of Token 1 (Last Name + First Initial of First Name + Gender + Date of Birth) and Token 2 (Last Name (soundex) + First Name (soundex) + Gender + Date of Birth). RESULTS: The sample size threshold to have at least 1 expected collision was found to be between N = 40,000 and N = 70,000. We observed that as sample size increased, the number of collisions and collision rate correspondingly increased. When Token 1 and Token 2 were used together, the resulting distinct combinations of PII were higher compared with when using Token 2 individually (5.65 billion vs. 1.845 billion using 100% of the dataset sample), and the collision rate was substantially lower. CONCLUSIONS: The results indicate that collision rate scales about linearly with sample size. This validation helps to further inform how the false positive rate of token-based matching algorithms may change with sample size.
Conference/Value in Health Info
2021-11, ISPOR Europe 2021, Copenhagen, Denmark
Value in Health, Volume 24, Issue 12, S2 (December 2021)
Code
POSB321
Topic
Methodological & Statistical Research
Disease
No Specific Disease