Meta-TAK, a Scalable Double-Clustering Method for Treatment Sequences Visualization: Case Study in Breast Cancer Using Claim DATA
Author(s)
Prodel M1, Laurent M2, De Oliveira H2, Lamarsalle L2, Vainchtock A3
1HEVA, LYON, 69, France, 2HEVA, Lyon, France, 3HEVA, LYON, France
OBJECTIVES. The study purpose was to cluster and extract patterns in treatment sequences of thousands of patients, from claim databases and using a data-science methodology. Studying treatment sequences is critical for guidelines enforcement and to define the best therapeutic strategies, for instance for patients treated for her2-positive early breast cancer (eBC). Still, time sequence modeling is highly combinatorial and requires robust algorithms to extract patterns. METHODS. From the French national hospital database (PMSI), we identified patients with eBC, undergoing a breast surgery, and with over one subsequent trastuzumab administration (N=2,477). The TAK algorithm (Time-sequence Analysis through K-clustering) is an unsupervised clustering method designed to represent treatment sequences. It results in a comprehensive image of these sequences. The presented work is an extension of TAK to ensure the computational scalability in the number of patients, referred as the k-meta-TAK. It is a double-clustering method, one in preliminary and one in the TAK. The preliminary clustering smartly samples the original dataset in a sub-cohort containing k% of the original cohort. The approach was assessed on its practical usability on thousands of patients (computation time) and on the quality of pattern representation (mean entropy value). RESULTS. The baseline of the TAK (full cohort) showed an entropy of 0.15 and a computation time of 48 seconds. The 50%-ratio meta-tak was the best compromise, with a very close entropy (0.16) and a much shorter computation time of 10 seconds (-79%). Regarding pattern extraction, qualitative performances were met and validated through the identification of the same 5 extracted patterns by medical experts. CONCLUSIONS. The meta-TAK was introduced and proved useful for analyzing and visualizing treatment sequences from claim data. It paves the way to scale-up the clustering method to large cohorts of tens of thousands of patients while maintaining a low entropy, which makes clinical interpretations still trustworthy.
Conference/Value in Health Info
2020-11, ISPOR Europe 2020, Milan, Italy
Value in Health, Volume 23, Issue S2 (December 2020)
Code
PCN273
Topic
Health Service Delivery & Process of Care, Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics, Disease Management
Disease
Drugs, Oncology