L1-DISTANCE-BASED SIMILARITY FRAMEWORK FOR CLUSTERING IRREGULAR LONGITUDINAL DATA
Author(s)
TANMOY MAJUMDAR, Research Scholar.
Department of Mathematics & Computing, Indian Institute of Technology Dhanbad, Dhanbad, India.
Department of Mathematics & Computing, Indian Institute of Technology Dhanbad, Dhanbad, India.
OBJECTIVES: Longitudinal studies often produce data where subjects are measured at individual-specific time points, creating irregularly spaced observation schedules. This irregularity limits conventional clustering methods, which typically require a common measurement grid or rely on restrictive parametric assumptions about trajectory shape. This study aims to develop a flexible, non-parametric clustering framework for grouping subjects based on irregularly timed longitudinal measurements, without imposing any predefined functional form on the underlying trajectories.
METHODS: The proposed approach defines a dissimilarity measure between two subjects as the L1 area between their linearly interpolated response trajectories, computed only over the interval where both subjects are observed. This dissimilarity is transformed into a similarity score using a Gaussian kernel, producing a similarity matrix suitable for standard clustering algorithms, including spectral clustering and partitioning around medoids (PAM). The method accommodates fully subject-specific observation schedules and yields an interpretable similarity index equal to one for identical trajectories, decreasing smoothly as trajectories diverge.
RESULTS: The proposed method was compared against several established approaches, including functional principal component analysis (FPCA) with k-means, latent class mixed-effects models (LCMM), dynamic time warping (DTW) with k-medoids, and group-based trajectory modeling (GBTM). Across simulation scenarios and the real-data application, the proposed framework showed robust and competitive performance in recovering true underlying trajectory groups, frequently outperforming alternative methods, especially under conditions of substantial irregularity in observation timing.
CONCLUSIONS: The L1-distance-based similarity framework provides an effective, interpretable solution for clustering longitudinal data with irregular observation schedules. By avoiding restrictive parametric assumptions and accommodating subject-specific measurement times, the method offers a practical and broadly applicable alternative to existing trajectory-clustering approaches, with strong potential across diverse longitudinal data settings.
METHODS: The proposed approach defines a dissimilarity measure between two subjects as the L1 area between their linearly interpolated response trajectories, computed only over the interval where both subjects are observed. This dissimilarity is transformed into a similarity score using a Gaussian kernel, producing a similarity matrix suitable for standard clustering algorithms, including spectral clustering and partitioning around medoids (PAM). The method accommodates fully subject-specific observation schedules and yields an interpretable similarity index equal to one for identical trajectories, decreasing smoothly as trajectories diverge.
RESULTS: The proposed method was compared against several established approaches, including functional principal component analysis (FPCA) with k-means, latent class mixed-effects models (LCMM), dynamic time warping (DTW) with k-medoids, and group-based trajectory modeling (GBTM). Across simulation scenarios and the real-data application, the proposed framework showed robust and competitive performance in recovering true underlying trajectory groups, frequently outperforming alternative methods, especially under conditions of substantial irregularity in observation timing.
CONCLUSIONS: The L1-distance-based similarity framework provides an effective, interpretable solution for clustering longitudinal data with irregular observation schedules. By avoiding restrictive parametric assumptions and accommodating subject-specific measurement times, the method offers a practical and broadly applicable alternative to existing trajectory-clustering approaches, with strong potential across diverse longitudinal data settings.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR179
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas