FIDELITY-BOUNDED SYNTHETIC COHORTS FOR TRIAL-DESIGN EXPLORATION: AN ENDPOINT-EFFICIENCY CASE STUDY IN PEDIATRIC RHINOPHARYNGITIS
Author(s)
Salma Barkaoui, PhD1, Mohammed BENNANI, PhD2, Hadhami Mejbri, Master3, Jerome Vetillard, PhD3.
1Qualees, PARIS, France, 2QUALEES, PARIS, France, 3Qualees, Paris, France.
1Qualees, PARIS, France, 2QUALEES, PARIS, France, 3Qualees, Paris, France.
OBJECTIVES: Pediatric trials often face recruitment constraints that limit sample size and evaluation of alternative endpoint strategies. Synthetic patient generation has been proposed for trial-design simulation, yet its ability to preserve design-relevant statistical properties is rarely validated against ground truth. Using a completed pediatric randomized controlled trial (RCT), we evaluated whether synthetic cohorts reproduce the original trial-design conclusions. We did not aim to establish treatment efficacy or increase evidentiary sample size.
METHODS: Reference data came from a multicenter, double-blind, placebo-controlled rhinopharyngitis trial (254 children, 6-15 years; 2:1 active). Four generators were compared using a pre-specified seven-metric framework assessing distributional fidelity, multivariable structure, treatment-group balance, and real-versus-synthetic distinguishability. The best-performing Gaussian Mixture Model (GMM) was retained. Bootstrap simulations (500 iterations; N=100-2,500) applied the trial ANCOVA (treatment, baseline score, centre, age) to five endpoints: ΔTNSS Day 5, ΔTNSS Day 3, AUC-TNSS, ΔSSCS, and Day-5 cure rate. Effect sizes, confidence intervals, power trajectories, and sample size for 80% power (N80) were estimated using observed and shrinkage-adjusted effects.
RESULTS: The GMM achieved the highest real-versus-synthetic indistinguishability (discriminator AUC=0.31) while preserving endpoint distributions, treatment allocation, and effect-size ranking. ΔTNSS Day 3 showed the largest treatment signal (Cohen's d=0.179; 95% CI −0.08 to +0.44) and lowest N80 (observed-effect N=880; shrinkage-adjusted N≈1,550). The original primary endpoint (ΔTNSS Day 5) showed a smaller, null-consistent effect (d=−0.040) and required N=2,080. Endpoint rankings remained concordant between real and synthetic cohorts. Treatment-group separation peaked early and diminished as symptom trajectories converged.
CONCLUSIONS: Using a trial with known ground truth, the GMM preserved design-relevant properties and reproduced the original trial-design conclusions, supporting synthetic cohorts as a bounded tool for endpoint selection and sample-size planning rather than a substitute for clinical evidence or statistical power. Earlier endpoint Day 3 appeared more statistically efficient than Day 5, requiring prospective confirmation. External validation of synthetic-derived N80 estimates remains necessary.
METHODS: Reference data came from a multicenter, double-blind, placebo-controlled rhinopharyngitis trial (254 children, 6-15 years; 2:1 active). Four generators were compared using a pre-specified seven-metric framework assessing distributional fidelity, multivariable structure, treatment-group balance, and real-versus-synthetic distinguishability. The best-performing Gaussian Mixture Model (GMM) was retained. Bootstrap simulations (500 iterations; N=100-2,500) applied the trial ANCOVA (treatment, baseline score, centre, age) to five endpoints: ΔTNSS Day 5, ΔTNSS Day 3, AUC-TNSS, ΔSSCS, and Day-5 cure rate. Effect sizes, confidence intervals, power trajectories, and sample size for 80% power (N80) were estimated using observed and shrinkage-adjusted effects.
RESULTS: The GMM achieved the highest real-versus-synthetic indistinguishability (discriminator AUC=0.31) while preserving endpoint distributions, treatment allocation, and effect-size ranking. ΔTNSS Day 3 showed the largest treatment signal (Cohen's d=0.179; 95% CI −0.08 to +0.44) and lowest N80 (observed-effect N=880; shrinkage-adjusted N≈1,550). The original primary endpoint (ΔTNSS Day 5) showed a smaller, null-consistent effect (d=−0.040) and required N=2,080. Endpoint rankings remained concordant between real and synthetic cohorts. Treatment-group separation peaked early and diminished as symptom trajectories converged.
CONCLUSIONS: Using a trial with known ground truth, the GMM preserved design-relevant properties and reproduced the original trial-design conclusions, supporting synthetic cohorts as a bounded tool for endpoint selection and sample-size planning rather than a substitute for clinical evidence or statistical power. Earlier endpoint Day 3 appeared more statistically efficient than Day 5, requiring prospective confirmation. External validation of synthetic-derived N80 estimates remains necessary.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR156
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
Pediatrics, Respiratory-Related Disorders (Allergy, Asthma, Smoking, Other Respiratory)