AGENTIC AI FOR DATASET QUALITY ASSURANCE AND LATENT CLASS ANALYSIS: A HUMAN-IN-THE-LOOP METHODOLOGICAL FRAMEWORK APPLIED TO THE SWAN COHORT

Author(s)

Elif Inan Eroglu, PhD1, Motahhareh Nadimi, MSc2, Ann-Kathrin Frenz, MSc2, Sukriti Poddar, BSc, MS3, Melanie Tuchardt, MASc2, Nils Schoof, PhD4, Leonardo Dambrosi, MSc5.
1Berlin, Germany, 2Bayer AG, Berlin, Germany, 3Axtria, Inc, Berkley Heights, NJ, USA, 4Bayer AG, Wiesbaden, Germany, 5Bayer, Berlin, Germany.
OBJECTIVES: Longitudinal cohort studies generate complex, high-dimensional datasets that challenge conventional analytic workflows. We aim to develop and evaluate a methodological framework to harmonize agentic artificial intelligence(AI) with human oversight for dataset quality assurance and latent class analysis(LCA) in women's health research.
METHODS: The framework was applied to 3,306 women in the longitudinal Study of Women's Health Across the Nation (SWAN) cohort, using the first 10 publicly available follow-up visits and symptom measures spanning vasomotor symptoms, sleep disturbance, mood/well-being, cognition. The methodological focus was a six-step agentic AI pipeline:(1)analytic dataset creation by a human data analyst, (2)AI-driven data ingestion and profiling, (3)automated quality checks for missing values, duplicates, out-of-range values, distributional anomalies, inconsistent data types, and unexpected categorical values, with findings documented in reproducible reports (4)AI-guided feature preparation and recoding for LCA, (5)iterative LCA model fitting with automated class enumeration using BIC/AIC/entropy criteria, (6)auto-generated reports with visualizations and summaries. Human-in-the-loop checkpoints were embedded throughout to validate specifications, review outputs incrementally, and ensure domain-appropriate interpretation.
RESULTS: The AI workflow completed the full LCA pipeline, including diagnostics, in approximately 45 minutes and systematically generated transparent quality-control outputs, automatically adapted when models failed to converge by modifying starting values, and produced class profiles, plots, code, and narrative summaries. Human review remained essential for planning, protocol definition and stepwise validation, preventing overreliance on fully autonomous execution. The workflow identified different latent classes representing differing levels and domains of symptom burden, demonstrating that agentic AI can support rapid and reproducible methodological execution when bounded by clear human guardrails.
CONCLUSIONS: A human-in-the-loop agentic AI framework can strengthen the efficiency, reproducibility, and transparency of LCA workflows in complex epidemiologic datasets. This approach is especially valuable for methodologic use cases where scalable automation is needed, but scientific validity depends on structured human oversight at predefined decision points.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR154

Topic

Epidemiology & Public Health, Methodological & Statistical Research, Real World Data & Information Systems

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

Reproductive & Sexual Health

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×