A GENERATIVE AI SUITE FOR REAL-WORLD EVIDENCE STUDY DOCUMENTATION: MULTI-MODULE VALIDATION ACROSS PHARMACOEPIDEMIOLOGICAL DELIVERABLES

Author(s)

Massoud Toussi, MBA, PhD, MD1, Motiur Rahman, MS, PhD2, Ayad Ali, PhD3, Laurie Jean Lambert, MPH, PhD4, Andrew Cooper, PhD5, Gorana Capkun, PhD6, Karl-Johan Myren, MSc7, François GAVINI, PhD8, Christopher M. Blanchette, MA, MBA, MSc, PhD9, Mats Rosenlund, PhD10.
1Evidence Mastery, Lailly en Val, France, 2US Food and Drug Administration, Silver Spring, MD, USA, 3BeOne Medicines, San Carlos, CA, USA, 4CADTH, Newington, ON, Canada, 5Shionogi, London, United Kingdom, 6Merck KGaA, Allschwil, Switzerland, 7Alexion, Stockholm, Sweden, 8Takeda, Zurich, France, 9Novo Nordisk, Doylestown, PA, USA, 10Daiichi Sankyo, Stokholm, Sweden.
OBJECTIVES: To compare the performance of EvidenceAi Suite, a generative AI platform comprising modules for generation of first draft of study documents, with human generation of a first draft of the same document.
METHODS: A prospective mixed-methods evaluation was conducted with RWE professionals across organizations, including regulators, HTA bodies and international pharmaceutical companies. Participants used one or more EvidenceAi Suite to generate or evaluate study documents. A ten-domain assessment survey was built inside the tool, based on the guidance by the World Health Organization (WHO) to compare the tool with humans: ease of use, time-to-render, editorial quality, scientific quality, error frequency, usefulness, workflow integrability, and overall satisfaction were rated on a 10-point comparative scale anchored against a human-generated equivalent. A score of 1 was defined as “significantly inferior to human” while a score of 10 was defined as “significantly superior to human”. Calendar time saved, human effort saved, and overall impression were collected as numerical and text values. Survey responses were analyzed using descriptive statistics.
RESULTS: So far, 39 users from 12 organizations used the Suite to generate 402 study documents in a pilot environment. The preliminary results on 15 survey responses show superior performance to humans in all domains. The mean score is 9.3 for ease of use (compared to explaining the task to a human), 9.3 for time to render, 8.6 for editorial quality, 8.3 for scientific quality, 9.2 for usefulness, 9.3 for workflow integration, and 9.5 for overall satisfaction. Human capital savings reached 58.7 hours per use across all module types. Per-module results will be presented during ISPOR.
CONCLUSIONS: The AI Suite demonstrated scientific and editorial quality exceeding human first-draft generation, while considerably reducing calendar time and human effort. These findings support the integration of purpose-built generative AI into RWE study design and reporting workflows.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

RWD21

Topic

Epidemiology & Public Health, Methodological & Statistical Research, Real World Data & Information Systems

Disease

Cardiovascular Disorders (including MI, Stroke, Circulatory), Diabetes/Endocrine/Metabolic Disorders (including obesity), Neurological Disorders, Oncology, Rare & Orphan Diseases

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×