REPRODUCIBILITY, AUDITABILITY, AND THE PROMPT DEPENDENCY PROBLEM: AN ASSESSMENT OF METHODOLOGICAL MATURITY OF THE AI EVIDENCE BASE UNDERPINNING RAISE

Author(s)

Angeline Babitha Dhas, BS1, Viji Queen V, Sr., PharmD2, Revanth M, B.E.3, Diwyashri Govindarajaperumal, B.Pharm3, Aditi Bajpai, PharmD1, Meghan Oates-Zalesky, MSc4.
1MadeAi, Cambridge, MA, USA, 2MadeAI, Nagercoil, India, 3MadeAi, Nagercoil, India, 4Chief Marketing Officer, MadeAi, Cambridge, MA, USA.
OBJECTIVES: To characterize the AI-system composition of RAISE's cited reference base; evaluate the transparency and reproducibility of foundational large language model (LLM) evaluations within it; and determine whether the current evidence base meets the standards required for regulatory-grade evidence generation.
METHODS: All 177 references cited across the RAISE papers were classified by AI-system type: foundational/general-purpose LLMs, LLM-based commercial platforms, purpose-built ML/NLP tools, classical ML methods, and non-AI documents. The 30 unique foundational-LLM evaluations were assessed across eight transparency and reproducibility dimensions: full-prompt disclosure; temperature/decoding parameter reporting; run-to-run consistency testing; open code/data availability; use of a human or benchmark reference standard; pre-registration; author ML/NLP expertise; and study maturity.
RESULTS: Of 86 references actively evaluating AI systems, 37 (43%) involve prompt-dependent foundational models—the category most susceptible to reproducibility concerns. Notable strengths were identified: 23/30 (77%) employed a human-reference standard, and half used pre-registered protocols (mean transparency score 6.5/10). However, critical reproducibility indicators were frequently absent: only 7/30 (23%) reported temperature/decoding settings; 10/30 (33%) tested output consistency across repeated runs; and 8/30 (27%) shared code and data openly. Eleven studies (37%) were preprints or basic-capability probes not subject to full peer review. Commercial platform evaluations were among the least transparent, with minimal disclosure of model versions, prompt configurations, or underlying technical infrastructure.
CONCLUSIONS: The AI evidence base underpinning RAISE is promising but does not yet consistently meet the reproducibility and auditability standards required for regulatory-grade evidence generation. We therefore call on the broader research community to adopt minimum reporting standards for LLM-based evaluations—including full prompt and model-version disclosure, temperature reporting, repeated-run consistency testing, data sharing, and pre-registered protocols—to accelerate evidence maturation and enable frameworks such as RAISE to produce stronger, more defensible, and ultimately more trustworthy evidence-based guidance across a range of clinical and regulatory decision-making contexts.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR115

Topic

Health Technology Assessment, Methodological & Statistical Research, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×