REPRODUCIBILITY, AUDITABILITY, AND THE PROMPT DEPENDENCY PROBLEM: AN ASSESSMENT OF METHODOLOGICAL MATURITY OF THE AI EVIDENCE BASE UNDERPINNING RAISE
Author(s)
Angeline Babitha Dhas, BS1, Viji Queen V, Sr., PharmD2, Revanth M, B.E.3, Diwyashri Govindarajaperumal, B.Pharm3, Aditi Bajpai, PharmD1, Meghan Oates-Zalesky, MSc4.
1MadeAi, Cambridge, MA, USA, 2MadeAI, Nagercoil, India, 3MadeAi, Nagercoil, India, 4Chief Marketing Officer, MadeAi, Cambridge, MA, USA.
1MadeAi, Cambridge, MA, USA, 2MadeAI, Nagercoil, India, 3MadeAi, Nagercoil, India, 4Chief Marketing Officer, MadeAi, Cambridge, MA, USA.
OBJECTIVES: To characterize the AI-system composition of RAISE's cited reference base; evaluate the transparency and reproducibility of foundational large language model (LLM) evaluations within it; and determine whether the current evidence base meets the standards required for regulatory-grade evidence generation.
METHODS: All 177 references cited across the RAISE papers were classified by AI-system type: foundational/general-purpose LLMs, LLM-based commercial platforms, purpose-built ML/NLP tools, classical ML methods, and non-AI documents. The 30 unique foundational-LLM evaluations were assessed across eight transparency and reproducibility dimensions: full-prompt disclosure; temperature/decoding parameter reporting; run-to-run consistency testing; open code/data availability; use of a human or benchmark reference standard; pre-registration; author ML/NLP expertise; and study maturity.
RESULTS: Of 86 references actively evaluating AI systems, 37 (43%) involve prompt-dependent foundational models—the category most susceptible to reproducibility concerns. Notable strengths were identified: 23/30 (77%) employed a human-reference standard, and half used pre-registered protocols (mean transparency score 6.5/10). However, critical reproducibility indicators were frequently absent: only 7/30 (23%) reported temperature/decoding settings; 10/30 (33%) tested output consistency across repeated runs; and 8/30 (27%) shared code and data openly. Eleven studies (37%) were preprints or basic-capability probes not subject to full peer review. Commercial platform evaluations were among the least transparent, with minimal disclosure of model versions, prompt configurations, or underlying technical infrastructure.
CONCLUSIONS: The AI evidence base underpinning RAISE is promising but does not yet consistently meet the reproducibility and auditability standards required for regulatory-grade evidence generation. We therefore call on the broader research community to adopt minimum reporting standards for LLM-based evaluations—including full prompt and model-version disclosure, temperature reporting, repeated-run consistency testing, data sharing, and pre-registered protocols—to accelerate evidence maturation and enable frameworks such as RAISE to produce stronger, more defensible, and ultimately more trustworthy evidence-based guidance across a range of clinical and regulatory decision-making contexts.
METHODS: All 177 references cited across the RAISE papers were classified by AI-system type: foundational/general-purpose LLMs, LLM-based commercial platforms, purpose-built ML/NLP tools, classical ML methods, and non-AI documents. The 30 unique foundational-LLM evaluations were assessed across eight transparency and reproducibility dimensions: full-prompt disclosure; temperature/decoding parameter reporting; run-to-run consistency testing; open code/data availability; use of a human or benchmark reference standard; pre-registration; author ML/NLP expertise; and study maturity.
RESULTS: Of 86 references actively evaluating AI systems, 37 (43%) involve prompt-dependent foundational models—the category most susceptible to reproducibility concerns. Notable strengths were identified: 23/30 (77%) employed a human-reference standard, and half used pre-registered protocols (mean transparency score 6.5/10). However, critical reproducibility indicators were frequently absent: only 7/30 (23%) reported temperature/decoding settings; 10/30 (33%) tested output consistency across repeated runs; and 8/30 (27%) shared code and data openly. Eleven studies (37%) were preprints or basic-capability probes not subject to full peer review. Commercial platform evaluations were among the least transparent, with minimal disclosure of model versions, prompt configurations, or underlying technical infrastructure.
CONCLUSIONS: The AI evidence base underpinning RAISE is promising but does not yet consistently meet the reproducibility and auditability standards required for regulatory-grade evidence generation. We therefore call on the broader research community to adopt minimum reporting standards for LLM-based evaluations—including full prompt and model-version disclosure, temperature reporting, repeated-run consistency testing, data sharing, and pre-registered protocols—to accelerate evidence maturation and enable frameworks such as RAISE to produce stronger, more defensible, and ultimately more trustworthy evidence-based guidance across a range of clinical and regulatory decision-making contexts.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR115
Topic
Health Technology Assessment, Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas