AI IN HEALTH ECONOMICS AND OUTCOMES RESEARCH SUFFERS FROM LOW EXPLAINABILITY, TRACEABILITY, REPRODUCIBILITY AND FROM HALLUCINATIONS: A TARGETED REVIEW OF CHALLENGES AND TECHNICAL MITIGATIONS
Author(s)
Elias Altrabsheh, MSc1, Oliver Peter Whitaker, PhD1, Robert Görke, PhD2.
1d-fine, London, United Kingdom, 2d-fine, Frankfurt am Main, Germany.
1d-fine, London, United Kingdom, 2d-fine, Frankfurt am Main, Germany.
OBJECTIVES: Adoption of artificial intelligence (AI) and machine learning in health economics and outcomes research (HEOR) is constrained by four recurring concerns: explainability, hallucinations, traceability, and reproducibility. We characterised how often each concern is raised in the HEOR-AI literature and mapped the technical approaches proposed to address them.
METHODS: A targeted, purposive review identified HEOR-relevant sources published January 2020 to June 2026, comprising peer-reviewed articles, ISPOR and Value in Health outputs (including working-group reports and conference abstracts), and evidence-synthesis methods literature applicable to HEOR. Each source was coded against the four challenges and against five mitigation categories: knowledge graphs/GraphRAG, retrieval-augmented generation (RAG), human-in-the-loop review, provenance/audit frameworks, and documentation/reporting standards. A problem-by-technique matrix quantified co-occurrence.
RESULTS: Twenty-four sources met criteria. Explainability and hallucinations were the most frequently discussed challenges (each 50%), followed by reproducibility (46%) and traceability (33%). Among mitigations, human-in-the-loop review was most common (58%), then documentation/reporting standards (38%), RAG (25%), provenance/audit frameworks (21%), and knowledge graphs/GraphRAG (8%). The matrix showed human-in-the-loop review spanning all four challenges, most densely for reproducibility. Reporting standards clustered around explainability and reproducibility; RAG clustered around hallucinations; provenance frameworks aligned with traceability. Knowledge graphs were rare and absent from the reproducibility cell. Formal quantitative evaluation of any mitigation was uncommon; most sources described approaches conceptually rather than benchmarking them.
CONCLUSIONS: HEOR discourse on trustworthy AI concentrates on explainability and hallucinations, with reproducibility comparatively under-operationalised and traceability least discussed. Human oversight is the default safeguard across all challenges, while retrieval and provenance methods target hallucinations and traceability specifically. For practice, validated grounding and audit tooling is the near-term priority; standardised reproducibility evidence remains the principal gap. Whilst LLMs are non-deterministic by nature, and as such there is no ‘solution’ to these challenges, using well selected technical implementations can enhance traceability and isolate areas of non-deterministic behaviors for human intervention
METHODS: A targeted, purposive review identified HEOR-relevant sources published January 2020 to June 2026, comprising peer-reviewed articles, ISPOR and Value in Health outputs (including working-group reports and conference abstracts), and evidence-synthesis methods literature applicable to HEOR. Each source was coded against the four challenges and against five mitigation categories: knowledge graphs/GraphRAG, retrieval-augmented generation (RAG), human-in-the-loop review, provenance/audit frameworks, and documentation/reporting standards. A problem-by-technique matrix quantified co-occurrence.
RESULTS: Twenty-four sources met criteria. Explainability and hallucinations were the most frequently discussed challenges (each 50%), followed by reproducibility (46%) and traceability (33%). Among mitigations, human-in-the-loop review was most common (58%), then documentation/reporting standards (38%), RAG (25%), provenance/audit frameworks (21%), and knowledge graphs/GraphRAG (8%). The matrix showed human-in-the-loop review spanning all four challenges, most densely for reproducibility. Reporting standards clustered around explainability and reproducibility; RAG clustered around hallucinations; provenance frameworks aligned with traceability. Knowledge graphs were rare and absent from the reproducibility cell. Formal quantitative evaluation of any mitigation was uncommon; most sources described approaches conceptually rather than benchmarking them.
CONCLUSIONS: HEOR discourse on trustworthy AI concentrates on explainability and hallucinations, with reproducibility comparatively under-operationalised and traceability least discussed. Human oversight is the default safeguard across all challenges, while retrieval and provenance methods target hallucinations and traceability specifically. For practice, validated grounding and audit tooling is the near-term priority; standardised reproducibility evidence remains the principal gap. Whilst LLMs are non-deterministic by nature, and as such there is no ‘solution’ to these challenges, using well selected technical implementations can enhance traceability and isolate areas of non-deterministic behaviors for human intervention
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR163
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas