ENGINEERING PERSISTENT CONTEXT FOR REPRODUCIBLE HEOR AI TOOLS
Author(s)
Hanan Irfan, MSc1, Tushar Srivastava, MSc1, Kunal Swami, MSc2.
1ConnectHEOR, London, United Kingdom, 2Connectheor Ltd, Delhi, India.
1ConnectHEOR, London, United Kingdom, 2Connectheor Ltd, Delhi, India.
OBJECTIVES: As generative AI tools are embedded into HEOR, a credibility problem emerges: the same task, run twice, can yield different outputs. We aimed to characterise why HEOR AI outputs drift between runs and to derive a framework for engineering persistent context so outputs are reproducible enough for regulated decision-making.
METHODS: We conducted a structured analysis of two HEOR AI tools developed in house, a cost-effectiveness model analyser and an automated report writer, decomposing each pipeline into context-assembly stages (ingestion, retrieval, prompt and skill assembly, inference, post-processing). At each stage we mapped mechanisms introducing run-to-run variance, distinguishing context-level drift (different information assembled across runs) from inference-level non-determinism (variation at fixed input). Failure modes were triangulated against context-engineering literature, retrieval-reproducibility evidence, and reproducibility practices in AI-mature domains. Reproducibility was operationalised through measurable constructs: retrieved-context overlap across runs and output agreement across runs (semantic and numeric concordance). Findings were synthesised into a persistent-context framework for HEOR and HTA.
RESULTS: The framework comprises five dimensions: (1) context provenance and versioning, pinning every source document, parameter file, prompt, and skill to immutable, hashed versions; (2) deterministic retrieval and assembly, fixing chunking, embedding model, ranking, and ordering so assembled context is stable; (3) context-state persistence, re-instantiating a stored context rather than rebuilding it, guarding against context collapse; (4) inference reproducibility controls, fixing decoding parameters and model snapshots while acknowledging residual system-level non-determinism; and (5) reproducibility verification and governance gates, quantifying context overlap and output agreement against thresholds stratified by decisional criticality, with drift monitoring and human sign-off.
CONCLUSIONS: Run-to-run reproducibility in HEOR AI is governed less by the model than by the stability of the assembled context. Engineering persistent context, verified and governed, reframes AI tooling as a reproducible analytical lifecycle fit for HTA scrutiny.
METHODS: We conducted a structured analysis of two HEOR AI tools developed in house, a cost-effectiveness model analyser and an automated report writer, decomposing each pipeline into context-assembly stages (ingestion, retrieval, prompt and skill assembly, inference, post-processing). At each stage we mapped mechanisms introducing run-to-run variance, distinguishing context-level drift (different information assembled across runs) from inference-level non-determinism (variation at fixed input). Failure modes were triangulated against context-engineering literature, retrieval-reproducibility evidence, and reproducibility practices in AI-mature domains. Reproducibility was operationalised through measurable constructs: retrieved-context overlap across runs and output agreement across runs (semantic and numeric concordance). Findings were synthesised into a persistent-context framework for HEOR and HTA.
RESULTS: The framework comprises five dimensions: (1) context provenance and versioning, pinning every source document, parameter file, prompt, and skill to immutable, hashed versions; (2) deterministic retrieval and assembly, fixing chunking, embedding model, ranking, and ordering so assembled context is stable; (3) context-state persistence, re-instantiating a stored context rather than rebuilding it, guarding against context collapse; (4) inference reproducibility controls, fixing decoding parameters and model snapshots while acknowledging residual system-level non-determinism; and (5) reproducibility verification and governance gates, quantifying context overlap and output agreement against thresholds stratified by decisional criticality, with drift monitoring and human sign-off.
CONCLUSIONS: Run-to-run reproducibility in HEOR AI is governed less by the model than by the stability of the assembled context. Engineering persistent context, verified and governed, reframes AI tooling as a reproducible analytical lifecycle fit for HTA scrutiny.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR299
Topic
Methodological & Statistical Research, Real World Data & Information Systems, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas