IDENTICAL OUTPUTS, DIVERGENT DEFENSIBILITY: LLM INTERACTION STYLE IN HEALTH-ECONOMIC MODELING
Author(s)
Benjamin White1, Anshul Shah, B.pharm, MSc2, Sandra Milev, MSc3.
1Director, HEOR, Red Nucleus, Yardley, PA, USA, 2Red Nucleus, yardley, PA, USA, 3Red Nucleus, Yardley, PA, USA.
1Director, HEOR, Red Nucleus, Yardley, PA, USA, 2Red Nucleus, yardley, PA, USA, 3Red Nucleus, Yardley, PA, USA.
OBJECTIVES: To evaluate whether LLM-assisted economic model development induces "cognitive debt": a state where a model runs but cannot be fully explained by the analyst or evidenced in code.
METHODS: Single-analyst case study using a validated Excel® cost-effectiveness model in a metabolic disease: a Markov model with 10 alive health states. It was rebuilt twice as an R package from identical specifications and inputs, varying only interaction style: 1) broad single-pass delegation (SIMPLE) vs. 2) modular, analyst-specified and analyst-owned components (MODULAR). Two outcomes were scored by independent panels of LLM graders (N = 10 per panel, per model): code quality (weighted dimensions) and analyst understanding (code-closed quiz).
RESULTS: Both builds agreed within 1% on incremental cost-effectiveness and differed by 5-8% from the Excel reference (likely due to implementation decisions unrelated to interaction style). Code-quality composites were high for both (99.5% MODULAR vs 86.9% SIMPLE, 0-1 scale), tying on numerical correctness dimensions. Analyst understanding diverged (82.5% MODULAR vs 66% SIMPLE); in SIMPLE, the analyst answered from specification and theory rather than implementation, restating the specifications 7 times (vs 0) and overclaiming (asserting unverified implementation behaviour) 14 times (vs. 1). Both LLM assessments located the cognitive debt in the implementation decisions left open by the specifications.
CONCLUSIONS: Both models produced nearly identical results and were well scored on numerical correctness. However, the incremental, modular build resulted in substantially higher analyst ownership, buying defensibility. The single-pass delegation model is useful for verifying results but not as a model that a team defends. Review should target the implementation itself (e.g., does it enforce the model's invariants such as probabilities summing to one, carry tests, trace to sources, and match its documentation), not just confirm the result reproduces. Findings are illustrative (single-analyst; analyst not blinded to arm; LLM graders measuring LLM-induced debt).
METHODS: Single-analyst case study using a validated Excel® cost-effectiveness model in a metabolic disease: a Markov model with 10 alive health states. It was rebuilt twice as an R package from identical specifications and inputs, varying only interaction style: 1) broad single-pass delegation (SIMPLE) vs. 2) modular, analyst-specified and analyst-owned components (MODULAR). Two outcomes were scored by independent panels of LLM graders (N = 10 per panel, per model): code quality (weighted dimensions) and analyst understanding (code-closed quiz).
RESULTS: Both builds agreed within 1% on incremental cost-effectiveness and differed by 5-8% from the Excel reference (likely due to implementation decisions unrelated to interaction style). Code-quality composites were high for both (99.5% MODULAR vs 86.9% SIMPLE, 0-1 scale), tying on numerical correctness dimensions. Analyst understanding diverged (82.5% MODULAR vs 66% SIMPLE); in SIMPLE, the analyst answered from specification and theory rather than implementation, restating the specifications 7 times (vs 0) and overclaiming (asserting unverified implementation behaviour) 14 times (vs. 1). Both LLM assessments located the cognitive debt in the implementation decisions left open by the specifications.
CONCLUSIONS: Both models produced nearly identical results and were well scored on numerical correctness. However, the incremental, modular build resulted in substantially higher analyst ownership, buying defensibility. The single-pass delegation model is useful for verifying results but not as a model that a team defends. Review should target the implementation itself (e.g., does it enforce the model's invariants such as probabilities summing to one, carry tests, trace to sources, and match its documentation), not just confirm the result reproduces. Findings are illustrative (single-analyst; analyst not blinded to arm; LLM graders measuring LLM-induced debt).
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
EE169
Topic
Economic Evaluation, Health Technology Assessment, Methodological & Statistical Research
Disease
Cardiovascular Disorders (including MI, Stroke, Circulatory), Diabetes/Endocrine/Metabolic Disorders (including obesity)