CONFIGURATION OVER CONTENT: HOW LARGE LANGUAGE MODEL SETTINGS, NOT QUESTION COMPLEXITY, GOVERN REFERENCE FABRICATION IN NARRATIVE-REVIEW-STYLE OUTPUTS
Author(s)
Matthew Campbell, PhD, Kevin Kallmes, MA, JD, Allie Cichewicz, MSc.
Nested Knowledge, St. Paul, MN, USA.
Nested Knowledge, St. Paul, MN, USA.
OBJECTIVES: Large language models (LLMs) are increasingly used to generate evidence summaries, but they often read as authoritative while resting on fabricated references. It remains unclear whether reference fabrication is driven mainly by the complexity of the question asked or by tool configuration. This exploratory audit tested whether question specificity or LLM configuration better explained reference fabrication in narrative-review-style outputs.
METHODS: One commercially available and widely used LLM was prompted to write a "comprehensive narrative review" answering a disease prevalence question at three levels of specificity: a broad national type 2 diabetes population, a South Asian female subgroup, and a narrow age band within that subgroup. The same prompt was run under five configurations: four paid-tier configurations varying reasoning ("thinking") mode and web search, plus one free-tier configuration. Each cited reference was checked against indexed literature and classified as real and accurate, real but degraded, or fabricated. Reference volume was determined by the model and was not capped.
RESULTS: Across 152 audited references, question specificity did not predict fabrication, with rates remaining flat or falling as questions narrowed; configuration was the determining factor. Fabrication occurred in 1% of references when reasoning was enabled (1/71), rising to 39% when reasoning was disabled (21/54), while the free-tier configuration fell between these extremes at 19% (5/27). Enabling web search did not reduce fabrication in non-reasoning outputs. Across all configurations, the model self-limited the number of references it produced despite being prompted for a comprehensive narrative review.
CONCLUSIONS: In this exploratory audit, LLM configuration mattered more than question complexity. Reasoning separated near-perfect referencing from frequent fabrication, while default, non-reasoning, and free-tier outputs carried substantially higher risk. These findings support LLM configuration disclosure and caution against unchecked LLM-generated references. Beyond disclosure, reliable evidence synthesis needs verified references and fit-for-purpose tools built around source-grounded retrieval.
METHODS: One commercially available and widely used LLM was prompted to write a "comprehensive narrative review" answering a disease prevalence question at three levels of specificity: a broad national type 2 diabetes population, a South Asian female subgroup, and a narrow age band within that subgroup. The same prompt was run under five configurations: four paid-tier configurations varying reasoning ("thinking") mode and web search, plus one free-tier configuration. Each cited reference was checked against indexed literature and classified as real and accurate, real but degraded, or fabricated. Reference volume was determined by the model and was not capped.
RESULTS: Across 152 audited references, question specificity did not predict fabrication, with rates remaining flat or falling as questions narrowed; configuration was the determining factor. Fabrication occurred in 1% of references when reasoning was enabled (1/71), rising to 39% when reasoning was disabled (21/54), while the free-tier configuration fell between these extremes at 19% (5/27). Enabling web search did not reduce fabrication in non-reasoning outputs. Across all configurations, the model self-limited the number of references it produced despite being prompted for a comprehensive narrative review.
CONCLUSIONS: In this exploratory audit, LLM configuration mattered more than question complexity. Reasoning separated near-perfect referencing from frequent fabrication, while default, non-reasoning, and free-tier outputs carried substantially higher risk. These findings support LLM configuration disclosure and caution against unchecked LLM-generated references. Beyond disclosure, reliable evidence synthesis needs verified references and fit-for-purpose tools built around source-grounded retrieval.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR288
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas