DISTINGUISHING VERIFICATION FROM JUDGMENT IN HUMAN-IN-THE-LOOP EVIDENCE SYNTHESIS
Author(s)
Andrew Frederickson, MS.
Vice President, Precision AQ, Gladstone, NJ, USA.
Vice President, Precision AQ, Gladstone, NJ, USA.
OBJECTIVES: “Human-in-the-loop” is widely cited as a safeguard for AI-assisted evidence synthesis, and agreement with human reviewers (e.g., sensitivity/specificity) is treated as evidence of validity. Both practices assume human involvement is a single thing. It is not. Human input performs two functions: verification, checking whether an output matches a specified target; and judgment, choosing which target or tradeoff suits the decision problem. The objective was to distinguish these functions and identify where practice conflates them.
METHODS: A structured conceptual analysis was conducted across common evidence synthesis tasks where AI classifies, extracts, summarizes, or compares evidence. For each task, decision points were classified as verification against a pre-specified target or as requiring selection among competing targets, thresholds, or error tradeoffs. Decision points requiring tradeoff specification were grouped by decision-problem type.
RESULTS: Verification applies where the human assesses whether an output matches a pre-specified target, such as confirming accurate data extraction or that a study meets an eligibility criterion. Judgment applies where the preferred decision depends on the consequences of alternative errors or the intended use of the evidence: borderline screening, where sensitivity and specificity are weighted by review purpose; population-similarity assessment, where acceptable heterogeneity depends on the comparison; and outcome harmonization, where completeness and comparability imply different synthesis choices. In such cases, adding a human reviewer does not resolve the problem unless the decision rule and tradeoff are made explicit.
CONCLUSIONS: Human involvement in AI-assisted evidence synthesis is not a single category, yet it is reported as one. Distinguishing verification from judgment clarifies when a human is checking an output against a target and when the human is defining what counts as a better answer for the decision. Reporting of AI workflows, and “human-in-the-loop” claims especially, should specify which function each human role performs.
METHODS: A structured conceptual analysis was conducted across common evidence synthesis tasks where AI classifies, extracts, summarizes, or compares evidence. For each task, decision points were classified as verification against a pre-specified target or as requiring selection among competing targets, thresholds, or error tradeoffs. Decision points requiring tradeoff specification were grouped by decision-problem type.
RESULTS: Verification applies where the human assesses whether an output matches a pre-specified target, such as confirming accurate data extraction or that a study meets an eligibility criterion. Judgment applies where the preferred decision depends on the consequences of alternative errors or the intended use of the evidence: borderline screening, where sensitivity and specificity are weighted by review purpose; population-similarity assessment, where acceptable heterogeneity depends on the comparison; and outcome harmonization, where completeness and comparability imply different synthesis choices. In such cases, adding a human reviewer does not resolve the problem unless the decision rule and tradeoff are made explicit.
CONCLUSIONS: Human involvement in AI-assisted evidence synthesis is not a single category, yet it is reported as one. Distinguishing verification from judgment clarifies when a human is checking an output against a target and when the human is defining what counts as a better answer for the decision. Reporting of AI workflows, and “human-in-the-loop” claims especially, should specify which function each human role performs.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR212
Topic
Methodological & Statistical Research, Organizational Practices, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas