WHERE MUST THE HUMAN REMAIN IN THE LOOP? EVALUATING EXPERT-VALIDATION CHECKPOINTS IN AI-ASSISTED HTA EVIDENCE WORKFLOWS
Author(s)
Ankit Dhaundiyal, MS1, Jagriti Prasad, MPH2, Tejesh S, MPH2, Mona Thangamma, MSc Health Economics2, Megha Tharad, PhD3, Molshree Pandey, MBA4.
1Evalueserve GmbH, Rheinbach, Germany, 2Evalueserve Pvt Ltd, Bengaluru, India, 3Evalueserve SEZ Pvt Ltd, Gurugram, India, 4Evalueserve, Singapore, Singapore.
1Evalueserve GmbH, Rheinbach, Germany, 2Evalueserve Pvt Ltd, Bengaluru, India, 3Evalueserve SEZ Pvt Ltd, Gurugram, India, 4Evalueserve, Singapore, Singapore.
OBJECTIVES: To evaluate how the placement of expert checkpoints affects residual error survival in AI-assisted HTA evidence workflows, and whether stage-gated review offers stronger governance than AI-only or final-output review.
METHODS: An exploratory controlled checkpoint evaluation was conducted across 10 oncology product-indication scenarios using 100 intermediate outputs from an AI-assisted HTA evidence workflow. Outputs covered PICOTS framing, evidence extraction, comparator/endpoint assessment, traceability/citation verification, and final materiality review. Each output was classified by expert review as correct or erroneous and assigned severity using a structured rubric. Prespecified catch assumptions were used to determine which checkpoint could reasonably identify each error type. Four strategies were compared: AI-only; final-review only; midpoint review after evidence extraction and comparator/endpoint alignment; and stage-gated review at decision-critical steps. Outcomes included residual errors, material error survival, and percentage error reduction versus AI-only.
RESULTS: Across 100 outputs, 60 were correct and 40 erroneous, including 14 material, 21 moderate, and 5 minor errors. In the AI-only strategy, all 40 errors survived. Final-review reduced residual errors to 24 and material errors to 1, corresponding to 40.0% total error reduction. Midpoint review reduced residual errors to 14 and eliminated material error survival (65.0% reduction). Stage-gated review eliminated residual errors within this controlled dataset; it required more expert touchpoints but concentrated oversight at decision-critical stages. This supports a Human-in-the-Lead configuration, where experts guide progression at key evidence-decision points rather than reviewing final outputs alone.
CONCLUSIONS: Structured checkpointing reduced error survival in AI-assisted HTA workflows. Final-output review mitigated most material risk but allowed non-material errors to persist, whereas midpoint and stage-gated approaches prevented material error survival. Results suggest that a Human-in-the-Lead model, with expert control before dossier-critical outputs are finalised, may be more appropriate than passive final review alone. Prospective validation should assess dual adjudication, non-oncology settings, real-world timing, and the trade-off between expert effort and errors avoided.
METHODS: An exploratory controlled checkpoint evaluation was conducted across 10 oncology product-indication scenarios using 100 intermediate outputs from an AI-assisted HTA evidence workflow. Outputs covered PICOTS framing, evidence extraction, comparator/endpoint assessment, traceability/citation verification, and final materiality review. Each output was classified by expert review as correct or erroneous and assigned severity using a structured rubric. Prespecified catch assumptions were used to determine which checkpoint could reasonably identify each error type. Four strategies were compared: AI-only; final-review only; midpoint review after evidence extraction and comparator/endpoint alignment; and stage-gated review at decision-critical steps. Outcomes included residual errors, material error survival, and percentage error reduction versus AI-only.
RESULTS: Across 100 outputs, 60 were correct and 40 erroneous, including 14 material, 21 moderate, and 5 minor errors. In the AI-only strategy, all 40 errors survived. Final-review reduced residual errors to 24 and material errors to 1, corresponding to 40.0% total error reduction. Midpoint review reduced residual errors to 14 and eliminated material error survival (65.0% reduction). Stage-gated review eliminated residual errors within this controlled dataset; it required more expert touchpoints but concentrated oversight at decision-critical stages. This supports a Human-in-the-Lead configuration, where experts guide progression at key evidence-decision points rather than reviewing final outputs alone.
CONCLUSIONS: Structured checkpointing reduced error survival in AI-assisted HTA workflows. Final-output review mitigated most material risk but allowed non-material errors to persist, whereas midpoint and stage-gated approaches prevented material error survival. Results suggest that a Human-in-the-Lead model, with expert control before dossier-critical outputs are finalised, may be more appropriate than passive final review alone. Prospective validation should assess dual adjudication, non-oncology settings, real-world timing, and the trade-off between expert effort and errors avoided.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
PT11
Topic
Health Technology Assessment, Methodological & Statistical Research, Organizational Practices
Topic Subcategory
Industry
Disease
No Additional Disease & Conditions/Specialized Treatment Areas