MOCK EAG-STYLE REVIEWS USING GENERATIVE AI TO ANTICIPATE HTA CRITIQUE
Author(s)
Hanan Irfan, MSc, Tushar Srivastava, MSc, Thaison Tong, PhD, Shilpi Swami, MSc.
ConnectHEOR, London, United Kingdom.
ConnectHEOR, London, United Kingdom.
OBJECTIVES: EAGs appraise company submissions to NICE under tight timelines, yet generative AI risks displacing the judgement appraisal credibility. We developed and validated a scope-grounded LLM system that drafts NICE EAG-style critique reports and clarification questions.
METHODS: We built a multi-stage, retrieval-grounded pipeline on a reasoning LLM (GPT-5.4) that ingests the company submission and NICE final scope. Stage 1 produced a structured critique outline per domain from exemplar tables of contents. Stage 2 generated issue-based critique section by section under strict grounding rules: facts were confined to the current submission section. Stage 3 produced a numbered overview of key issues, with a parallel clarification-question module. The case study was the NICE appraisal of lorlatinib for untreated ALK-positive advanced NSCLC. The case study was chosen such that its release date was after the pretraining data cutoff of GPT 5.4 to ensure the LLM could not have been pretrained on this public data.
RESULTS: The tool recommended 19 issues out of which 12 matched with the actual EAG report that cited 14 issues in the original report (86%). Key issues includes immature overall survival, investigator-assessed progression after central review ceased, and a weak indirect comparison. Of 19 issues identified by tool, 16 (84%) were valid and 12 decision-critical as found in EAG report; three were false positives requiring human interpretation. No fabricated trials, comparators or values were identified; all traced to the submission. The clarification module produced 34 questions out of which 27 (79%) material and answerable as seen in EAG clarification questions as well. A first-draft critique was produced in under two hours.
CONCLUSIONS: A scope-grounded LLM system produced an auditable first-draft EAG critique that recovered most material issues without fabricating evidence. It supports, rather than replaces, expert EAG review, positioning AI as a governed, human-in-the-loop component of HTA appraisal.
METHODS: We built a multi-stage, retrieval-grounded pipeline on a reasoning LLM (GPT-5.4) that ingests the company submission and NICE final scope. Stage 1 produced a structured critique outline per domain from exemplar tables of contents. Stage 2 generated issue-based critique section by section under strict grounding rules: facts were confined to the current submission section. Stage 3 produced a numbered overview of key issues, with a parallel clarification-question module. The case study was the NICE appraisal of lorlatinib for untreated ALK-positive advanced NSCLC. The case study was chosen such that its release date was after the pretraining data cutoff of GPT 5.4 to ensure the LLM could not have been pretrained on this public data.
RESULTS: The tool recommended 19 issues out of which 12 matched with the actual EAG report that cited 14 issues in the original report (86%). Key issues includes immature overall survival, investigator-assessed progression after central review ceased, and a weak indirect comparison. Of 19 issues identified by tool, 16 (84%) were valid and 12 decision-critical as found in EAG report; three were false positives requiring human interpretation. No fabricated trials, comparators or values were identified; all traced to the submission. The clarification module produced 34 questions out of which 27 (79%) material and answerable as seen in EAG clarification questions as well. A first-draft critique was produced in under two hours.
CONCLUSIONS: A scope-grounded LLM system produced an auditable first-draft EAG critique that recovered most material issues without fabricating evidence. It supports, rather than replaces, expert EAG review, positioning AI as a governed, human-in-the-loop component of HTA appraisal.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR62
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas