A MINIMUM PERFORMANCE-REPORTING STANDARD FOR AI IN SYSTEMATIC LITERATURE REVIEW IN ABSTRACTS PRESENTED IN ISPOR
Author(s)
Shilpi Swami, MSc, Hanan Irfan, MSc, Raju Gautam, PhD, Tushar Srivastava, MSc.
ConnectHEOR, London, United Kingdom.
ConnectHEOR, London, United Kingdom.
OBJECTIVES: AI-assisted literature review is becoming central to SLRs, yet difficult to compare or trust for HTA. We analysed recent ISPOR AI abstracts to quantify this reporting gap and define the minimum performance information to make AI-SLR screening claims auditable.
METHODS: We conducted a content analysis of 86 evidence-synthesis and SLR-automation abstracts from the ISPOR Philadelphia 2026 , the largest AI theme within 314 AI-related abstracts. Each abstract was coded for SLR sub-task, method family, reported metrics and values, reference standard, and evidence of error. Reporting completeness was assessed against the metrics required to judge screening under emerging HTA expectations.
RESULTS: Reporting was led by efficiency, not safety. Overall, 56% (48/86) of abstracts claimed time savings, ranging from approximately 40% to 500%, but used no common baseline. Accuracy was reported in 42% (36/86), while recall or sensitivity was reported in only 30% (26/86) and precision in 15% (13/86). Only 12% reported recall and precision together, and 49% (42/86) reported no core performance metric. Work Saved over Sampling, the standard screening-efficiency metric, was not reported in any abstract. Only 35% named a reference standard, often on incompatible bases, and only 10% audited missed includes. Reported performance ranges were too wide for comparison: sensitivity ranged from 10% to 100% and accuracy from 40% to 100%. One validation reported sensitivity 0.99 with accuracy 0.34 and precision 0.21, showing why accuracy alone can mislead under class imbalance. We derived a six-item minimum reporting standard requiring recall at a pre-specified threshold, WSS at that recall, precision and missed includes, reference standard, dataset prevalence, and screening stage.
CONCLUSIONS: Current AI-SLR screening evidence is not yet comparable enough for HTA decision-making. A concise itemised reporting standard, aligned with the logic of CONSORT and PRISMA, would turn AI screening claims into auditable evidence that reviewers can compare, challenge, and use.
METHODS: We conducted a content analysis of 86 evidence-synthesis and SLR-automation abstracts from the ISPOR Philadelphia 2026 , the largest AI theme within 314 AI-related abstracts. Each abstract was coded for SLR sub-task, method family, reported metrics and values, reference standard, and evidence of error. Reporting completeness was assessed against the metrics required to judge screening under emerging HTA expectations.
RESULTS: Reporting was led by efficiency, not safety. Overall, 56% (48/86) of abstracts claimed time savings, ranging from approximately 40% to 500%, but used no common baseline. Accuracy was reported in 42% (36/86), while recall or sensitivity was reported in only 30% (26/86) and precision in 15% (13/86). Only 12% reported recall and precision together, and 49% (42/86) reported no core performance metric. Work Saved over Sampling, the standard screening-efficiency metric, was not reported in any abstract. Only 35% named a reference standard, often on incompatible bases, and only 10% audited missed includes. Reported performance ranges were too wide for comparison: sensitivity ranged from 10% to 100% and accuracy from 40% to 100%. One validation reported sensitivity 0.99 with accuracy 0.34 and precision 0.21, showing why accuracy alone can mislead under class imbalance. We derived a six-item minimum reporting standard requiring recall at a pre-specified threshold, WSS at that recall, precision and missed includes, reference standard, dataset prevalence, and screening stage.
CONCLUSIONS: Current AI-SLR screening evidence is not yet comparable enough for HTA decision-making. A concise itemised reporting standard, aligned with the logic of CONSORT and PRISMA, would turn AI screening claims into auditable evidence that reviewers can compare, challenge, and use.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR231
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas