PERFORMANCE OF NESTED KNOWLEDGE CRITERIA-BASED SCREENING AND ADAPTIVE SMART TAGGING AGAINST A MANUALLY CONDUCTED SYSTEMATIC LITERATURE REVIEW IN HEREDITARY ANGIOEDEMA
Author(s)
Alexandra M.E. Zuckermann, PhD, Amrita Debnath, MSc, Elizabeth M. Halloran, MSc, Sumeet Singh, MSc.
Value & Evidence, EVERSANA, Victoria, BC, Canada.
Value & Evidence, EVERSANA, Victoria, BC, Canada.
OBJECTIVES: AI-assisted tools offer the potential to improve the efficiency of systematic literature reviews (SLRs). This study aimed to validate two AI tools available on the Nested Knowledge platform, Criteria-based screening (CBS) and Adaptive Smart Tagging (AST) for data extraction, against a published SLR conducted manually.
METHODS: A published SLR of clinical trials evaluating prophylactic therapies for hereditary angioedema (HAE) served as the reference standard (PROSPERO, CRD42022359207). The original review was conducted without AI, with screening and extraction decisions retained internally by the authors. For CBS at abstract stage, the record set matched the original SLR. At full-text stage, the record set comprised all records for which full-text was originally sourced. AI decisions were compared against original human decisions to calculate performance metrics. For AST, all records originally included were assessed. AST was applied in high-fidelity mode to extract pre-specified elements, with outputs compared descriptively against the manually extracted dataset.
RESULTS: At abstract screening, CBS achieved 65% precision, 46% recall, 97% specificity, and 92% accuracy; notably, all false negatives at this stage were ultimately excluded from the published SLR, yielding a final decision-adjusted recall of 100%. At full-text screening, performance depended on the selected threshold: requiring all criteria yielded 74% recall, 44% precision, 73% specificity, and 73% accuracy, whereas requiring five of six criteria improved recall to 97%, with modest decreases in precision (39%) and specificity (56%). For data extraction, considering patient characteristics and clinical outcomes, AST was fully accurate for 64.9% of data elements (66.3% including partially accurate extractions). Some data elements were missed; further discrepancies were attributable to occasional mismatching of "not reported" and zero values, inconsistent unit conversions, and difficulty handling conference abstract booklets as well as aggregate or per-patient data formats.
CONCLUSIONS: Validated against a manually conducted SLR in HAE, the AI-assisted screening and extraction tools demonstrated promising performance.
METHODS: A published SLR of clinical trials evaluating prophylactic therapies for hereditary angioedema (HAE) served as the reference standard (PROSPERO, CRD42022359207). The original review was conducted without AI, with screening and extraction decisions retained internally by the authors. For CBS at abstract stage, the record set matched the original SLR. At full-text stage, the record set comprised all records for which full-text was originally sourced. AI decisions were compared against original human decisions to calculate performance metrics. For AST, all records originally included were assessed. AST was applied in high-fidelity mode to extract pre-specified elements, with outputs compared descriptively against the manually extracted dataset.
RESULTS: At abstract screening, CBS achieved 65% precision, 46% recall, 97% specificity, and 92% accuracy; notably, all false negatives at this stage were ultimately excluded from the published SLR, yielding a final decision-adjusted recall of 100%. At full-text screening, performance depended on the selected threshold: requiring all criteria yielded 74% recall, 44% precision, 73% specificity, and 73% accuracy, whereas requiring five of six criteria improved recall to 97%, with modest decreases in precision (39%) and specificity (56%). For data extraction, considering patient characteristics and clinical outcomes, AST was fully accurate for 64.9% of data elements (66.3% including partially accurate extractions). Some data elements were missed; further discrepancies were attributable to occasional mismatching of "not reported" and zero values, inconsistent unit conversions, and difficulty handling conference abstract booklets as well as aggregate or per-patient data formats.
CONCLUSIONS: Validated against a manually conducted SLR in HAE, the AI-assisted screening and extraction tools demonstrated promising performance.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR280
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas