DO TRADITIONAL SYSTEMATIC LITERATURE REVIEW (SLR) METHODS MEET THE STANDARDS OF TRANSPARENCY AND INCREASING AI-ASSISTED SCREENING? A CASE STUDY COMPARING INDIVIDUAL VS. HIERARCHICAL POPULATION, INTERVENTION/COMPARATOR, OUTCOMES, STUDY DESIGN (PICOS...
Author(s)
Rozee Liu, MSc1, Stacy Grieve, PhD2, Mihaela Musat, PhD2, Anna Forsythe, MBA, MSc, PharmD2.
1Senior director, Product and Business, Oncoscope AI, Miamia, FL, USA, 2Oncoscope, Miami, FL, USA.
1Senior director, Product and Business, Oncoscope AI, Miamia, FL, USA, 2Oncoscope, Miami, FL, USA.
OBJECTIVES: SLR guidelines recommend hierarchical assessment of PICOS criteria. In practice, reviewers frequently exclude records based on a single criterion despite multiple PICOS criteria being unmet. This approach may reduce transparency, complicate scope changes, and limit efficient AI-assisted evidence synthesis. We evaluated whether independent PICOS-level annotation improves adaptability, reviewer consistency, and AI training efficiency.
METHODS: Traditional hierarchical screening workflows were compared with a Real-Time AI-assisted systematic literature review (REAL-SLR) framework that independently annotate decisions for each PICOS criterion. A prostate cancer (PC) scope change re-introducing active surveillance served as a case study. Outcomes included inter-reviewer discordance, re-review effort following scope change, and AI training requirements.
RESULTS: Among 9,302 prostate cancer records, 10-70% were assigned different exclusion reasons by reviewers despite reaching the same exclusion decision. This discordance was substantially reduced with independent PICOS annotation, with discordance rates of 0.4%, 1.25%, 4.5%, and 1.7% for population, intervention/comparator, outcomes, and study design, respectively. Individual PICOS annotation also reduced AI training effort. Initial AI training achieved accuracies of 99.64%, 92.34%, 86.86%, and 94.53% for each PICOS criterion, respectively. Preserving individual PICOS annotations allowed subsequent training iterations to focus only on lower-performing criteria, requiring 3, 3, and 2 iterations, respectively, to achieve accuracies >95%. When active surveillance was added as an intervention, independent PICOS annotation reduced re-review effort by 98.75% compared with traditional hierarchical screening, as only 0.7% of records were excluded by intervention while meeting all remaining PICOS criteria.
CONCLUSIONS: Preserving independent PICOS-level decisions improves transparency, reproducibility, and adaptability while substantially reducing scope-change burden. With the increasing use of AI in PICO review, the current hierarchical screening process may not be sufficient to ensure that accuracy and transparency are maintained, suggesting current SLR guidelines should consider individual PICOS annotation to support living evidence generation, AI-assisted review, and evolving HTA evidence requirements.
METHODS: Traditional hierarchical screening workflows were compared with a Real-Time AI-assisted systematic literature review (REAL-SLR) framework that independently annotate decisions for each PICOS criterion. A prostate cancer (PC) scope change re-introducing active surveillance served as a case study. Outcomes included inter-reviewer discordance, re-review effort following scope change, and AI training requirements.
RESULTS: Among 9,302 prostate cancer records, 10-70% were assigned different exclusion reasons by reviewers despite reaching the same exclusion decision. This discordance was substantially reduced with independent PICOS annotation, with discordance rates of 0.4%, 1.25%, 4.5%, and 1.7% for population, intervention/comparator, outcomes, and study design, respectively. Individual PICOS annotation also reduced AI training effort. Initial AI training achieved accuracies of 99.64%, 92.34%, 86.86%, and 94.53% for each PICOS criterion, respectively. Preserving individual PICOS annotations allowed subsequent training iterations to focus only on lower-performing criteria, requiring 3, 3, and 2 iterations, respectively, to achieve accuracies >95%. When active surveillance was added as an intervention, independent PICOS annotation reduced re-review effort by 98.75% compared with traditional hierarchical screening, as only 0.7% of records were excluded by intervention while meeting all remaining PICOS criteria.
CONCLUSIONS: Preserving independent PICOS-level decisions improves transparency, reproducibility, and adaptability while substantially reducing scope-change burden. With the increasing use of AI in PICO review, the current hierarchical screening process may not be sufficient to ensure that accuracy and transparency are maintained, suggesting current SLR guidelines should consider individual PICOS annotation to support living evidence generation, AI-assisted review, and evolving HTA evidence requirements.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
SA52
Topic
Study Approaches
Topic Subcategory
Literature Review & Synthesis
Disease
Oncology