Evaluation of Artificial Intelligence within DistillerSR Software As a Second Reviewer for a Systematic Literature Review

Author(s)

Patel A1, Hallock C2, Fusco N2, Cadarette SM2, Mody L1
1Xcenda, L.L.C., Tampa, FL, USA, 2Xcenda, L.L.C., Palm Harbor, FL, USA

OBJECTIVES: To evaluate the performance of artificial intelligence (AI) as a second reviewer within the DistillerSR platform for a systematic literature review (SLR).

METHODS: Originally, a total of 2,613 references identified in the SLR were assessed by 2 independent human reviewers; a third analyst resolved conflicts. For this evaluation, the AI acted as a second reviewer in the screening process (2 possible outcomes: include or exclude) across 3 single-reviewer manually screened reference groups from the original total references (n=300, 400, and 500). Results were analyzed for 3 training and test proportions in each set (% AI train/% AI test: 5%/80%, 20%/65%, and 80%/20%). The AI screening results of the remainder of the 2,613 references were then compared with the original screened references. The accuracy, sensitivity, and specificity were calculated and compared.

RESULTS: The accuracy of AI-reviewed references increased with increasing number of manually screened references for the training sets. Overall, the AI screening accuracy was consistently lower using 300 vs 500 manually screened references across the 3 training sets (5%/80%, 93.3% vs 94.6%; 20%/65%, 93.2% vs 94.3%; 80%/20%, 92.5% vs 94.3%). The sensitivity also improved with increasing numbers of manually screened references and with a higher percentage of references used for training (300 references, range: 0.27 to 0.35 vs 500 references, range: 0.31 to 0.37). The specificity was high across all 3 training sets and manually screened reference sets (range, 0.97 to 0.99).

CONCLUSIONS: Although the accuracy and specificity of the included and excluded references were high, the sensitivity was low across all training sets. AI within DistillerSR may be useful in streamlining the reference-screening process by prioritizing likely inclusions and providing an additional level of security by verifying exclusions; however, more research is needed before substituting AI as a second reviewer.

Conference/Value in Health Info

2021-05, ISPOR 2021, Montreal, Canada

Value in Health, Volume 24, Issue 5, S1 (May 2021)

Code

PNS101

Topic

Methodological & Statistical Research

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Specific Disease

Explore Related HEOR by Topic


Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×