VALIDATION OF DATA EXTRACTION BY AN AGENTIC AI LITERATURE REVIEW PLATFORM AGAINST MANUAL REVIEW
Author(s)
Shahrzad Salmasi, MSc, PhD1, James M. Gwinnutt, PhD2, Lenon Mendes Pereira, PhD3, Julien Heidt, BS, MS4, Niamh McGuinness, PhD5, Emily Bratton, PhD, MSPH6, Christina Mack, MSPH, PhD7.
1IQVIA, Vancouver, BC, Canada, 2IQVIA, Reading, United Kingdom, 3IQVIA, NA, DC, USA, 4IQVIA, Durham, NC, USA, 5IQVIA, Ottawa, ON, Canada, 6IQVIA, Hillsborough, NC, USA, 7IQVIA, Chapel Hill, NC, USA.
1IQVIA, Vancouver, BC, Canada, 2IQVIA, Reading, United Kingdom, 3IQVIA, NA, DC, USA, 4IQVIA, Durham, NC, USA, 5IQVIA, Ottawa, ON, Canada, 6IQVIA, Hillsborough, NC, USA, 7IQVIA, Chapel Hill, NC, USA.
OBJECTIVES: Robust validation of AI tools used to assist with data extraction is crucial to ensure scientific validity of literature review. We validated the accuracy of data extracted by an agentic AI literature review platform (“AI”) against manual data extraction by an epidemiologist and compared the time taken for each.
METHODS: Data extraction was performed by epi and AI for 16 variables across 27 pregnancy studies on infant developmental delay (DD). Percentage of full concordance (complete factual alignment between AI and epi despite differences in formatting or language) was reported.
RESULTS: Concordance was highest (85-100%) for structured and consistently reported variables, including study metadata (title, authors, journal), and study descriptors (region, objectives). Full concordance ranged between 67-100% across all variables regardless of structure or consistency; factors impacting extraction accuracy including inconsistent terminology for the same concept (e.g., developmental domains labeled as “examinations”, “skills,” “subscales,” “indices”), information that was implied but not explicitly stated (e.g., head control indicating motor skills), atypical presentation of information (e. g. exposures of interest presented as “non-overlapping subgroups for analysis”), and ambiguity in the prompt (e.g. whether “sample size” refers to the number enrolled or the number completing the study). In studies using multiple instruments to assess DD at varying timepoints, AI accurately matched DD assessment to timepoints but sometimes did not capture all instruments. AI distinguished well between similar concepts (e.g. infant age, gestational age, age at assessment) to identify the variable of interest. In some instances, AI identified and extracted information missed by the epidemiologist, providing a complement to human work. AI reduced extraction time by 73.7% (5 hours vs 19 hours for the epidemiologist).
CONCLUSIONS: Our findings demonstrate that AI can perform data extraction accurately and transparently, when used with appropriate human oversight (particularly for more complex variables), while offering substantial time savings.
METHODS: Data extraction was performed by epi and AI for 16 variables across 27 pregnancy studies on infant developmental delay (DD). Percentage of full concordance (complete factual alignment between AI and epi despite differences in formatting or language) was reported.
RESULTS: Concordance was highest (85-100%) for structured and consistently reported variables, including study metadata (title, authors, journal), and study descriptors (region, objectives). Full concordance ranged between 67-100% across all variables regardless of structure or consistency; factors impacting extraction accuracy including inconsistent terminology for the same concept (e.g., developmental domains labeled as “examinations”, “skills,” “subscales,” “indices”), information that was implied but not explicitly stated (e.g., head control indicating motor skills), atypical presentation of information (e. g. exposures of interest presented as “non-overlapping subgroups for analysis”), and ambiguity in the prompt (e.g. whether “sample size” refers to the number enrolled or the number completing the study). In studies using multiple instruments to assess DD at varying timepoints, AI accurately matched DD assessment to timepoints but sometimes did not capture all instruments. AI distinguished well between similar concepts (e.g. infant age, gestational age, age at assessment) to identify the variable of interest. In some instances, AI identified and extracted information missed by the epidemiologist, providing a complement to human work. AI reduced extraction time by 73.7% (5 hours vs 19 hours for the epidemiologist).
CONCLUSIONS: Our findings demonstrate that AI can perform data extraction accurately and transparently, when used with appropriate human oversight (particularly for more complex variables), while offering substantial time savings.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
SA71
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Literature Review & Synthesis
Disease
Neurological Disorders