EVALUATING AI-ASSISTED VERSUS MANUAL TARGETED LITERATURE REVIEWS: A CASE STUDY ON THE BURDEN OF ILLNESS IN FACIOSCAPULOHUMERAL MUSCULAR DYSTROPHY
Author(s)
Nita Santpurkar, MSc1, Diwakar Jha, MSc1, Hugo Dubucq, MPH2, Kevin H. Li, MS, PharmD3, Duygu Bozkaya, MBA, MSc3.
1Sanofi, Hyderabad, India, 2Sanofi, Barcelona, Spain, 3Sanofi, Cambridge, MA, USA.
1Sanofi, Hyderabad, India, 2Sanofi, Barcelona, Spain, 3Sanofi, Cambridge, MA, USA.
OBJECTIVES: Facioscapulohumeral muscular dystrophy (FSHD) is a rare, progressive neuromuscular disorder with existing burden data limited by small, heterogeneous populations, inconsistent outcome measures, predominance of cross-sectional designs, and under-reporting of patient-reported outcomes and economic burden. This research aimed to conduct an AI-assisted targeted literature review (aiTLR) to accelerate and improve the comprehensiveness of evidence synthesis, characterize the global burden of illness (BoI) associated with FSHD, and systematically identify gaps to inform clinical development.
METHODS: Leveraging the protocol-derived prompt, the generative AI tool screened ~1,000 publications from electronic databases, grey literature, and conference proceedings across all BoI domains. Further, the tool systematically extracted predefined data on epidemiological, clinical, humanistic, and economic outcomes. Data extraction was conducted using a prioritization matrix to ensure inclusion of the best available quality of evidence. AI performance was evaluated against manual review across key performance metrics: sensitivity, specificity, accuracy, completeness, and time saved. 100% human quality check validated all AI screening decisions and extracted data.
RESULTS: From ~1,000 publications, AI screening achieved 93% sensitivity and 70% specificity for study inclusion. Expert validation in 12 days confirmed 243 studies provided complete BoI data, compared to an estimated 22 days for manual full-text review. For data extraction, AI reduced expert time by 50% (30 vs. 60 days manual), achieving 80% accuracy across verified data points and 70% completeness. Systematic evidence mapping identified critical data gaps across all domains, with particularly limited evidence on epidemiology and economic burden — highlighting priority areas for future evidence generation.
CONCLUSIONS: AI-assisted workflow efficiently screened and extracted BoI data from FSHD literature, revealing evidence gaps to inform clinical trial design. However, the results of AI-assisted workflows should be interpreted with caution, as they are contingent on both the complexity of the project and user expertise.
METHODS: Leveraging the protocol-derived prompt, the generative AI tool screened ~1,000 publications from electronic databases, grey literature, and conference proceedings across all BoI domains. Further, the tool systematically extracted predefined data on epidemiological, clinical, humanistic, and economic outcomes. Data extraction was conducted using a prioritization matrix to ensure inclusion of the best available quality of evidence. AI performance was evaluated against manual review across key performance metrics: sensitivity, specificity, accuracy, completeness, and time saved. 100% human quality check validated all AI screening decisions and extracted data.
RESULTS: From ~1,000 publications, AI screening achieved 93% sensitivity and 70% specificity for study inclusion. Expert validation in 12 days confirmed 243 studies provided complete BoI data, compared to an estimated 22 days for manual full-text review. For data extraction, AI reduced expert time by 50% (30 vs. 60 days manual), achieving 80% accuracy across verified data points and 70% completeness. Systematic evidence mapping identified critical data gaps across all domains, with particularly limited evidence on epidemiology and economic burden — highlighting priority areas for future evidence generation.
CONCLUSIONS: AI-assisted workflow efficiently screened and extracted BoI data from FSHD literature, revealing evidence gaps to inform clinical trial design. However, the results of AI-assisted workflows should be interpreted with caution, as they are contingent on both the complexity of the project and user expertise.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
SA12
Topic
Health Technology Assessment, Methodological & Statistical Research, Study Approaches
Topic Subcategory
Literature Review & Synthesis
Disease
No Additional Disease & Conditions/Specialized Treatment Areas, Rare & Orphan Diseases