ACCELERATING RWE TREATMENT PATTERN SYNTHESIS THROUGH AI-ASSISTED LITERATURE REVIEW
Author(s)
Adeline ABBE, PhD1, Hugo Dubucq, MPH2, Paulo Carita, DrPH3.
1Aixial, Sevres, France, 2Sanofi, Barcelona, Spain, 3Sanofi, Gentilly, France.
1Aixial, Sevres, France, 2Sanofi, Barcelona, Spain, 3Sanofi, Gentilly, France.
OBJECTIVES: Real-world evidence (RWE) treatment patterns inform trial comparators and PICOS frameworks. However, treatment utilization data remains fragmented across diverse sources (registries, claims databases, observational cohorts) with inconsistent reporting formats. We evaluated whether an AI-assisted literature review could identify colorectal cancer treatment patterns across geographies with reduced time and maintained accuracy versus manual expert curation.
METHODS: We used a generative AI tool with documented prompts to screen publications using pre-specified PICOS criteria for treatment. Then AI extracted three standardized data points (treatment line, treatment type, percentage of use) with complete provenance metadata. Finally, experts reviewed all AI-flagged studies and validated extracted data, with discrepancies resolved through documented adjudication. We compared AI vs manual expert review for screening sensitivity/specificity, extraction accuracy, completeness of treatment pattern reporting, and time efficiency.
RESULTS: From 500+ publications, AI screening achieved 84% sensitivity, 94% specificity, 79% precision. Expert validation of the screening, guided by AI explainability, was completed in 3 versus 12 days estimated manual review (75% time reduction). This efficiency stems from focused validation of included studies, rapid triage of excluded studies, and AI-guided navigation to relevant text spans. AI-assisted extraction required less than 10 minutes versus one hour manually, demonstrating 96% accuracy for treatment pattern data, despite the complexity of distinguishing treatment patterns from baseline characteristics. All extractions included complete provenance metadata enabling independent verification.
CONCLUSIONS: AI-assisted workflow achieved reduction while maintaining accuracy and generating governance artifacts for good practice. This demonstrates that credibility-by-design through transparency, traceability, and human oversight can enable rapid evidence synthesis from heterogeneous sources. AI performance depends on tool version, prompt engineering, data complexity, and complete documentation for reproducibility. These results support proportionality frameworks where AI-assisted methods are appropriate for exploratory RWE synthesis when benchmarked against manual review and paired with human validation of critical data points.
METHODS: We used a generative AI tool with documented prompts to screen publications using pre-specified PICOS criteria for treatment. Then AI extracted three standardized data points (treatment line, treatment type, percentage of use) with complete provenance metadata. Finally, experts reviewed all AI-flagged studies and validated extracted data, with discrepancies resolved through documented adjudication. We compared AI vs manual expert review for screening sensitivity/specificity, extraction accuracy, completeness of treatment pattern reporting, and time efficiency.
RESULTS: From 500+ publications, AI screening achieved 84% sensitivity, 94% specificity, 79% precision. Expert validation of the screening, guided by AI explainability, was completed in 3 versus 12 days estimated manual review (75% time reduction). This efficiency stems from focused validation of included studies, rapid triage of excluded studies, and AI-guided navigation to relevant text spans. AI-assisted extraction required less than 10 minutes versus one hour manually, demonstrating 96% accuracy for treatment pattern data, despite the complexity of distinguishing treatment patterns from baseline characteristics. All extractions included complete provenance metadata enabling independent verification.
CONCLUSIONS: AI-assisted workflow achieved reduction while maintaining accuracy and generating governance artifacts for good practice. This demonstrates that credibility-by-design through transparency, traceability, and human oversight can enable rapid evidence synthesis from heterogeneous sources. AI performance depends on tool version, prompt engineering, data complexity, and complete documentation for reproducibility. These results support proportionality frameworks where AI-assisted methods are appropriate for exploratory RWE synthesis when benchmarked against manual review and paired with human validation of critical data points.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR138
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
Oncology