NO REVIEWER LEFT BEHIND: AN OPEN-SOURCE LLM APPLICATION AUTOMATING THE FULL SYSTEMATIC REVIEW PIPELINE WITH STARR PROTOCOL COMPLIANCE

Author(s)

Valentina D. Batorova, MA, Kirill Sapozhnikov, BA, BEc, BS, BSc, MA, MD, Daria Tolkacheva, MS.
Russian Presidential Academy of National Economy and Public Administration, Moscow, Russian Federation.
OBJECTIVES: Systematic literature reviews (SLRs) are a cornerstone of health technology assessment (HTA), yet remain resource-intensive and prone to inter-reviewer variability. We developed an open-source application that automates the complete SLR pipeline — from search query generation to full-text screening — using locally deployable or cloud-based large language models (LLMs).
METHODS: The application integrates LLMs (deployable via Ollama for local, privacy-preserving use, or via API to commercial providers) across three sequential modules: (1) search query generation from PICO frameworks, PROSPERO protocols, or free-text specifications; (2) title/abstract screening following the STARR Protocol — with dual-pass cross-validation and mandatory re-evaluation of excluded records to minimise false negatives; (3) full-text screening via GLM-OCR for structured extraction from complex PDFs, with classification by intervention, population, endpoint, assessment timepoint, and user-defined categories. Outputs are exported as Excel and RIS files with per-article decisions and inclusion rationale.
RESULTS: The tool was validated on an SLR in oncology (total citations screened: n=2,847). At the title/abstract stage, LLM-assisted screening achieved sensitivity of 96.3% and specificity of 89.1%, with the STARR re-evaluation step recovering 31% of initially excluded relevant records. Full-text GLM-OCR screening demonstrated 93.7% concordance with independent dual-reviewer decisions. Local Ollama deployment (Llama 3.1 70B) performed within 4% of cloud-based GPT-4o on sensitivity metrics.
CONCLUSIONS: This open-source solution substantially reduces the time and human resource burden of SLRs while maintaining methodological rigor through protocol-aligned, multi-step LLM screening. Local deployment via Ollama addresses data confidentiality requirements critical in pharmaceutical and HTA contexts. The tool is positioned as a human-in-the-loop assistant, augmenting rather than replacing expert reviewer judgment. Code and documentation are publicly available at GitHub.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

HTA160

Topic

Health Technology Assessment, Organizational Practices

Topic Subcategory

Systems & Structure

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×