GOLDILOCKS AND THE THREE RETRIEVAL APPROACHES: HUMAN-REVIEWED VALIDATION OF LARGE LANGUAGE MODEL (LLM) EMBEDDING-DRIVEN RETRIEVAL FOR TRUSTWORTHY EVIDENCE GAP ANALYSIS AND VALUE DOSSIER DRAFTING

Author(s)

Katelyn Keyloun, BS, MS, PharmD1, Tyler Reinsch, PharmD2, Gavin J. Outteridge, MA3, Illia Sydun, MA4, Anwar Sabir, MA5, Joseph Shaheen, MA3, Gabriel Bishop, MS, MA6.
1Director, Product Innovation & Development, Arysana, Carson City, NV, USA, 2Arysana, Springfield, MO, USA, 3Arysana, Durham, NC, USA, 4Arysana, Hollywood, FL, USA, 5Arysana, Boston, MA, USA, 6Arysana, Palo Alto, CA, USA.
OBJECTIVES: While AI-assisted research may accelerate evidence synthesis for Evidence Gap Analyses or Value Dossier development, little is known about the impact of different retrieval approaches (RA). The purpose of this study was to evaluate deterministic and LLM-driven retrieval methods.
METHODS: RA performance was assessed in two separate therapeutics areas (Immunology and Psychiatry) with small (n=28 articles) and large (n=82 articles) content sizes and with specific Predefined retrieval queries (PRQ; 26, 29, respectively) representing economic, humanistic, and clinical sections. Articles were extracted into tagged chunks. Baseline RA (deterministic, tagging, section rules) was compared to advanced RAs: RA1 (OpenAI text-embedding-3-large alone), RA2 (added BM25 keyword search and reciprocal-rank fusion), RA3.1 (added ranking weight for scispaCy-derived terms, using en_core_sci_sm model) and RA3.2 (weighted specific ~5.3% scispaCY-derived terms). Automated performance metrics included: % useful chunks among the top-20 chunks, ≥1 useful chunk in the top-20 chunks, first useful chunk rank, where human-assessed performance was compared for a random 10% sample of PRQs.
RESULTS: 401 Immunology and 6,626 psychiatry chunks were retrieved across PRQs. Baseline RA exhibited the lowest performance for Psychiatry PRQs across sections (77.8%-100.0% of PRQs with ≥1 useful top-20 chunk; 51.1%-76.7% useful top-20 chunks; mean first useful rank 1.5-18.7). RA1 and RA2 generally experienced higher performance versus baseline. RA3.1 performed well for Immunology in the Disease Burden/Unmet Need section (96.1% useful top-20 chunks), yet performance was reduced for Psychiatry (75.0%-85.0% useful top-20 chunks). RA2 and RA3.2 demonstrated the strongest performance (85.8%-100% useful top-20 chunks, 100% of PRQs with ≥1 useful top-20 chunk, and mean first useful rank 1.0-1.6). Automated performance metrics were generally higher than human-assessed performance; however, both assessments supported improved performance of advanced retrieval approaches.
CONCLUSIONS: Combining deterministic and LLM-driven approaches improved retrieval for AI-assisted Evidence Gap Analysis and Value Dossier development, although broad tagging approaches may reduce performance, especially as content increases.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR75

Topic

Methodological & Statistical Research, Real World Data & Information Systems, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

Mental Health (including addiction), Systemic Disorders/Conditions (Anesthesia, Auto-Immune Disorders (n.e.c.), Hematological Disorders (non-oncologic), Pain)

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×