GOLDILOCKS AND THE THREE RETRIEVAL APPROACHES: HUMAN-REVIEWED VALIDATION OF LARGE LANGUAGE MODEL (LLM) EMBEDDING-DRIVEN RETRIEVAL FOR TRUSTWORTHY EVIDENCE GAP ANALYSIS AND VALUE DOSSIER DRAFTING
Author(s)
Katelyn Keyloun, BS, MS, PharmD1, Tyler Reinsch, PharmD2, Gavin J. Outteridge, MA3, Illia Sydun, MA4, Anwar Sabir, MA5, Joseph Shaheen, MA3, Gabriel Bishop, MS, MA6.
1Director, Product Innovation & Development, Arysana, Carson City, NV, USA, 2Arysana, Springfield, MO, USA, 3Arysana, Durham, NC, USA, 4Arysana, Hollywood, FL, USA, 5Arysana, Boston, MA, USA, 6Arysana, Palo Alto, CA, USA.
1Director, Product Innovation & Development, Arysana, Carson City, NV, USA, 2Arysana, Springfield, MO, USA, 3Arysana, Durham, NC, USA, 4Arysana, Hollywood, FL, USA, 5Arysana, Boston, MA, USA, 6Arysana, Palo Alto, CA, USA.
OBJECTIVES: While AI-assisted research may accelerate evidence synthesis for Evidence Gap Analyses or Value Dossier development, little is known about the impact of different retrieval approaches (RA). The purpose of this study was to evaluate deterministic and LLM-driven retrieval methods.
METHODS: RA performance was assessed in two separate therapeutics areas (Immunology and Psychiatry) with small (n=28 articles) and large (n=82 articles) content sizes and with specific Predefined retrieval queries (PRQ; 26, 29, respectively) representing economic, humanistic, and clinical sections. Articles were extracted into tagged chunks. Baseline RA (deterministic, tagging, section rules) was compared to advanced RAs: RA1 (OpenAI text-embedding-3-large alone), RA2 (added BM25 keyword search and reciprocal-rank fusion), RA3.1 (added ranking weight for scispaCy-derived terms, using en_core_sci_sm model) and RA3.2 (weighted specific ~5.3% scispaCY-derived terms). Automated performance metrics included: % useful chunks among the top-20 chunks, ≥1 useful chunk in the top-20 chunks, first useful chunk rank, where human-assessed performance was compared for a random 10% sample of PRQs.
RESULTS: 401 Immunology and 6,626 psychiatry chunks were retrieved across PRQs. Baseline RA exhibited the lowest performance for Psychiatry PRQs across sections (77.8%-100.0% of PRQs with ≥1 useful top-20 chunk; 51.1%-76.7% useful top-20 chunks; mean first useful rank 1.5-18.7). RA1 and RA2 generally experienced higher performance versus baseline. RA3.1 performed well for Immunology in the Disease Burden/Unmet Need section (96.1% useful top-20 chunks), yet performance was reduced for Psychiatry (75.0%-85.0% useful top-20 chunks). RA2 and RA3.2 demonstrated the strongest performance (85.8%-100% useful top-20 chunks, 100% of PRQs with ≥1 useful top-20 chunk, and mean first useful rank 1.0-1.6). Automated performance metrics were generally higher than human-assessed performance; however, both assessments supported improved performance of advanced retrieval approaches.
CONCLUSIONS: Combining deterministic and LLM-driven approaches improved retrieval for AI-assisted Evidence Gap Analysis and Value Dossier development, although broad tagging approaches may reduce performance, especially as content increases.
METHODS: RA performance was assessed in two separate therapeutics areas (Immunology and Psychiatry) with small (n=28 articles) and large (n=82 articles) content sizes and with specific Predefined retrieval queries (PRQ; 26, 29, respectively) representing economic, humanistic, and clinical sections. Articles were extracted into tagged chunks. Baseline RA (deterministic, tagging, section rules) was compared to advanced RAs: RA1 (OpenAI text-embedding-3-large alone), RA2 (added BM25 keyword search and reciprocal-rank fusion), RA3.1 (added ranking weight for scispaCy-derived terms, using en_core_sci_sm model) and RA3.2 (weighted specific ~5.3% scispaCY-derived terms). Automated performance metrics included: % useful chunks among the top-20 chunks, ≥1 useful chunk in the top-20 chunks, first useful chunk rank, where human-assessed performance was compared for a random 10% sample of PRQs.
RESULTS: 401 Immunology and 6,626 psychiatry chunks were retrieved across PRQs. Baseline RA exhibited the lowest performance for Psychiatry PRQs across sections (77.8%-100.0% of PRQs with ≥1 useful top-20 chunk; 51.1%-76.7% useful top-20 chunks; mean first useful rank 1.5-18.7). RA1 and RA2 generally experienced higher performance versus baseline. RA3.1 performed well for Immunology in the Disease Burden/Unmet Need section (96.1% useful top-20 chunks), yet performance was reduced for Psychiatry (75.0%-85.0% useful top-20 chunks). RA2 and RA3.2 demonstrated the strongest performance (85.8%-100% useful top-20 chunks, 100% of PRQs with ≥1 useful top-20 chunk, and mean first useful rank 1.0-1.6). Automated performance metrics were generally higher than human-assessed performance; however, both assessments supported improved performance of advanced retrieval approaches.
CONCLUSIONS: Combining deterministic and LLM-driven approaches improved retrieval for AI-assisted Evidence Gap Analysis and Value Dossier development, although broad tagging approaches may reduce performance, especially as content increases.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR75
Topic
Methodological & Statistical Research, Real World Data & Information Systems, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
Mental Health (including addiction), Systemic Disorders/Conditions (Anesthesia, Auto-Immune Disorders (n.e.c.), Hematological Disorders (non-oncologic), Pain)