FIT FOR PURPOSE? PURPOSE BUILT ARTIFICIAL INTELLIGENCE (AI) TOOLS VS GENERAL PURPOSE LARGE-LANGUAGE MODEL (LLM) FOR DATA EXTRACTION IN EVIDENCE SYNTHESIS

Author(s)

Allie Cichewicz, MSc1, Marius Sauca, BSc, MSc2, Kevin Kallmes, MA, JD3.
1Nested Knowledge, Boston, MA, USA, 2Nested Knowledge, UTRECHT, Netherlands, 3Nested Knowledge, Saint Paul, MN, USA.
OBJECTIVES: The widespread availability and ease of use of general-purpose LLMs have raised questions about whether purpose-built evidence synthesis tools provide meaningful advantages for structured data extraction. This study compared a fit-for-purpose AI tool with a general-purpose LLM, assessing the accuracy, completeness, and consistency of structured data extraction.
METHODS: Five trials in NSCLC were selected for extraction and carried out in Nested Knowledge (NK; v1.111, high fidelity) and Claude (Max, Opus4.7). A tag hierarchy was developed in NK, with prompts for each extraction element. The tabular NK output for one human-validated benchmark study served as Claude’s extraction template, with column-level instructions applied across all studies in a single batch to match NK’s multi-study extraction workflow while simulating a human, template-guided process. Each tool extracted study and patient characteristics and efficacy outcomes; extraction outputs were compared against the human-extracted reference for accuracy and completeness, with results summarized overall, by tool, and by extraction domain.
RESULTS: Both tools exhibited high accuracy for study and patient characteristics, but efficacy extraction was less complete and less reliable. Claude achieved higher accuracy for study characteristics than NK (90.0% vs 86.2%), whereas NK performed slightly better for patient characteristics (86.8% vs 84.6%). For observed efficacy fields by treatment arm, 79.6% and 68.3% of data were correct for NK and Claude, respectively. After accounting for missed subgroups by treatment arm, efficacy performance declined to 39.3% (41 omitted rows) and 19.4% (55 omitted rows), respectively. Claude also produced 38 cell-merging errors and showed greater variation across studies, with weaker performance in later-positioned outputs.
CONCLUSIONS: General-purpose LLMs performed similarly for simpler descriptive fields but were less reliable for complex efficacy extraction. These findings suggest that general-purpose LLMs may support routine evidence extraction, whereas complex HEOR datasets may require more specialized workflows and rigorous quality assurance, where human review remains essential.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR81

Topic

Methodological & Statistical Research, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

Oncology

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×