FIT FOR PURPOSE? PURPOSE BUILT ARTIFICIAL INTELLIGENCE (AI) TOOLS VS GENERAL PURPOSE LARGE-LANGUAGE MODEL (LLM) FOR DATA EXTRACTION IN EVIDENCE SYNTHESIS
Author(s)
Allie Cichewicz, MSc1, Marius Sauca, BSc, MSc2, Kevin Kallmes, MA, JD3.
1Nested Knowledge, Boston, MA, USA, 2Nested Knowledge, UTRECHT, Netherlands, 3Nested Knowledge, Saint Paul, MN, USA.
1Nested Knowledge, Boston, MA, USA, 2Nested Knowledge, UTRECHT, Netherlands, 3Nested Knowledge, Saint Paul, MN, USA.
OBJECTIVES: The widespread availability and ease of use of general-purpose LLMs have raised questions about whether purpose-built evidence synthesis tools provide meaningful advantages for structured data extraction. This study compared a fit-for-purpose AI tool with a general-purpose LLM, assessing the accuracy, completeness, and consistency of structured data extraction.
METHODS: Five trials in NSCLC were selected for extraction and carried out in Nested Knowledge (NK; v1.111, high fidelity) and Claude (Max, Opus4.7). A tag hierarchy was developed in NK, with prompts for each extraction element. The tabular NK output for one human-validated benchmark study served as Claude’s extraction template, with column-level instructions applied across all studies in a single batch to match NK’s multi-study extraction workflow while simulating a human, template-guided process. Each tool extracted study and patient characteristics and efficacy outcomes; extraction outputs were compared against the human-extracted reference for accuracy and completeness, with results summarized overall, by tool, and by extraction domain.
RESULTS: Both tools exhibited high accuracy for study and patient characteristics, but efficacy extraction was less complete and less reliable. Claude achieved higher accuracy for study characteristics than NK (90.0% vs 86.2%), whereas NK performed slightly better for patient characteristics (86.8% vs 84.6%). For observed efficacy fields by treatment arm, 79.6% and 68.3% of data were correct for NK and Claude, respectively. After accounting for missed subgroups by treatment arm, efficacy performance declined to 39.3% (41 omitted rows) and 19.4% (55 omitted rows), respectively. Claude also produced 38 cell-merging errors and showed greater variation across studies, with weaker performance in later-positioned outputs.
CONCLUSIONS: General-purpose LLMs performed similarly for simpler descriptive fields but were less reliable for complex efficacy extraction. These findings suggest that general-purpose LLMs may support routine evidence extraction, whereas complex HEOR datasets may require more specialized workflows and rigorous quality assurance, where human review remains essential.
METHODS: Five trials in NSCLC were selected for extraction and carried out in Nested Knowledge (NK; v1.111, high fidelity) and Claude (Max, Opus4.7). A tag hierarchy was developed in NK, with prompts for each extraction element. The tabular NK output for one human-validated benchmark study served as Claude’s extraction template, with column-level instructions applied across all studies in a single batch to match NK’s multi-study extraction workflow while simulating a human, template-guided process. Each tool extracted study and patient characteristics and efficacy outcomes; extraction outputs were compared against the human-extracted reference for accuracy and completeness, with results summarized overall, by tool, and by extraction domain.
RESULTS: Both tools exhibited high accuracy for study and patient characteristics, but efficacy extraction was less complete and less reliable. Claude achieved higher accuracy for study characteristics than NK (90.0% vs 86.2%), whereas NK performed slightly better for patient characteristics (86.8% vs 84.6%). For observed efficacy fields by treatment arm, 79.6% and 68.3% of data were correct for NK and Claude, respectively. After accounting for missed subgroups by treatment arm, efficacy performance declined to 39.3% (41 omitted rows) and 19.4% (55 omitted rows), respectively. Claude also produced 38 cell-merging errors and showed greater variation across studies, with weaker performance in later-positioned outputs.
CONCLUSIONS: General-purpose LLMs performed similarly for simpler descriptive fields but were less reliable for complex efficacy extraction. These findings suggest that general-purpose LLMs may support routine evidence extraction, whereas complex HEOR datasets may require more specialized workflows and rigorous quality assurance, where human review remains essential.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR81
Topic
Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
Oncology