PROMPT SPECIFICATION AND EXPERT OVERSIGHT SHAPE AI-ASSISTED LITERATURE REVIEW OUTPUTS ACROSS DISEASE AREAS
Author(s)
Maxime Gobin, PhD.
Biolevate, SUCY EN BRIE, France.
Biolevate, SUCY EN BRIE, France.
OBJECTIVES: To examine how prompt specification and expert oversight affect AI-assisted literature-review workflows, and to identify workflow components required for scalable, reproducible, and auditable HEOR evidence generation.
METHODS: We conducted a retrospective methodological analysis of three anonymized literature-review datasets comprising 1,243, 1,393, and 246 initial records. AI-assisted outputs were assessed across title/abstract screening, full-text screening, data extraction, PRISMA flow generation, and narrative synthesis by several user groups to each dataset. Prompts were decomposed into methodological components including population, disease scope, outcomes, study design, inclusion/exclusion rules, uncertainty handling, output format, and source-evidence requirements. Broad and strict prompting strategies were compared using retained records, included full texts, and extracted values across user-generated outputs. Available expected-response, reference, or human-review candidate fields were used to explore AI-versus-reference screening alignment when same-stage comparisons were possible.
RESULTS: Prompt structure materially affected screening outputs. Across repeated user-generated title/abstract screening outputs, broad versus strict prompts retained 248.3±113.0 versus 159.8±17.7 records in Dataset A, 537.8±33.0 versus 288.0±200.5 in Dataset B, and 36.4±7.3 versus 25.6±5.7 in Dataset C. Similar prompt-dependent variation was observed at full-text screening and extraction. More restrictive outputs were associated with narrower population definitions, tighter disease scope, explicit study-design constraints, stricter outcome requirements, and stronger source-evidence thresholds. In available AI-versus-reference screening checks covering 2,908 decisions, exact agreement reached 75.0% (2,182/2,908), with most disagreements linked to uncertainty or full-text-triage decisions.. PRISMA diagrams and narrative syntheses were reproducible in structure, while counts, exclusion reasons, emphasis, and disease-specific content reflected upstream prompt and screening decisions.
CONCLUSIONS: AI-assisted reviews require methodological governance rather than automation alone. Reliability depends on expert-defined prompts, reusable criteria presets, source-traceable extraction, uncertainty rules, and human verification. These components can support scalable, reproducible, and audit-ready literature-review workflows for HEOR evidence generation.
METHODS: We conducted a retrospective methodological analysis of three anonymized literature-review datasets comprising 1,243, 1,393, and 246 initial records. AI-assisted outputs were assessed across title/abstract screening, full-text screening, data extraction, PRISMA flow generation, and narrative synthesis by several user groups to each dataset. Prompts were decomposed into methodological components including population, disease scope, outcomes, study design, inclusion/exclusion rules, uncertainty handling, output format, and source-evidence requirements. Broad and strict prompting strategies were compared using retained records, included full texts, and extracted values across user-generated outputs. Available expected-response, reference, or human-review candidate fields were used to explore AI-versus-reference screening alignment when same-stage comparisons were possible.
RESULTS: Prompt structure materially affected screening outputs. Across repeated user-generated title/abstract screening outputs, broad versus strict prompts retained 248.3±113.0 versus 159.8±17.7 records in Dataset A, 537.8±33.0 versus 288.0±200.5 in Dataset B, and 36.4±7.3 versus 25.6±5.7 in Dataset C. Similar prompt-dependent variation was observed at full-text screening and extraction. More restrictive outputs were associated with narrower population definitions, tighter disease scope, explicit study-design constraints, stricter outcome requirements, and stronger source-evidence thresholds. In available AI-versus-reference screening checks covering 2,908 decisions, exact agreement reached 75.0% (2,182/2,908), with most disagreements linked to uncertainty or full-text-triage decisions.. PRISMA diagrams and narrative syntheses were reproducible in structure, while counts, exclusion reasons, emphasis, and disease-specific content reflected upstream prompt and screening decisions.
CONCLUSIONS: AI-assisted reviews require methodological governance rather than automation alone. Reliability depends on expert-defined prompts, reusable criteria presets, source-traceable extraction, uncertainty rules, and human verification. These components can support scalable, reproducible, and audit-ready literature-review workflows for HEOR evidence generation.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR53
Topic
Clinical Outcomes, Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas