FASTER, MOSTLY FAITHFUL: CAN AI REPRODUCE MANUAL PICO EXTRACTION IN PREPARATION FOR EU JCA?
Author(s)
Fadel Shoughari, MSc1, Patrick Adams, BA1, Allie Cichewicz, MSc2, Lisa Nicole Cox, BSc1, Kevin Kallmes, MA, JD3.
1Acumetis, London, United Kingdom, 2Consultant, Pittsfield, MA, USA, 3Nested Knowledge, St. Paul, MN, USA.
1Acumetis, London, United Kingdom, 2Consultant, Pittsfield, MA, USA, 3Nested Knowledge, St. Paul, MN, USA.
OBJECTIVES: Preparation for EU Joint Clinical Assessment (JCA) requires PICO (population, intervention, comparator, outcome) extraction from national health technology assessment (HTA) reports and clinical practice guidelines across member states, which is a labour-intensive task. Using a second-line-plus NSCLC asset as a case study, we assessed whether complete artificial intelligence (AI)-driven extraction using Nested Knowledge could reproduce manual PICO extraction and evaluated the associated time saving.
METHODS: PICO elements were extracted manually and by AI from eight HTA reports and 16 country-level guidelines (four PICO fields each; 96 total datapoint comparisons). Each AI output was scored against the manual reference as "Same", "Similar", "Needs review", "Missing", or "Different". Concordance was summarised by PICO field and turnaround time was recorded for both approaches.
RESULTS: For HTA reports, 28/32 datapoints (88%) were concordant (Same/Similar). Intervention matched perfectly (100%); population (87.5%) and comparator (87.5%) were strong, with gaps where AI truncated comparator lists or dropped population qualifiers. 75% of outcomes were scored Same/Similar, and the single true error was one report where AI replaced the clinical endpoint with economic outcomes. For guidelines, 7/16 sources (44%) were concordant. However, this mainly reflected more comprehensive AI extraction: manual extraction exclusively extracted the second-line-plus data whereas AI returned the full guidelines. Manual extraction took two weeks, whereas AI took two hours (~40-fold faster).
CONCLUSIONS: AI reproduced manual HTA PICO extractions closely while cutting turnaround by ~40-fold. Apparent guideline extraction disagreement largely reflected AI's comprehensiveness versus a narrow manual reference: the AI prompts were general and did not constrain extraction to the relevant second-line-plus setting. Refining prompts to specify line of therapy and asset-relevant scope is therefore key to more accurate, fit-for-purpose extractions. These findings support deploying AI as a first-pass engine for JCA PICO scoping, validated and refined by human experts, shifting human effort from manual extraction to higher-value strategic review.
METHODS: PICO elements were extracted manually and by AI from eight HTA reports and 16 country-level guidelines (four PICO fields each; 96 total datapoint comparisons). Each AI output was scored against the manual reference as "Same", "Similar", "Needs review", "Missing", or "Different". Concordance was summarised by PICO field and turnaround time was recorded for both approaches.
RESULTS: For HTA reports, 28/32 datapoints (88%) were concordant (Same/Similar). Intervention matched perfectly (100%); population (87.5%) and comparator (87.5%) were strong, with gaps where AI truncated comparator lists or dropped population qualifiers. 75% of outcomes were scored Same/Similar, and the single true error was one report where AI replaced the clinical endpoint with economic outcomes. For guidelines, 7/16 sources (44%) were concordant. However, this mainly reflected more comprehensive AI extraction: manual extraction exclusively extracted the second-line-plus data whereas AI returned the full guidelines. Manual extraction took two weeks, whereas AI took two hours (~40-fold faster).
CONCLUSIONS: AI reproduced manual HTA PICO extractions closely while cutting turnaround by ~40-fold. Apparent guideline extraction disagreement largely reflected AI's comprehensiveness versus a narrow manual reference: the AI prompts were general and did not constrain extraction to the relevant second-line-plus setting. Refining prompts to specify line of therapy and asset-relevant scope is therefore key to more accurate, fit-for-purpose extractions. These findings support deploying AI as a first-pass engine for JCA PICO scoping, validated and refined by human experts, shifting human effort from manual extraction to higher-value strategic review.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR218
Topic
Health Policy & Regulatory, Health Technology Assessment, Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas