CLOSING THE LANGUAGE GAP IN HEOR REVIEWS: VALIDATING AI-ASSISTED SCREENING OF CHINESE-LANGUAGE EVIDENCE
Author(s)
Caroline von Wilamowitz-Moellendorff, PhD1, Ziyi Li2, Jade Thurnham3, Aiswarya Shree, MSc4, Andreas Freitag, MSc1, Caoimhe Treise Rice, MA, MSc, MD5, Mireia Raluy Callado, MSc6, Lu Ban, PhD7.
1Thermo Fisher, London, United Kingdom, 2Beijing, China, 3Nested Knowledge, London, United Kingdom, 4Thermo Fisher, Bhubaneswar, India, 5Thermo Fisher, Bristol, United Kingdom, 6Thermo Fisher, Stockholm, Sweden, 7Thermo Fisher, Beijing, China.
1Thermo Fisher, London, United Kingdom, 2Beijing, China, 3Nested Knowledge, London, United Kingdom, 4Thermo Fisher, Bhubaneswar, India, 5Thermo Fisher, Bristol, United Kingdom, 6Thermo Fisher, Stockholm, Sweden, 7Thermo Fisher, Beijing, China.
OBJECTIVES: European health technology assessment and market access decisions increasingly rely on global real-world evidence (RWE), but Asian-language literature may be underrepresented in evidence reviews because of language barriers. This study evaluated whether the Nested Knowledge Robot Screener could support AI-assisted, human-in-the-loop screening of Chinese-language abstracts using a targeted literature review (TLR) of RWE on the direct and indirect costs of lung, colorectal, and breast cancer care in mainland China as a case study.
METHODS: Chinese-language publications (CKNI and Wanfang) were searched for this TLR. Within Nested Knowledge, the Robot Screener model was trained using known relevant Chinese-language articles and assigned an advancement probability score between 0 (irrelevant) and 1 (highly likely to be relevant) to each remaining abstract. Human reviewers used these scores to prioritize abstracts, focusing on records scoring ≥0.70 and validating a sample of lower-scoring records to assess threshold-based exclusions. These validations were used to determine whether advancement probability scores supported efficient identification of relevant evidence while maintaining human oversight of inclusion decisions.
RESULTS: In total, 3,053 Chinese-language records were identified. Using the AI-assisted, human-in-the-loop workflow, 109 abstracts were prioritized for human screening, and approximately one-third of AI-prioritised abstracts were considered suitable for answering the TLR research questions.. Although abstracts scoring ≥0.70 were initially prioritized, the most relevant evidence was concentrated among records scoring ≥0.80. Records scoring 0.70-0.79 were generally less relevant. Performance was broadly consistent with prior experience using the tool for English-language abstract screening.
CONCLUSIONS: AI-assisted, human-in-the-loop screening supported efficient identification of relevant Chinese-language RWE while substantially reducing the volume of abstracts requiring manual review. These findings suggest that AI-enabled screening may help European and global HEOR researchers incorporate Asian-language evidence more systematically into literature reviews. Human validation remains important when applying threshold-based prioritization or exclusion to non-English evidence sources.
METHODS: Chinese-language publications (CKNI and Wanfang) were searched for this TLR. Within Nested Knowledge, the Robot Screener model was trained using known relevant Chinese-language articles and assigned an advancement probability score between 0 (irrelevant) and 1 (highly likely to be relevant) to each remaining abstract. Human reviewers used these scores to prioritize abstracts, focusing on records scoring ≥0.70 and validating a sample of lower-scoring records to assess threshold-based exclusions. These validations were used to determine whether advancement probability scores supported efficient identification of relevant evidence while maintaining human oversight of inclusion decisions.
RESULTS: In total, 3,053 Chinese-language records were identified. Using the AI-assisted, human-in-the-loop workflow, 109 abstracts were prioritized for human screening, and approximately one-third of AI-prioritised abstracts were considered suitable for answering the TLR research questions.. Although abstracts scoring ≥0.70 were initially prioritized, the most relevant evidence was concentrated among records scoring ≥0.80. Records scoring 0.70-0.79 were generally less relevant. Performance was broadly consistent with prior experience using the tool for English-language abstract screening.
CONCLUSIONS: AI-assisted, human-in-the-loop screening supported efficient identification of relevant Chinese-language RWE while substantially reducing the volume of abstracts requiring manual review. These findings suggest that AI-enabled screening may help European and global HEOR researchers incorporate Asian-language evidence more systematically into literature reviews. Human validation remains important when applying threshold-based prioritization or exclusion to non-English evidence sources.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR195
Topic
Methodological & Statistical Research
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
Oncology