A SENTENCE-LEVEL ASSESSMENT FRAMEWORK FOR EVALUATING AI-GENERATED STUDY SUMMARIES AGAINST EXPERT-CURATED CLINICAL OVERVIEW
Author(s)
Viji Queen V, PharmD1, George Alisha, B.E.2, Angeline Babitha Dhas, BS3, Swathirajan C R, Ph.D2, Revanth M, B.E.2, Meghan Oates-Zalesky, MSc4.
1MadeAI, Nagercoil, India, 2MadeAi, Nagercoil, India, 3MadeAi, Cambridge, MA, USA, 4Chief Marketing Officer, MadeAi, Cambridge, MA, USA.
1MadeAI, Nagercoil, India, 2MadeAi, Nagercoil, India, 3MadeAi, Cambridge, MA, USA, 4Chief Marketing Officer, MadeAi, Cambridge, MA, USA.
OBJECTIVES: Artificial intelligence (AI) is increasingly used to generate study summaries for evidence synthesis and regulatory writing. However, existing evaluations often rely on lexical similarity and document-level metrics that may overlook important errors. This study developed and applied a sentence-level framework to evaluate AI-generated study summaries against expert-curated clinical evidence narratives.
METHODS: A comparative evaluation was conducted using 38 published clinical studies referenced in Clinical Overview (CO) documents. Full-text publications were processed through the MadeAi-LR Platform to generate study summaries. Corresponding CO summaries served as the reference standard. A sentence-level framework assessed relevance, accuracy, completeness, and information order, yielding a maximum score of 7 points per sentence. Aggregate and domain-specific scores were calculated, and inaccuracies and omissions were identified.
RESULTS: A total of 180 generated summary sentences were evaluated. The mean overall score was 6.12/7 (87.4% of the maximum score). Mean domain scores were 1.96/2 for relevance, 1.69/2 for accuracy, 1.47/2 for completeness, and 1/1 for information order. Overall sentence-level accuracy was 82.2%, with 69.4% of sentences receiving the maximum accuracy score. Despite strong aggregate performance, critical inaccuracies and omissions were identified in 7.2% and 8.9% of sentences, respectively, most commonly involving subgroup analyses, nuanced safety findings, and contextual interpretation of outcomes. Initial summary generation time decreased human effort by 95.7%, while end-to-end summarization time, including reviewer validation, decreased by 47.8%.
CONCLUSIONS: A sentence-level evaluation framework offers a robust method for identifying inaccuracies and omissions in AI-generated study summaries that are often masked by high aggregate performance scores. Through granular assessment, it enables targeted optimization and validation of AI systems, facilitating continuous improvements in output accuracy and completeness. This may strengthen confidence in AI-assisted evidence synthesis and accelerate its adoption in high-accuracy medical writing applications.
METHODS: A comparative evaluation was conducted using 38 published clinical studies referenced in Clinical Overview (CO) documents. Full-text publications were processed through the MadeAi-LR Platform to generate study summaries. Corresponding CO summaries served as the reference standard. A sentence-level framework assessed relevance, accuracy, completeness, and information order, yielding a maximum score of 7 points per sentence. Aggregate and domain-specific scores were calculated, and inaccuracies and omissions were identified.
RESULTS: A total of 180 generated summary sentences were evaluated. The mean overall score was 6.12/7 (87.4% of the maximum score). Mean domain scores were 1.96/2 for relevance, 1.69/2 for accuracy, 1.47/2 for completeness, and 1/1 for information order. Overall sentence-level accuracy was 82.2%, with 69.4% of sentences receiving the maximum accuracy score. Despite strong aggregate performance, critical inaccuracies and omissions were identified in 7.2% and 8.9% of sentences, respectively, most commonly involving subgroup analyses, nuanced safety findings, and contextual interpretation of outcomes. Initial summary generation time decreased human effort by 95.7%, while end-to-end summarization time, including reviewer validation, decreased by 47.8%.
CONCLUSIONS: A sentence-level evaluation framework offers a robust method for identifying inaccuracies and omissions in AI-generated study summaries that are often masked by high aggregate performance scores. Through granular assessment, it enables targeted optimization and validation of AI systems, facilitating continuous improvements in output accuracy and completeness. This may strengthen confidence in AI-assisted evidence synthesis and accelerate its adoption in high-accuracy medical writing applications.
Conference/Value in Health Info
2026-11, ISPOR Europe 2026, Vienna, Austria
Value in Health, Volume 29, Issue 12S
Code
MSR77
Topic
Health Technology Assessment, Methodological & Statistical Research, Study Approaches
Topic Subcategory
Artificial Intelligence, Machine Learning, Predictive Analytics
Disease
No Additional Disease & Conditions/Specialized Treatment Areas