A SENTENCE-LEVEL ASSESSMENT FRAMEWORK FOR EVALUATING AI-GENERATED STUDY SUMMARIES AGAINST EXPERT-CURATED CLINICAL OVERVIEW

Author(s)

Viji Queen V, PharmD1, George Alisha, B.E.2, Angeline Babitha Dhas, BS3, Swathirajan C R, Ph.D2, Revanth M, B.E.2, Meghan Oates-Zalesky, MSc4.
1MadeAI, Nagercoil, India, 2MadeAi, Nagercoil, India, 3MadeAi, Cambridge, MA, USA, 4Chief Marketing Officer, MadeAi, Cambridge, MA, USA.
OBJECTIVES: Artificial intelligence (AI) is increasingly used to generate study summaries for evidence synthesis and regulatory writing. However, existing evaluations often rely on lexical similarity and document-level metrics that may overlook important errors. This study developed and applied a sentence-level framework to evaluate AI-generated study summaries against expert-curated clinical evidence narratives.
METHODS: A comparative evaluation was conducted using 38 published clinical studies referenced in Clinical Overview (CO) documents. Full-text publications were processed through the MadeAi-LR Platform to generate study summaries. Corresponding CO summaries served as the reference standard. A sentence-level framework assessed relevance, accuracy, completeness, and information order, yielding a maximum score of 7 points per sentence. Aggregate and domain-specific scores were calculated, and inaccuracies and omissions were identified.
RESULTS: A total of 180 generated summary sentences were evaluated. The mean overall score was 6.12/7 (87.4% of the maximum score). Mean domain scores were 1.96/2 for relevance, 1.69/2 for accuracy, 1.47/2 for completeness, and 1/1 for information order. Overall sentence-level accuracy was 82.2%, with 69.4% of sentences receiving the maximum accuracy score. Despite strong aggregate performance, critical inaccuracies and omissions were identified in 7.2% and 8.9% of sentences, respectively, most commonly involving subgroup analyses, nuanced safety findings, and contextual interpretation of outcomes. Initial summary generation time decreased human effort by 95.7%, while end-to-end summarization time, including reviewer validation, decreased by 47.8%.
CONCLUSIONS: A sentence-level evaluation framework offers a robust method for identifying inaccuracies and omissions in AI-generated study summaries that are often masked by high aggregate performance scores. Through granular assessment, it enables targeted optimization and validation of AI systems, facilitating continuous improvements in output accuracy and completeness. This may strengthen confidence in AI-assisted evidence synthesis and accelerate its adoption in high-accuracy medical writing applications.

Conference/Value in Health Info

2026-11, ISPOR Europe 2026, Vienna, Austria

Value in Health, Volume 29, Issue 12S

Code

MSR77

Topic

Health Technology Assessment, Methodological & Statistical Research, Study Approaches

Topic Subcategory

Artificial Intelligence, Machine Learning, Predictive Analytics

Disease

No Additional Disease & Conditions/Specialized Treatment Areas

Your browser is out-of-date

ISPOR recommends that you update your browser for more security, speed and the best experience on ispor.org. Update my browser now

×