ARTIFICIAL INTELLIGENCE-BASED AUTOMATED ASSESSMENT OF THORACENTESIS SKILLS: DEVELOPMENT AND REAL-WORLD EVALUATION
Author(s)
Yu Qing, MD1, Min Zhang2, Zhang Wen, MD1.
1Zhongshan Hospital , Fudan University, shanghai, China, 2assistant researcher, Zhongshan Hospital , Fudan University, SHANGHAI, China.
1Zhongshan Hospital , Fudan University, shanghai, China, 2assistant researcher, Zhongshan Hospital , Fudan University, SHANGHAI, China.
OBJECTIVES: Thoracentesis is an essential clinical procedure requiring standardized training and assessment. Conventional evaluation methods depend heavily on expert supervision and are associated with substantial time and personnel requirements. This study aimed to develop and evaluate an artificial intelligence (AI)-based automated assessment system for thoracentesis skills and explore its potential to improve training efficiency and optimize resource utilization in medical education.
METHODS: A temporal medical action knowledge graph was constructed to define 12 procedural stages, 39 sub-actions, and 47 potential error actions in thoracentesis. Video data were collected from 9 volunteers using dual first-person cameras, generating a dataset comprising approximately 740,000 frames (6.9 hours; 52.1 GB). A temporal action segmentation framework integrating temporal clustering attention and diffusion mechanisms was developed to recognize procedural actions and detect errors. The dataset was divided into training and testing sets at a ratio of 60:40. Performance was evaluated using frame-wise accuracy (Acc), Edit Score, and F1 scores at overlap thresholds of 10%, 25%, and 50%. Real-world deployment was conducted with two medical interns.
RESULTS: The proposed framework achieved 73.51% Acc, 76.08% Edit Score, and F1 scores of 78.36%, 72.79%, and 61.54% at overlap thresholds of 10%, 25%, and 50%, respectively. Incorrect action recognition accuracy reached 65.73% during deployment. The system enabled automated, real-time assessment of procedural performance and reduced dependence on continuous expert observation. Standardized evaluation and timely feedback improved training efficiency and demonstrated potential to reduce personnel requirements and resource utilization.
CONCLUSIONS: The AI-based thoracentesis assessment system demonstrated promising performance in procedural action recognition and error detection. By supporting automated evaluation and reducing expert supervision requirements, the system may improve training efficiency, optimize educational resource allocation, and provide a scalable approach for procedural skills assessment in medical education.
METHODS: A temporal medical action knowledge graph was constructed to define 12 procedural stages, 39 sub-actions, and 47 potential error actions in thoracentesis. Video data were collected from 9 volunteers using dual first-person cameras, generating a dataset comprising approximately 740,000 frames (6.9 hours; 52.1 GB). A temporal action segmentation framework integrating temporal clustering attention and diffusion mechanisms was developed to recognize procedural actions and detect errors. The dataset was divided into training and testing sets at a ratio of 60:40. Performance was evaluated using frame-wise accuracy (Acc), Edit Score, and F1 scores at overlap thresholds of 10%, 25%, and 50%. Real-world deployment was conducted with two medical interns.
RESULTS: The proposed framework achieved 73.51% Acc, 76.08% Edit Score, and F1 scores of 78.36%, 72.79%, and 61.54% at overlap thresholds of 10%, 25%, and 50%, respectively. Incorrect action recognition accuracy reached 65.73% during deployment. The system enabled automated, real-time assessment of procedural performance and reduced dependence on continuous expert observation. Standardized evaluation and timely feedback improved training efficiency and demonstrated potential to reduce personnel requirements and resource utilization.
CONCLUSIONS: The AI-based thoracentesis assessment system demonstrated promising performance in procedural action recognition and error detection. By supporting automated evaluation and reducing expert supervision requirements, the system may improve training efficiency, optimize educational resource allocation, and provide a scalable approach for procedural skills assessment in medical education.
Conference/Value in Health Info
2026-09, ISPOR Asia Pacific 2026, Bangkok, Thailand
Value in Health, Volume 55, Issue S1
Code
CO12
Topic
Clinical Outcomes
Topic Subcategory
Comparative Effectiveness or Efficacy
Disease
No Additional Disease & Conditions/Specialized Treatment Areas