Temporal HeartNet: Towards Human-Level Automatic Analysis of Fetal Cardiac Screening Video
We present an automatic method to describe clinically useful information about scanning, and to guide image interpretation in ultrasound (US) videos of the fetal heart. Our method is able to jointly predict the visibility, viewing plane, location and orientation of the fetal heart at the frame level. The contributions of the paper are three-fold: (i) a convolutional neural network architecture is developed for a multi-task prediction, which is computed by sliding a 3x3 window spatially through convolutional maps. (ii) an anchor mechanism and Intersection over Union (IoU) loss are applied for improving localization accuracy. (iii) a recurrent architecture is designed to recursively compute regional convolutional features temporally over sequential frames, allowing each prediction to be conditioned on the whole video. This results in a spatial-temporal model that precisely describes detailed heart parameters in challenging US videos. We report results on a real-world clinical dataset, where our method achieves performance on par with expert annotations.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Bodily Behaviors in Social Interaction: Novel Annotations and State-of-the-Art Evaluation
Body language is an eye-catching social signal and its automatic analysis can significantly advance artificial intelligence systems to understand and actively participate in social interactions. While computer vision has…
Action DetectionDescriptivePose EstimationHuman vs Automatic Metrics: on the Importance of Correlation Design
This paper discusses two existing approaches to the correlation analysis between automatic evaluation metrics and human scores in the area of natural language generation. Our experiments show that depending on the usage …
SentenceText GenerationRepresenting Videos Using Mid-level Discriminative Patches
representation for videos based on mid-level discriminative spatio-temporal patches. These spatio-temporal patches might correspond to a primitive human action, a semantic object, or perhaps a random but informative spat…
Action ClassificationGeneral ClassificationA Deep Learning Based Automatic Defect Analysis Framework for In-situ TEM Ion Irradiations
Videos captured using Transmission Electron Microscopy (TEM) can encode details regarding the morphological and temporal evolution of a material by taking snapshots of the microstructure sequentially. However, manual ana…
Defect Detectionobject-detectionObject DetectionTwo-stage Temporal Modelling Framework for Video-based Depression Recognition using Graph Representation
Video-based automatic depression analysis provides a fast, objective and repeatable self-assessment solution, which has been widely developed in recent years. While depression clues may be reflected by human facial behav…