paper-with-me

Papers

Spatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos

2026-06-16 · Bo Gou, Jicheng Zhang, Jianlong Xiong, Tao He, Bentian Liu, Hai Wu, Yijiao Wang, Yu Zhang, Yujia Yang, Yun Dai, Jian Liu, Jie Wang arxiv

Automated classification of standard echocardiographic views is crucial for efficient clinical workflow but faces three main challenges. First, publicly available datasets are scarce and limited in scale and view coverage. Second, the performance of some modern video-level architectures for echocardiographic view classification remains underexplored. Third, some view categories exhibit highly similar spatial appearances, making single-frame features insufficient for discrimination, while heterogeneous frame quality complicates robust temporal information fusion. To address these challenges, we release the Echocardiographic Videos of Nine Views (EV9V) dataset, comprising 5,138 videos, 910,579 frames, and 9 standard views, which is, to the best of our knowledge, the largest publicly available echocardiography video dataset. Using EV9V, we systematically benchmark representative video classification architectures, including Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and Transformers. Furthermore, we propose a Spatio-Temporal Fusion Model (STFM), an efficient dual-stream CNN-LSTM (Long Short-Term Memory) framework that jointly captures spatial anatomical structures and temporal cardiac dynamics. The proposed framework leverages uncertainty-aware learning to preferentially sample representative video segments during training and evidence-based fusion during inference, improving robustness to variations in frame quality across echocardiographic videos. Extensive experiments demonstrate that our method achieves competitive performance across diverse video classification models, validating the effectiveness of uncertainty-aware spatio-temporal learning for echocardiographic view classification. The code is available at https://github.com/bgx666/stfm.

📄 PDF Abstract BibTeX arXiv:2606.17437

Code (0)

등록된 구현이 없습니다.

Tasks

Video Classification

Similar Papers 제목 키워드 기반

ConvFormer3D-TAP: Phase/Uncertainty-Aware Front-End Fusion for Cine CMR View Classification Pipelines

2026-04-13 · Nafiseh Ghaffar Nia, Vinesh Appadurai, Suchithra V., Chinmay Rane 외 arxiv

Reliable recognition of standard cine cardiac MRI views is essential because each view determines which cardiac anatomy is visualized and which quantitative analyses can be performed. Incorrect view identification, wheth…

A Decade of Deep Learning for Remote Sensing Spatiotemporal Fusion: Advances, Challenges, and Opportunities

2025-04-01 · Enzhe Sun, Yongchuan Cui, Peng Liu, Jining Yan

Hardware limitations and satellite launch costs make direct acquisition of high temporal-spatial resolution remote sensing imagery challenging. Remote sensing spatiotemporal fusion (STF) technology addresses this problem…

A Survey on Diffusion Models for Time Series and Spatio-Temporal Data

2024-04-29 · Yiyuan Yang, Ming Jin, Haomin Wen, Chaoli Zhang 외

The study of time series is crucial for understanding trends and anomalies over time, enabling predictive insights across various sectors. Spatio-temporal data, on the other hand, is vital for analyzing phenomena in both…

Anomaly DetectionImputationTime Series

AnyView: Synthesizing Any Novel View in Dynamic Scenes

2026-01-23 · Basile Van Hoorick, Dian Chen, Shun Iwase, Pavel Tokmakov 외 arxiv

Modern generative video models excel at producing convincing, high-quality outputs, but struggle to maintain multi-view and spatiotemporal consistency in highly dynamic real-world environments. In this work, we introduce…

Video Generation

STDD: Spatio-Temporal Dual Diffusion for Video Generation

2025-01-01 · CVPR 2025 1 · Shuaizhen Yao, Xiaoya Zhang, Xin Liu, Mengyi Liu 외

Diffusion probabilistic model is becoming the cornerstone of data generation, especially generating high-quality images. As an extension, video diffusion generation is in urgent need of a principled temporal-sequence…

Text-to-Video GenerationVideo Generation