paper-with-me

홈 › Papers

How Much Temporal Modeling is Enough? A Systematic Study of Hybrid CNN-RNN Architectures for Multi-Label ECG Classification

2026-01-25 · Alireza Jafari, Fatemeh Jafari arxiv

Accurate multi-label classification of electrocardiogram (ECG) signals remains challenging due to the coexistence of multiple cardiac conditions, pronounced class imbalance, and long-range temporal dependencies in multi-lead recordings. Although recent studies increasingly rely on deep and stacked recurrent architectures, the necessity and clinical justification of such architectural complexity have not been rigorously examined. In this work, we perform a systematic comparative evaluation of convolutional neural networks (CNNs) combined with multiple recurrent configurations, including LSTM, GRU, Bidirectional LSTM (BiLSTM), and their stacked variants, for multi-label ECG classification on the PTB-XL dataset comprising 23 diagnostic categories. The CNN component serves as a morphology-driven baseline, while recurrent layers are progressively integrated to assess their contribution to temporal modeling and generalization performance. Experimental results indicate that a CNN integrated with a single BiLSTM layer achieves the most favorable trade-off between predictive performance and model complexity. This configuration attains superior Hamming loss (0.0338), macro-AUPRC (0.4715), micro-F1 score (0.6979), and subset accuracy (0.5723) compared with deeper recurrent combinations. Although stacked recurrent models occasionally improve recall for specific rare classes, our results provide empirical evidence that increasing recurrent depth yields diminishing returns and may degrade generalization due to reduced precision and overfitting. These findings suggest that architectural alignment with the intrinsic temporal structure of ECG signals, rather than increased recurrent depth, is a key determinant of robust performance and clinically relevant deployment.

📄 PDF Abstract BibTeX arXiv:2601.18830

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Label ClassificationECG Classification

Similar Papers 제목 키워드 기반

Implicit Temporal Modeling with Learnable Alignment for Video Recognition

2023-04-20 · ICCV 2023 1 · Shuyuan Tu, Qi Dai, Zuxuan Wu, Zhi-Qi Cheng 외

Contrastive language-image pretraining (CLIP) has demonstrated remarkable success in various image tasks. However, how to extend CLIP with effective temporal modeling is still an open and crucial problem. Existing factor…

Action ClassificationAction RecognitionVideo Recognition

Generalization and Feature Attribution in Machine Learning Models for Crop Yield and Anomaly Prediction in Germany

2025-12-17 · Roland Baatz arxiv

This study examines the generalization performance and interpretability of machine learning (ML) models used for predicting crop yield and yield anomalies in Germany's NUTS-3 regions. Using a high-quality, long-term data…

Feature Importance

EMoG: Synthesizing Emotive Co-speech 3D Gesture with Diffusion Model

2023-06-20 · Lianying Yin, Yijun Wang, Tianyu He, Jinming Liu 외

Although previous co-speech gesture generation methods are able to synthesize motions in line with speech content, it is still not enough to handle diverse and complicated motion distribution. The key challenges are: 1) …

DenoisingGesture Generation

Revisiting Temporal Modeling for Video-based Person ReID

2018-05-05 · Jiyang Gao, Ram Nevatia

Video-based person reID is an important task, which has received much attention in recent years due to the increasing demand in surveillance and camera networks. A typical video-based person reID system consists of three…

Mislearning from Censored Data: The Gambler's Fallacy and Other Correlational Mistakes in Optimal-Stopping Problems

2018-03-21 · Kevin He

I study endogenous learning dynamics for people who misperceive intertemporal correlations in random sequences. Biased agents face an optimal-stopping problem. They are uncertain about the underlying distribution and lea…