paper-with-me

홈 › Papers

TCNet: Continuous Sign Language Recognition from Trajectories and Correlated Regions

2024-03-18 · Hui Lu, Albert Ali Salah, Ronald Poppe

A key challenge in continuous sign language recognition (CSLR) is to efficiently capture long-range spatial interactions over time from the video input. To address this challenge, we propose TCNet, a hybrid network that effectively models spatio-temporal information from Trajectories and Correlated regions. TCNet's trajectory module transforms frames into aligned trajectories composed of continuous visual tokens. In addition, for a query token, self-attention is learned along the trajectory. As such, our network can also focus on fine-grained spatio-temporal patterns, such as finger movements, of a specific region in motion. TCNet's correlation module uses a novel dynamic attention mechanism that filters out irrelevant frame regions. Additionally, it assigns dynamic key-value tokens from correlated regions to each query. Both innovations significantly reduce the computation cost and memory. We perform experiments on four large-scale datasets: PHOENIX14, PHOENIX14-T, CSL, and CSL-Daily, respectively. Our results demonstrate that TCNet consistently achieves state-of-the-art performance. For example, we improve over the previous state-of-the-art by 1.5% and 1.0% word error rate on PHOENIX14 and PHOENIX14-T, respectively.

📄 PDF Abstract BibTeX arXiv:2403.11818

Code (1)

hotfinda/tcnet 공식 구현 pytorch

Tasks

Sign Language Recognition

Methods 이 논문이 사용한 방법론

Focus 설명 없음
CSL Circular Smooth Label (CSL) is a classification-based rotation detection technique for arbitrary-oriented object detection. It is used for circularly distributed angle…

Similar Papers 제목 키워드 기반

GM-TCNet: Gated Multi-scale Temporal Convolutional Network using Emotion Causality for Speech Emotion Recognition

2022-10-28 · Jia-Xin Ye, Xin-Cheng Wen, Xuan-Ze Wang, Yong Xu 외

In human-computer interaction, Speech Emotion Recognition (SER) plays an essential role in understanding the user's intent and improving the interactive experience. While similar sentimental speeches own diverse speaker …

Emotion RecognitionRepresentation LearningSpeech Emotion Recognition

Explainable Continuous-Time Mask Refinement with Local Self-Similarity Priors for Medical Image Segmentation

2026-02-28 · Rajdeep Chatterjee, Sudip Chakrabarty, Trishaani Acharjee arxiv

Accurate semantic segmentation of foot ulcers is essential for automated wound monitoring, yet boundary delineation remains challenging due to tissue heterogeneity and poor contrast with surrounding skin. To overcome the…

Medical Image SegmentationSemantic Segmentation

Beyond First-Order: A Multi-Scale Approach to Finger Knuckle Print Biometrics

2024-06-28 · Chengrui Gao, Ziyuan Yang, Andrew Beng Jin Teoh, Min Zhu

Recently, finger knuckle prints (FKPs) have gained attention due to their rich textural patterns, positioning them as a promising biometric for identity recognition. Prior FKP recognition methods predominantly leverage f…

BiTimeCrossNet: Time-Aware Self-Supervised Learning for Pediatric Sleep

2026-02-02 · Saurav Raj Pandey, Harlin Lee arxiv

We present BiTimeCrossNet (BTCNet), a multimodal self-supervised learning framework for long physiological recordings such as overnight sleep studies. While many existing approaches train on short segments treated as ind…

Self-Supervised Learning

MTCNet: Motion and Topology Consistency Guided Learning for Mitral Valve Segmentationin 4D Ultrasound

2025-07-01 · Rusi Chen, Yuanting Yang, Jiezhi Yao, Hongning Song 외 arxiv

Mitral regurgitation is one of the most prevalent cardiac disorders. Four-dimensional (4D) ultrasound has emerged as the primary imaging modality for assessing dynamic valvular morphology. However, 4D mitral valve (MV) a…