paper-with-me

홈 › Papers

Multi-Scale Local-Temporal Similarity Fusion for Continuous Sign Language Recognition

2021-07-27 · Pan Xie, Zhi Cui, Yao Du, Mengyi Zhao, Jianwei Cui, Bin Wang, Xiaohui Hu

Continuous sign language recognition (cSLR) is a public significant task that transcribes a sign language video into an ordered gloss sequence. It is important to capture the fine-grained gloss-level details, since there is no explicit alignment between sign video frames and the corresponding glosses. Among the past works, one promising way is to adopt a one-dimensional convolutional network (1D-CNN) to temporally fuse the sequential frames. However, CNNs are agnostic to similarity or dissimilarity, and thus are unable to capture local consistent semantics within temporally neighboring frames. To address the issue, we propose to adaptively fuse local features via temporal similarity for this task. Specifically, we devise a Multi-scale Local-Temporal Similarity Fusion Network (mLTSF-Net) as follows: 1) In terms of a specific video frame, we firstly select its similar neighbours with multi-scale receptive regions to accommodate different lengths of glosses. 2) To ensure temporal consistency, we then use position-aware convolution to temporally convolve each scale of selected frames. 3) To obtain a local-temporally enhanced frame-wise representation, we finally fuse the results of different scales using a content-dependent aggregator. We train our model in an end-to-end fashion, and the experimental results on RWTH-PHOENIX-Weather 2014 datasets (RWTH) demonstrate that our model achieves competitive performance compared with several state-of-the-art models.

📄 PDF Abstract BibTeX arXiv:2107.12762

Code (0)

등록된 구현이 없습니다.

Tasks

Sign Language Recognition

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Repetitive Action Counting with Hybrid Temporal Relation Modeling

2024-12-10 · Kun Li, Xinge Peng, Dan Guo, Xun Yang 외

Repetitive Action Counting (RAC) aims to count the number of repetitive actions occurring in videos. In the real world, repetitive actions have great diversity and bring numerous challenges (e.g., viewpoint changes, non-…

RelationRepetitive Action Counting

ATSS: Detecting AI-Generated Videos via Anomalous Temporal Self-Similarity

2026-04-05 · Hang Wang, Chao Shen, Lei Zhang, Zhi-Qi Cheng arxiv

AI-generated videos (AIGVs) have achieved unprecedented photorealism, posing severe threats to digital forensics. Existing AIGV detectors focus mainly on localized artifacts or short-term temporal inconsistencies, thus o…

Video Generation

Multiscale Fusion for Abnormality Detection and Localization of Distributed Parameter Systems

2023-10-10 · Peng Wei, Han-Xiong Li

Numerous industrial thermal processes and fluid processes can be described by distributed parameter systems (DPSs), wherein many process parameters and variables vary in space and time. Early internal abnormalities in th…

Anomaly DetectionFault Detection

Local Spatiotemporal Representation Learning for Longitudinally-consistent Neuroimage Analysis

2022-06-09 · Mengwei Ren, Neel Dey, Martin A. Styner, Kelly Botteron 외

Recent self-supervised advances in medical computer vision exploit global and local anatomical self-similarity for pretraining prior to downstream tasks such as segmentation. However, current methods assume i.i.d. image …

One-Shot SegmentationRepresentation LearningSegmentation

MFF-EINV2: Multi-scale Feature Fusion across Spectral-Spatial-Temporal Domains for Sound Event Localization and Detection

2024-06-13 · Da Mu, Zhicheng Zhang, Haobo Yue

Sound Event Localization and Detection (SELD) involves detecting and localizing sound events using multichannel sound recordings. Previously proposed Event-Independent Network V2 (EINV2) has achieved outstanding performa…

Sound Event Localization and Detection