paper-with-me

홈 › Papers

Aligning Subtitles in Sign Language Videos

2021-05-06 · ICCV 2021 10 · Hannah Bull, Triantafyllos Afouras, Gül Varol, Samuel Albanie, Liliane Momeni, Andrew Zisserman

The goal of this work is to temporally align asynchronous subtitles in sign language videos. In particular, we focus on sign-language interpreted TV broadcast data comprising (i) a video of continuous signing, and (ii) subtitles corresponding to the audio content. Previous work exploiting such weakly-aligned data only considered finding keyword-sign correspondences, whereas we aim to localise a complete subtitle text in continuous signing. We propose a Transformer architecture tailored for this task, which we train on manually annotated alignments covering over 15K subtitles that span 17.7 hours of video. We use BERT subtitle embeddings and CNN video representations learned for sign recognition to encode the two signals, which interact through a series of attention layers. Our model outputs frame-level predictions, i.e., for each video frame, whether it belongs to the queried subtitle or not. Through extensive evaluations, we show substantial improvements over existing alignment baselines that do not make use of subtitle text embeddings for learning. Our automatic alignment model opens up possibilities for advancing machine translation of sign languages via providing continuously synchronized video-text data.

📄 PDF Abstract BibTeX arXiv:2105.02877

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing

2025-12-08 · Zifan Jiang, Youngjoon Jang, Liliane Momeni, Gül Varol 외 arxiv

The goal of this work is to develop a universal approach for aligning subtitles (i.e., spoken language text with corresponding timestamps) to continuous sign language videos. Prior approaches typically rely on end-to-end…

MEDIAPI-SKEL - A 2D-Skeleton Video Database of French Sign Language With Aligned French Subtitles

2020-05-01 · LREC 2020 5 · Hannah Bull, Annelies Braffort, Mich{\`e}le Gouiff{\`e}s

This paper presents MEDIAPI-SKEL, a 2D-skeleton database of French Sign Language videos aligned with French subtitles. The corpus contains 27 hours of video of body, face and hand keypoints, aligned to subtitles with a v…

Cross-Modal RetrievalRetrievalSemantic SegmentationVideo Semantic Segmentation

HowToCaption: Prompting LLMs to Transform Video Annotations at Scale

2023-10-07 · Nina Shvetsova, Anna Kukleva, Xudong Hong, Christian Rupprecht 외

Instructional videos are a common source for learning text-video or even multimodal representations by leveraging subtitles extracted with automatic speech recognition systems (ASR) from the audio signal in the videos. H…

Automatic Speech RecognitionVideo CaptioningVideo RetrievalZero-Shot Video-Audio Retrieval+1

Read and Attend: Temporal Localisation in Sign Language Videos

2021-03-30 · CVPR 2021 1 · Gül Varol, Liliane Momeni, Samuel Albanie, Triantafyllos Afouras 외

The objective of this work is to annotate sign instances across a broad vocabulary in continuous sign language. We train a Transformer model to ingest a continuous signing stream and output a sequence of written tokens o…

Sign Language Recognition

Visual Subtitles for Internet Videos

2013-08-01 · WS 2013 8 · Chitralekha Bhat, Imran Ahmed, Vikram Saxena, Sunil Kumar Kopparapu
Speech Recognition