paper-with-me

홈 › Papers

Selective Contrastive Learning For Gloss Free Sign Language Translation

2026-04-24 · Changhao Lai, Rui Zhao, Xuewen Zhong, Jinsong Su, Yidong Chen arxiv

Sign language translation (SLT) converts continuous sign videos into spoken-language text, yet it remains challenging due to the intrinsic modality mismatch between visual signs and written text, particularly in gloss-free settings. Recent SLT systems increasingly adopt CLIP-like Vision-Language pretraining (VLP) for cross-modal alignment, but the random in-batch contrast provides few, batch-dependent negatives and may mislabel semantically similar (or even identical) pairs as negatives, introducing noisy and potentially inconsistent alignment supervision. In this work, we first conduct a preliminary trajectory-based analysis that tracks negative video-text similarity over training. The results show that only a small subset of negatives exhibits the desired behavior of being consistently pushed away, while the remaining negatives display heterogeneous and often non-decreasing similarity dynamics, suggesting that random in-batch negatives are frequently uninformative for effective alignment. Inspired by this, we propose Selective Contrastive Learning for SLT (SCL-SLT) with a Pair Selection (PS) strategy. PS scores candidate negatives using similarity dynamics from reference checkpoints and constructs mini-batches via a curriculum that progressively emphasizes more challenging negatives, thereby strengthening contrastive supervision while reducing the influence of noisy or semantically invalid negatives.

📄 PDF Abstract BibTeX arXiv:2604.22374

Code (0)

등록된 구현이 없습니다.

Tasks

Sign Language TranslationContrastive Learning

Similar Papers 제목 키워드 기반

Contrastive Pretraining with Dual Visual Encoders for Gloss-Free Sign Language Translation

2025-07-14 · Ozge Mercanoglu Sincan, Richard Bowden arxiv

Sign Language Translation (SLT) aims to convert sign language videos into spoken or written text. While early systems relied on gloss annotations as an intermediate supervision, such annotations are costly to obtain and …

Sign Language Translation

Hierarchical Feature Alignment for Gloss-Free Sign Language Translation

2025-07-09 · Sobhan Asasi, Mohamed Ilyes Lakhal, Richard Bowden arxiv

Sign Language Translation (SLT) attempts to convert sign language videos into spoken sentences. However, many existing methods struggle with the disparity between visual and textual representations during end-to-end lear…

Sign Language Translation

Beyond Gloss: A Hand-Centric Framework for Gloss-Free Sign Language Translation

2025-07-31 · Sobhan Asasi, Mohamed Ilyas Lakhal, Ozge Mercanoglu Sincan, Richard Bowden arxiv

Sign Language Translation (SLT) is a challenging task that requires bridging the modality gap between visual and linguistic information while capturing subtle variations in hand shapes and movements. To address these cha…

Sign Language Translation

Gloss-Free End-to-End Sign Language Translation

2023-05-22 · Kezhou Lin, Xiaohan Wang, Linchao Zhu, Ke Sun 외

In this paper, we tackle the problem of sign language translation (SLT) without gloss annotations. Although intermediate representation like gloss has been proven effective, gloss annotations are hard to acquire, especia…

Gloss-free Sign Language TranslationSign Language TranslationTranslation

Improving Gloss-free Sign Language Translation by Reducing Representation Density

2024-05-23 · Jinhui Ye, Xing Wang, Wenxiang Jiao, Junwei Liang 외

Gloss-free sign language translation (SLT) aims to develop well-performing SLT systems with no requirement for the costly gloss annotations, but currently still lags behind gloss-based approaches significantly. In this p…

Contrastive LearningGloss-free Sign Language TranslationSign Language TranslationTranslation