paper-with-me

홈 › Papers

Automatic dense annotation of large-vocabulary sign language videos

2022-08-04 · Liliane Momeni, Hannah Bull, K R Prajwal, Samuel Albanie, Gül Varol, Andrew Zisserman

Recently, sign language researchers have turned to sign language interpreted TV broadcasts, comprising (i) a video of continuous signing and (ii) subtitles corresponding to the audio content, as a readily available and large-scale source of training data. One key challenge in the usability of such data is the lack of sign annotations. Previous work exploiting such weakly-aligned data only found sparse correspondences between keywords in the subtitle and individual signs. In this work, we propose a simple, scalable framework to vastly increase the density of automatic annotations. Our contributions are the following: (1) we significantly improve previous annotation methods by making use of synonyms and subtitle-signing alignment; (2) we show the value of pseudo-labelling from a sign recognition model as a way of sign spotting; (3) we propose a novel approach for increasing our annotations of known and unknown classes based on in-domain exemplars; (4) on the BOBSL BSL sign language corpus, we increase the number of confident automatic annotations from 670K to 5M. We make these annotations publicly available to support the sign language research community.

📄 PDF Abstract BibTeX arXiv:2208.02802

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OpenAnnotate3D: Open-Vocabulary Auto-Labeling System for Multi-modal 3D Data

2023-10-20 · Yijie Zhou, Likun Cai, Xianhui Cheng, Zhongxue Gan 외

In the era of big data and large models, automatic annotating functions for multi-modal data are of great significance for real-world AI-driven applications, such as autonomous driving and embodied AI. Unlike traditional…

Autonomous Driving

Open-Vocabulary Camouflaged Object Segmentation

2023-11-19 · Youwei Pang, Xiaoqi Zhao, Jiaming Zuo, Lihe Zhang 외

Recently, the emergence of the large-scale vision-language model (VLM), such as CLIP, has opened the way towards open-world object perception. Many works have explored the utilization of pre-trained VLM for the challengi…

Camouflaged Object SegmentationImage SegmentationLanguage ModellingObject+2

GenCAMO: Scene-Graph Contextual Decoupling for Environment-aware and Mask-free Camouflage Image-Dense Annotation Generation

2026-01-03 · Chenglizhao Chen, Shaojiang Yuan, Xiaoxue Lu, Mengke Song 외 arxiv

Conceal dense prediction (CDP), especially RGB-D camouflage object detection and open-vocabulary camouflage object segmentation, plays a crucial role in advancing the understanding and reasoning of complex camouflage sce…

Object SegmentationObject Detection

FreeOcc: Training-Free Embodied Open-Vocabulary Occupancy Prediction

2026-04-30 · Zeyu Jiang, Changqing Zhou, Xingxing Zuo, Changhao Chen arxiv

Existing learning-based occupancy prediction methods rely on large-scale 3D annotations and generalize poorly across environments. We present FreeOcc, a training-free framework for open-vocabulary occupancy prediction fr…

Natural Vocabulary Emerges from Free-Form Annotations

2019-06-04 · Jordi Pont-Tuset, Michael Gygli, Vittorio Ferrari

We propose an approach for annotating object classes using free-form text written by undirected and untrained annotators. Free-form labeling is natural for annotators, they intuitively provide very specific and exhaustiv…

Form