paper-with-me

홈 › Papers

ProgressLabeller: Visual Data Stream Annotation for Training Object-Centric 3D Perception

2022-03-01 · Xiaotong Chen, Huijie Zhang, Zeren Yu, Stanley Lewis, Odest Chadwicke Jenkins

Visual perception tasks often require vast amounts of labelled data, including 3D poses and image space segmentation masks. The process of creating such training data sets can prove difficult or time-intensive to scale up to efficacy for general use. Consider the task of pose estimation for rigid objects. Deep neural network based approaches have shown good performance when trained on large, public datasets. However, adapting these networks for other novel objects, or fine-tuning existing models for different environments, requires significant time investment to generate newly labelled instances. Towards this end, we propose ProgressLabeller as a method for more efficiently generating large amounts of 6D pose training data from color images sequences for custom scenes in a scalable manner. ProgressLabeller is intended to also support transparent or translucent objects, for which the previous methods based on depth dense reconstruction will fail. We demonstrate the effectiveness of ProgressLabeller by rapidly create a dataset of over 1M samples with which we fine-tune a state-of-the-art pose estimation network in order to markedly improve the downstream robotic grasp success rates. ProgressLabeller is open-source at https://github.com/huijieZH/ProgressLabeller.

📄 PDF Abstract BibTeX arXiv:2203.00283

Code (1)

huijiezh/progresslabeller 공식 구현

Tasks

Pose Estimation

Similar Papers 제목 키워드 기반

Contrastive Pretraining with Dual Visual Encoders for Gloss-Free Sign Language Translation

2025-07-14 · Ozge Mercanoglu Sincan, Richard Bowden arxiv

Sign Language Translation (SLT) aims to convert sign language videos into spoken or written text. While early systems relied on gloss annotations as an intermediate supervision, such annotations are costly to obtain and …

Sign Language Translation

Pre-trained Visual Dynamics Representations for Efficient Policy Learning

2024-11-05 · Hao Luo, Bohan Zhou, Zongqing Lu

Pre-training for Reinforcement Learning (RL) with purely video data is a valuable yet challenging problem. Although in-the-wild videos are readily available and inhere a vast amount of prior world knowledge, the absence …

Reinforcement Learning (RL)Video PredictionWorld Knowledge

VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models

2026-04-08 · Pavan Kumar Anasosalu Vasu, Cem Koc, Fartash Faghri, Chun-Liang Li 외 arxiv

Streaming vision-language models (VLMs) continuously generate responses given an instruction prompt and an online stream of input frames. This is a core mechanism for real-time visual assistants. Existing VLM frameworks …

Question Answering

LocTex: Learning Data-Efficient Visual Representations from Localized Textual Supervision

2021-08-26 · ICCV 2021 10 · Zhijian Liu, Simon Stent, Jie Li, John Gideon 외

Computer vision tasks such as object detection and semantic/instance segmentation rely on the painstaking annotation of large training datasets. In this paper, we propose LocTex that takes advantage of the low-cost local…

image-classificationImage ClassificationInstance Segmentationobject-detection+2

PixelWeb: The First Web GUI Dataset with Pixel-Wise Labels

2025-04-23 · Qi Yang, Weichen Bi, Haiyang Shen, Yaoqi Guo 외

Graphical User Interface (GUI) datasets are crucial for various downstream tasks. However, GUI datasets often generate annotation information through automatic labeling, which commonly results in inaccurate GUI element B…

GUI Element Detection