paper-with-me

Papers

Mask as Supervision: Leveraging Unified Mask Information for Unsupervised 3D Pose Estimation

2023-12-12 · Yuchen Yang, Yu Qiao, Xiao Sun

Automatic estimation of 3D human pose from monocular RGB images is a challenging and unsolved problem in computer vision. In a supervised manner, approaches heavily rely on laborious annotations and present hampered generalization ability due to the limited diversity of 3D pose datasets. To address these challenges, we propose a unified framework that leverages mask as supervision for unsupervised 3D pose estimation. With general unsupervised segmentation algorithms, the proposed model employs skeleton and physique representations that exploit accurate pose information from coarse to fine. Compared with previous unsupervised approaches, we organize the human skeleton in a fully unsupervised way which enables the processing of annotation-free data and provides ready-to-use estimation results. Comprehensive experiments demonstrate our state-of-the-art pose estimation performance on Human3.6M and MPI-INF-3DHP datasets. Further experiments on in-the-wild datasets also illustrate the capability to access more data to boost our model. Code will be available at https://github.com/Charrrrrlie/Mask-as-Supervision.

📄 PDF Abstract BibTeX arXiv:2312.07051

Code (1)

charrrrrlie/mask-as-supervision 공식 구현

Tasks

3D Pose EstimationDiversityPose EstimationUnsupervised 3D Human Pose Estimation

Similar Papers 제목 키워드 기반

Mask-free OVIS: Open-Vocabulary Instance Segmentation without Manual Mask Annotations

2023-03-29 · CVPR 2023 1 · Vibashan VS, Ning Yu, Chen Xing, Can Qin 외

Existing instance segmentation models learn task-specific information using manual mask annotations from base (training) categories. These mask annotations require tremendous human effort, limiting the scalability to ann…

Image CaptioningInstance SegmentationLanguage ModelingLanguage Modelling+1

Learning Segmentation from Radiology Reports

2025-07-08 · Pedro R. A. S. Bassi, Wenxuan Li, Jieneng Chen, Zheren Zhu 외 arxiv

Tumor segmentation in CT scans is key for diagnosis, surgery, and prognosis, yet segmentation masks are scarce because their creation requires time and expertise. Public abdominal CT datasets have from dozens to a couple…

Tumor Segmentation

Text-Guided Video Masked Autoencoder

2024-08-01

Recent video masked autoencoder (MAE) works have designed improved masking algorithms focused on saliency. These works leverage visual cues such as motion to mask the most salient regions. However, the robustness of such…

Leveraging Multi-Modal Information to Enhance Dataset Distillation

2025-05-13 · Zhe Li, Hadrien Reynaud, Bernhard Kainz

Dataset distillation aims to create a compact and highly representative synthetic dataset that preserves the knowledge of a larger real dataset. While existing methods primarily focus on optimizing visual representations…

Dataset DistillationObject

META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR

2024-09-18 · Jinhan Wang, Weiqing Wang, Kunal Dhawan, Taejin Park 외

We propose a novel end-to-end multi-talker automatic speech recognition (ASR) framework that enables both multi-speaker (MS) ASR and target-speaker (TS) ASR. Our proposed model is trained in a fully end-to-end manner, in…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+2