paper-with-me

Papers

Leveraging Category Information for Single-Frame Visual Sound Source Separation

2020-07-15 · Lingyu Zhu, Esa Rahtu

Visual sound source separation aims at identifying sound components from a given sound mixture with the presence of visual cues. Prior works have demonstrated impressive results, but with the expense of large multi-stage architectures and complex data representations (e.g. optical flow trajectories). In contrast, we study simple yet efficient models for visual sound separation using only a single video frame. Furthermore, our models are able to exploit the information of the sound source category in the separation process. To this end, we propose two models where we assume that i) the category labels are available at the training time, or ii) we know if the training sample pairs are from the same or different category. The experiments with the MUSIC dataset show that our model obtains comparable or better performance compared to several recent baseline methods. The code is available at https://github.com/ly-zhu/Leveraging-Category-Information-for-Single-Frame-Visual-Sound-Source-Separation

📄 PDF Abstract BibTeX arXiv:2007.07984

Code (3)

ly-zhu/Leveraging-Category-Information-for-Single-Frame-Visual-Sound-Source-Separation 공식 구현 pytorch
ly-zhu/Separating-Sounds-from-a-Single-Image pytorch
ly-zhu/ly-zhu.github.io

Tasks

Optical Flow Estimation

Similar Papers 제목 키워드 기반

Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery

2024-03-15 · Enguang Wang, Zhimao Peng, Zhengyuan Xie, Fei Yang 외

Given unlabelled datasets containing both old and new categories, generalized category discovery (GCD) aims to accurately discover new classes while correctly classifying old classes, leveraging the class concepts learne…

Generalized Fine-Grained Category Discovery with Multi-Granularity Conceptual Experts

2025-09-30 · Haiyang Zheng, Nan Pu, Wenjing Li, Nicu Sebe 외 arxiv

Generalized Category Discovery (GCD) is an open-world problem that clusters unlabeled data by leveraging knowledge from partially labeled categories. A key challenge is that unlabeled data may contain both known and nove…

Fine-Grained Visual RecognitionRepresentation LearningContrastive Learning

Leveraging SE(3) Equivariance for Self-Supervised Category-Level Object Pose Estimation

2021-10-30 · NeurIPS 2021 12 · Xiaolong Li, Yijia Weng, Li Yi, Leonidas Guibas 외

Category-level object pose estimation aims to find 6D object poses of previously unseen object instances from known categories without access to object CAD models. To reduce the huge amount of pose annotations needed for…

ObjectPose EstimationSelf-Supervised Learning

Leveraging SE(3) Equivariance for Self-supervised Category-Level Object Pose Estimation from Point Clouds

2021-05-21 · NeurIPS 2021 12 · Xiaolong Li, Yijia Weng, Li Yi, Leonidas Guibas 외

Category-level object pose estimation aims to find 6D object poses of previously unseen object instances from known categories without access to object CAD models. To reduce the huge amount of pose annotations needed fo…

ObjectPose EstimationSelf-Supervised Learning

Learning Semantic-Aware Representation in Visual-Language Models for Multi-Label Recognition with Partial Labels

2024-12-14 · Haoxian Ruan, Zhihua Xu, Zhijing Yang, Yongyi Lu 외

Multi-label recognition with partial labels (MLR-PL), in which only some labels are known while others are unknown for each image, is a practical task in computer vision, since collecting large-scale and complete multi-l…