paper-with-me

Papers

OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup

2024-10-28 · Xize Cheng, Siqi Zheng, Zehan Wang, Minghui Fang, Ziang Zhang, Rongjie Huang, Ziyang Ma, Shengpeng Ji, Jialong Zuo, Tao Jin, Zhou Zhao

The scaling up has brought tremendous success in the fields of vision and language in recent years. When it comes to audio, however, researchers encounter a major challenge in scaling up the training data, as most natural audio contains diverse interfering signals. To address this limitation, we introduce Omni-modal Sound Separation (OmniSep), a novel framework capable of isolating clean soundtracks based on omni-modal queries, encompassing both single-modal and multi-modal composed queries. Specifically, we introduce the Query-Mixup strategy, which blends query features from different modalities during training. This enables OmniSep to optimize multiple modalities concurrently, effectively bringing all modalities under a unified framework for sound separation. We further enhance this flexibility by allowing queries to influence sound separation positively or negatively, facilitating the retention or removal of specific sounds as desired. Finally, OmniSep employs a retrieval-augmented approach known as Query-Aug, which enables open-vocabulary sound separation. Experimental evaluations on MUSIC, VGGSOUND-CLEAN+, and MUSIC-CLEAN+ datasets demonstrate effectiveness of OmniSep, achieving state-of-the-art performance in text-, image-, and audio-queried sound separation tasks. For samples and further information, please visit the demo page at \url{https://omnisep.github.io/}.

📄 PDF Abstract BibTeX arXiv:2410.21269

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Omni2Sound: Towards Unified Video-Text-to-Audio Generation

2026-01-06 · Yusheng Dai, Zehua Chen, Yuxuan Jiang, Baolong Gao 외 arxiv

Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers significant application flexibility, yet faces two unexplored foundational challenges…

Audio Generation

Ming-Omni: A Unified Multimodal Model for Perception and Generation

2025-06-11 · Inclusion AI, Biao Gong, Cheng Zou, Chuanyang Zheng 외

We propose Ming-Omni, a unified multimodal model capable of processing images, text, audio, and video, while demonstrating strong proficiency in both speech and image generation. Ming-Omni employs dedicated encoders to e…

Image Generationtext-to-speechText to Speech

Dynin-Omni: Omnimodal Unified Large Diffusion Language Model

2026-03-09 · Jaeik Kim, Woojin Kim, Jihwan Hong, Yejoon Lee 외 arxiv

We present Dynin-Omni, the first masked-diffusion-based omnimodal foundation model that unifies text, image, and speech understanding and generation, together with video understanding, within a single architecture. Unlik…

Cross-Modal RetrievalSpeech RecognitionImage Generation

Modality Unified Attack for Omni-Modality Person Re-Identification

2025-01-22 · Yuan Bian, Min Liu, Yunqi Yi, Xueping Wang 외

Deep learning based person re-identification (re-id) models have been widely employed in surveillance systems. Recent studies have demonstrated that black-box single-modality and cross-modality re-id models are vulnerabl…

Person Re-Identification

OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder

2026-05-02 · Detao Bai, Shimin Yao, Weixuan Chen, Chengen Lai 외 arxiv

Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectures rely on modality-specific encoders with a \emph{video-coarse, au…

Sign Language RecognitionComputational EfficiencySpeaker Identification