paper-with-me

Papers

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation

2025-10-10 · Akira Takahashi, Shusuke Takahashi, Yuki Mitsufuji arxiv

We introduce MMAudioSep, a generative model for video/text-queried sound separation that is founded on a pretrained video-to-audio model. By leveraging knowledge about the relationship between video/text and audio learned through a pretrained audio generative model, we can train the model more efficiently, i.e., the model does not need to be trained from scratch. We evaluate the performance of MMAudioSep by comparing it to existing separation models, including models based on both deterministic and generative approaches, and find it is superior to the baseline models. Furthermore, we demonstrate that even after acquiring functionality for sound separation via fine-tuning, the model retains the ability for original video-to-audio generation. This highlights the potential of foundational sound generation models to be adopted for sound-related downstream tasks. Our code is available at https://github.com/sony/mmaudiosep.

📄 PDF Abstract BibTeX arXiv:2510.09065

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Generation

Similar Papers 제목 키워드 기반

MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

2024-12-19 · CVPR 2025 1 · Ho Kei Cheng, Masato Ishii, Akio Hayakawa, Takashi Shibuya 외

We propose to synthesize high-quality and synchronized audio, given video and optional text conditions, using a novel multimodal joint training framework MMAudio. In contrast to single-modality training conditioned on (l…

Audio GenerationAudio SynthesisAudio-Visual SynchronizationVideo-to-Sound Generation

Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

2024-09-04 · Jianwen Jiang, Chao Liang, Jiaqi Yang, Gaojie Lin 외

With the introduction of diffusion-based video generation techniques, audio-conditioned human video generation has recently achieved significant breakthroughs in both the naturalness of motion and the synthesis of portra…

Video Generation

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

2025-10-03 · Kaisi Guan, Xihua Wang, Zhengfeng Lai, Xin Cheng 외 arxiv

This study focuses on a challenging yet promising task, Text-to-Sounding-Video (T2SV) generation, which aims to generate a video with synchronized audio from text conditions, meanwhile ensuring both modalities are aligne…

Video Generation

Taming Transformer for Emotion-Controllable Talking Face Generation

2025-08-20 · Ziqi Zhang, Cheng Deng arxiv

Talking face generation is a novel and challenging generation task, aiming at synthesizing a vivid speaking-face video given a specific audio. To fulfill emotion-controllable talking face generation, current methods need…

Talking Face Generation

Hiding Video in Audio via Reversible Generative Models

2019-10-01 · ICCV 2019 10 · Hyukryul Yang, Hao Ouyang, Vladlen Koltun, Qifeng Chen

We present a method for hiding video content inside audio files while preserving the perceptual fidelity of the cover audio. This is a form of cross-modal steganography and is particularly challenging due to the high bit…