paper-with-me

Papers

DeepSound-V1: Start to Think Step-by-Step in the Audio Generation from Videos

2025-03-28 · Yunming Liang, Zihao Chen, Chaofan Ding, Xinhan Di

Currently, high-quality, synchronized audio is synthesized from video and optional text inputs using various multi-modal joint learning frameworks. However, the precise alignment between the visual and generated audio domains remains far from satisfactory. One key factor is the lack of sufficient temporal and semantic alignment annotations in open-source video-audio and text-audio benchmarks. Therefore, we propose a framework for audio generation from videos, leveraging the internal chain-of-thought (CoT) of a multi-modal large language model (MLLM) to enable step-by-step reasoning without requiring additional annotations. Additionally, a corresponding multi-modal reasoning dataset is constructed to facilitate the learning of initial reasoning in audio generation. In the experiments, we demonstrate the effectiveness of the proposed framework in reducing misalignment (voice-over) in generated audio and achieving competitive performance compared to various state-of-the-art models. The evaluation results show that the proposed method outperforms state-of-the-art approaches across multiple metrics. Specifically, the F DP aSST indicator is reduced by up to 10.07%, the F DP AN N s indicator by up to 11.62%, and the F DV GG indicator by up to 38.61%. Furthermore, the IS indicator improves by up to 4.95%, the IB-score indicator increases by up to 6.39%, and the DeSync indicator is reduced by up to 0.89%.

📄 PDF Abstract BibTeX arXiv:2503.22208

Code (0)

등록된 구현이 없습니다.

Tasks

Audio GenerationLarge Language Model

Similar Papers 제목 키워드 기반

StepAudio 3 Realtime Technical Report

2026-09-12 · Bin Lin, Bo Zhao, Boyang Zhang, Boyong Wu 외 hf

Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loo…

Step-Audio-R1 Technical Report

2025-11-19 · Fei Tian, Xiangyu Tony Zhang, Yuxin Zhang, Haoyang Zhang 외 arxiv

Recent advances in reasoning models have demonstrated remarkable success in text and vision domains through extended chain-of-thought deliberation. However, a perplexing phenomenon persists in audio language models: they…

Multimodal Reasoning

ProcessThinker: Enhancing Multi-modal Large Language Models Reasoning via Rollout-based Process Reward

2026-04-23 · Jingpei Wu, Xiao Han, Weixiang Shen, Boer Zhang 외 arxiv

Visual question answering increasingly requires multi-step reasoning. Recent post-training with reinforcement learning under verifiable rewards (RLVR) and Group Relative Policy Optimization (GRPO) can improve multimodal …

Visual Question AnsweringReinforcement LearningMultimodal ReasoningLogical Reasoning

ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing

2025-06-26 · Huadai Liu, Jialei Wang, Kaicheng Luo, Wen Wang 외

While end-to-end video-to-audio generation has greatly improved, producing high-fidelity audio that authentically captures the nuances of visual content remains challenging. Like professionals in the creative industries,…

Audio GenerationLarge Language ModelMultimodal Large Language Model

DeepFake Detection: Current Challenges and Next Steps

2020-03-11 · Siwei Lyu

High quality fake videos and audios generated by AI-algorithms (the deep fakes) have started to challenge the status of videos and audios as definitive evidence of events. In this paper, we highlight a few of these chall…

DeepFake DetectionFace Swapping