paper-with-me

Papers

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation

2026-05-09 · Shihao Cheng, Jiaxu Zhang, Quanyue Song, Shansong Liu, Zhizhi Guo, Xiaolei Zhang, Chi Zhang, Xuelong Li, Zhigang Tu arxiv

Motion, speech, and sound effects are fundamental elements of human-centric videos, yet their heterogeneous temporal characteristics make joint generation highly challenging. Existing audio-video generation models often fail to maintain consistent alignment across these modalities, leading to noticeable mismatches between motion, speech, and environmental sounds. We present Unison, a unified framework that explicitly promotes coherence across the motion, speech, and sound modalities. Within the audio stream, Unison employs a semantic-guided harmonization strategy that decouples the generation of speech and sound-effect components. Leveraging bidirectional audio cross-attention and semantic-conditioned gating for semantic-driven adaptive recomposition, this approach effectively mitigates speech dominance and enhances acoustic clarity. For audio-motion synchronization, we propose a bidirectional cross-modal forcing strategy where the cleaner modality guides the noisier one through decoupled denoising schedules, reinforced by a progressive stabilization strategy. Extensive experiments demonstrate that Unison achieves state-of-the-art performance in both audio perceptual quality and cross-modal synchronization, highlighting the importance of explicit multimodal harmonization in human-centric video generation.

📄 PDF Abstract BibTeX arXiv:2605.08729

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

2026-04-24 · Chunyu Qiang, Xiaopeng Wang, Kang Yin, Yuzhe Liang 외 arxiv

Generative audio modeling has largely been fragmented into specialized tasks, text-to-speech (TTS), text-to-music (TTM), and text-to-audio (TTA), each operating under heterogeneous control paradigms. Unifying these modal…

Penambahan emosi menggunakan metode manipulasi prosodi untuk sistem text to speech bahasa Indonesia

2016-06-29 · Salita Ulitia Prini, Ary Setijadi Prihatmanto

Adding an emotions using prosody manipulation method for Indonesian text to speech system. Text To Speech (TTS) is a system that can convert text in one language into speech, accordance with the reading of the text in th…

Sentencetext-to-speechText to Speech

Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation

2026-06-25 · Jinyu Liu, Xincheng Shuai, Henghui Ding, Yu-Gang Jiang arxiv

Unified multimodal models capable of both understanding and generation have achieved remarkable strides. However, despite their unified designs, existing evaluations typically assess understanding and generation capabili…

Detection and Analysis of Human Emotions through Voice and Speech Pattern Processing

2017-10-27 · Poorna Banerjee Dasgupta

The ability to modulate vocal sounds and generate speech is one of the features which set humans apart from other living beings. The human voice can be characterized by several attributes such as pitch, timbre, loudness,…

DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches

2025-08-20 · Tai Vu arxiv

Speech emotion recognition (SER) has been a challenging problem in spoken language processing research, because it is unclear how human emotions are connected to various components of sounds such as pitch, loudness, and …

Speech Emotion RecognitionTransfer LearningData Augmentation