paper-with-me

홈 › Papers

VividVoice: A Unified Framework for Scene-Aware Visually-Driven Speech Synthesis

2026-02-01 · Chengyuan Ma, Jiawei Jin, Ruijie Xiong, Chunxiang Jin, Canxiang Yan, Wenming Yang arxiv

We introduce and define a novel task-Scene-Aware Visually-Driven Speech Synthesis, aimed at addressing the limitations of existing speech generation models in creating immersive auditory experiences that align with the real physical world. To tackle the two core challenges of data scarcity and modality decoupling, we propose VividVoice, a unified generative framework. First, we constructed a large-scale, high-quality hybrid multimodal dataset, Vivid-210K, which, through an innovative programmatic pipeline, establishes a strong correlation between visual scenes, speaker identity, and audio for the first time. Second, we designed a core alignment module, D-MSVA, which leverages a decoupled memory bank architecture and a cross-modal hybrid supervision strategy to achieve fine-grained alignment from visual scenes to timbre and environmental acoustic features. Both subjective and objective experimental results provide strong evidence that VividVoice significantly outperforms existing baseline models in terms of audio fidelity, content clarity, and multimodal consistency. Our demo is available at https://chengyuann.github.io/VividVoice/.

📄 PDF Abstract BibTeX arXiv:2602.02591

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

PhyMix: Towards Physically Consistent Single-Image 3D Indoor Scene Generation with Implicit--Explicit Optimization

2026-04-11 · Dongli Wu, Jingyu Hu, Ka-Hei Hui, Xiaobao Wei 외 arxiv

Existing single-image 3D indoor scene generators often produce results that look visually plausible but fail to obey real-world physics, limiting their reliability in robotics, embodied AI, and design. To examine this ga…

Scene Generation

AiSDF: Structure-aware Neural Signed Distance Fields in Indoor Scenes

2024-03-04 · Jaehoon Jang, Inha Lee, Minje Kim, Kyungdon Joo

Indoor scenes we are living in are visually homogenous or textureless, while they inherently have structural forms and provide enough structural priors for 3D scene reconstruction. Motivated by this fact, we propose a st…

3D Scene Reconstruction

PAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene Understanding

2026-06-04 · Shaohui Dai, Yansong Qu, You Shen, Shengchuan Zhang 외 arxiv

Recent advances in 3D multimodal large language models (3D-MLLMs) have enabled unified solutions for 3D scene understanding tasks, including visual question answering, captioning, and referring segmentation. However, exi…

Visual Question AnsweringRepresentation LearningScene Understanding

Drag4D: Align Your Motion with Text-Driven 3D Scene Generation

2025-09-26 · Minjun Kang, Inkyu Shin, Taeyeop Lee, In So Kweon 외 arxiv

We introduce Drag4D, an interactive framework that integrates object motion control within text-driven 3D scene generation. This framework enables users to define 3D trajectories for the 3D objects generated from a singl…

Scene Generation

Scene-Aware Vectorized Memory Multi-Agent Framework with Cross-Modal Differentiated Quantization VLMs for Visually Impaired Assistance

2025-08-25 · Xiangxiang Wang, Xuanyu Wang, YiJia Luo, Yongbin Yu 외 arxiv

Visually impaired individuals face significant challenges in environmental perception. Traditional assistive technologies often lack adaptive intelligence, focusing on individual components rather than integrated systems…

Computational Efficiency