paper-with-me

홈 › Papers

Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech

2024-10-18 · Shuwei He, Rui Liu

Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize reverberant speech for the spoken content. Previous works focus on the RGB modality for global environmental modeling, overlooking the potential of multi-source spatial knowledge like depth, speaker position, and environmental semantics. To address these issues, we propose a novel multi-source spatial knowledge understanding scheme for immersive VTTS, termed MS2KU-VTTS. Specifically, we first prioritize RGB image as the dominant source and consider depth image, speaker position knowledge from object detection, and Gemini-generated semantic captions as supplementary sources. Afterwards, we propose a serial interaction mechanism to effectively integrate both dominant and supplementary sources. The resulting multi-source knowledge is dynamically integrated based on the respective contributions of each source.This enriched interaction and integration of multi-source spatial knowledge guides the speech generation model, enhancing the immersive speech experience. Experimental results demonstrate that the MS$^2$KU-VTTS surpasses existing baselines in generating immersive speech. Demos and code are available at: https://github.com/AI-S2-Lab/MS2KU-VTTS.

📄 PDF Abstract BibTeX arXiv:2410.14101

Code (1)

ms2ku-vtts/ms2ku-vtts 공식 구현

Tasks

object-detectionObject DetectionPositiontext-to-speechText to Speech

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech

2024-12-16 · Rui Liu, Shuwei He, Yifan Hu, Haizhou Li

Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize the reverberant speech for the spoken content. The challenge of this task lies in understanding the spatial environment from t…

text-to-speechText to Speech

ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting

2025-04-29 · Yu Zhang, Wenxiang Guo, Changhao Pan, Zhiyuan Zhu 외

Multimodal immersive spatial drama generation focuses on creating continuous multi-speaker binaural speech with dramatic prosody based on multimodal prompts, with potential applications in AR, VR, and others. This task r…

Contrastive LearningMamba

ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?

2025-10-13 · Liu Yang, Huiyu Duan, Ran Tao, Juntao Cheng 외 arxiv

Omnidirectional images (ODIs) provide full 360x180 view which are widely adopted in VR, AR and embodied intelligence applications. While multi-modal large language models (MLLMs) have demonstrated remarkable performance …

Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes

2025-03-28 · Binh Thien Nguyen, Masahiro Yasuda, Daiki Takeuchi, Daisuke Niizumi 외

Immersive communication has made significant advancements, especially with the release of the codec for Immersive Voice and Audio Services. Aiming at its further realization, the DCASE 2025 Challenge has recently introdu…

Audio TaggingSemantic Segmentation

Realizing Immersive Volumetric Video: A Multimodal Framework for 6-DoF VR Engagement

2026-04-10 · Zhengxian Yang, Shengqi Wang, Shi Pan, Hongshuai Li 외 arxiv

Fully immersive experiences that tightly integrate 6-DoF visual and auditory interaction are essential for virtual and augmented reality. While such experiences can be achieved through computer-generated content, constru…