paper-with-me

홈 › Papers

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech

2024-12-16 · Rui Liu, Shuwei He, Yifan Hu, Haizhou Li

Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize the reverberant speech for the spoken content. The challenge of this task lies in understanding the spatial environment from the image. Many attempts have been made to extract global spatial visual information from the RGB space of an spatial image. However, local and depth image information are crucial for understanding the spatial environment, which previous works have ignored. To address the issues, we propose a novel multi-modal and multi-scale spatial environment understanding scheme to achieve immersive VTTS, termed M2SE-VTTS. The multi-modal aims to take both the RGB and Depth spaces of the spatial image to learn more comprehensive spatial information, and the multi-scale seeks to model the local and global spatial knowledge simultaneously. Specifically, we first split the RGB and Depth images into patches and adopt the Gemini-generated environment captions to guide the local spatial understanding. After that, the multi-modal and multi-scale features are integrated by the local-aware global spatial understanding. In this way, M2SE-VTTS effectively models the interactions between local and global spatial contexts in the multi-modal spatial environment. Objective and subjective evaluations suggest that our model outperforms the advanced baselines in environmental speech generation. The code and audio samples are available at: https://github.com/AI-S2-Lab/M2SE-VTTS.

📄 PDF Abstract BibTeX arXiv:2412.11409

Code (1)

ai-s2-lab/m2se-vtts 공식 구현

Tasks

text-to-speechText to Speech

Methods 이 논문이 사용한 방법론

ADOPT Please enter a description about the method here

Similar Papers 제목 키워드 기반

Multi-Scale and Multimodal Species Distribution Modeling

2024-11-06 · Nina van Tiel, Robin Zbinden, Emanuele Dalsasso, Benjamin Kellenberger 외

Species distribution models (SDMs) aim to predict the distribution of species by relating occurrence data with environmental variables. Recent applications of deep learning to SDMs have enabled new avenues, specifically …

A Geolocation-Aware Multimodal Approach for Ecological Prediction

2026-01-13 · Valerie Zermatten, Chiara Vanalli, Gencer Sumbul, Diego Marcos 외 arxiv

While integrating multiple modalities has the potential to improve environmental monitoring, current approaches struggle to combine data sources with heterogeneous formats or contents. A central difficulty arises when co…

SIGMMA: Hierarchical Graph-Based Multi-Scale Multi-modal Contrastive Alignment of Histopathology Image and Spatial Transcriptome

2025-11-19 · Dabin Jeong, Amirhossein Vahidi, Ciro Ramírez-Suástegui, Marie Moullet 외 arxiv

Recent advances in computational pathology have leveraged vision-language models to learn joint representations of Hematoxylin and Eosin (HE) images with spatial transcriptomic (ST) profiles. However, existing approaches…

Cross-Modal Retrieval

Multimodal and Multiscale Spatial-Temporal Semantic Search and Recommendation with AI Foundation Models

2026-06-15 · Yuanyuan Tian, Wenwen Li, Xiao Chen, Michael Brook 외 arxiv

Semantic search and recommendation of similar documents, such as news and reports about unusual environmental events (e.g., a dead whale washed ashore in Alaska) that contain spatial and temporal information, is a critic…

Information RetrievalSemantic Similarity

GeoViSTA: Geospatial Vision-Tabular Transformer for Multimodal Environment Representation

2026-05-14 · Yuhao Liu, Sadeer Al-Kindi, Ashok Veeraraghavan, Guha Balakrishnan arxiv

Large-scale pretraining on Earth observation imagery has yielded powerful representations of the natural and built environment. However, most existing geospatial foundation models do not directly model the structured soc…