paper-with-me

홈 › Papers

Beyond Descriptions: A Generative Scene2Audio Framework for Blind and Low-Vision Users to Experience Vista Landscapes

2026-03-28 · Chitralekha Gupta, Jing Peng, Ashwin Ram, Shreyas Sridhar, Christophe Jouffrais, Suranga Nanayakkara arxiv

Current scene perception tools for Blind and Low Vision (BLV) individuals rely on spoken descriptions but lack engaging representations of visually pleasing distant environmental landscapes (Vista spaces). Our proposed Scene2Audio framework generates comprehensible and enjoyable nonverbal audio using generative models informed by psychoacoustics, and principles of scene audio composition. Through a user study with 11 BLV participants, we found that combining the Scene2Audio sounds with speech creates a better experience than speech alone, as the sound effects complement the speech making the scene easier to imagine. A mobile app "in-the-wild" study with 7 BLV users for more than a week further showed the potential of Scene2Audio in enhancing outdoor scene experiences. Our work bridges the gap between visual and auditory scene perception by moving beyond purely descriptive aids, addressing the aesthetic needs of BLV users.

📄 PDF Abstract BibTeX arXiv:2603.27295

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Generating Multi-Sentence Lingual Descriptions of Indoor Scenes

2015-02-28 · Dahua Lin, Chen Kong, Sanja Fidler, Raquel Urtasun

This paper proposes a novel framework for generating lingual descriptions of indoor scenes. Whereas substantial efforts have been made to tackle this problem, previous approaches focusing primarily on generating a single…

SentenceText Generation

Recomposer: Event-roll-guided generative audio editing

2025-09-05 · Daniel P. W. Ellis, Eduardo Fonseca, Ron J. Weiss, Kevin Wilson 외 arxiv

Editing complex real-world sound scenes is difficult because individual sound sources overlap in time. Generative models can fill-in missing or corrupted details based on their strong prior understanding of the data doma…

Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation

2024-10-14 · Peiwen Sun, Sitong Cheng, Xiangtai Li, Zhen Ye 외

Recently, diffusion models have achieved great success in mono-channel audio generation. However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions. C…

Audio Generationmultimodal generation

From Sound to Sight: Towards AI-authored Music Videos

2025-08-20 · Leo Vitasovic, Stella Graßhof, Agnes Mercedes Kloft, Ville V. Lehtola 외 arxiv

Conventional music visualisation systems rely on handcrafted ad hoc transformations of shapes and colours that offer only limited expressiveness. We propose two novel pipelines for automatically generating music videos f…

VAInpaint: Zero-Shot Video-Audio inpainting framework with LLMs-driven Module

2025-09-21 · Kam Man Wu, Zeyue Tian, Liya Ji, Qifeng Chen arxiv

Video and audio inpainting for mixed audio-visual content has become a crucial task in multimedia editing recently. However, precisely removing an object and its corresponding audio from a video without affecting the res…

Video Inpainting