paper-with-me

홈 › Papers

AudioViewer: Learning to Visualize Sounds

2020-12-22 · Chunjin Song, Yuchi Zhang, Willis Peng, Parmis Mohaghegh, Bastian Wandt, Helge Rhodin

A long-standing goal in the field of sensory substitution is to enable sound perception for deaf and hard of hearing (DHH) people by visualizing audio content. Different from existing models that translate to hand sign language, between speech and text, or text and images, we target immediate and low-level audio to video translation that applies to generic environment sounds as well as human speech. Since such a substitution is artificial, without labels for supervised learning, our core contribution is to build a mapping from audio to video that learns from unpaired examples via high-level constraints. For speech, we additionally disentangle content from style, such as gender and dialect. Qualitative and quantitative results, including a human study, demonstrate that our unpaired translation approach maintains important audio features in the generated video and that videos of faces and numbers are well suited for visualizing high-dimensional audio features that can be parsed by humans to match and distinguish between sounds and words. Code and models are available at https://chunjinsong.github.io/audioviewer

📄 PDF Abstract BibTeX arXiv:2012.13341

Code (1)

ChunjinSong/audioviewer_code 공식 구현 pytorch

Tasks

Translation

Similar Papers 제목 키워드 기반

Generating Realistic Images from In-the-wild Sounds

2023-09-05 · ICCV 2023 1 · Taegyeong Lee, Jeonghun Kang, Hyeonyu Kim, Taehwan Kim

Representing wild sounds as images is an important but challenging task due to the lack of paired datasets between sound and images and the significant differences in the characteristics of these two modalities. Previous…

Audio captioningSentence

SoundSil-DS: Deep Denoising and Segmentation of Sound-field Images with Silhouettes

2024-11-12 · Risako Tanigawa, Kenji Ishikawa, Noboru Harada, Yasuhiro Oikawa

Development of optical technology has enabled imaging of two-dimensional (2D) sound fields. This acousto-optic sensing enables understanding of the interaction between sound and objects such as reflection and diffraction…

DenoisingObject

Low-dimensional representation of infant and adult vocalization acoustics

2022-04-25 · Silvia Pagliarini, Sara Schneider, Christopher T. Kello, Anne S. Warlaumont

During the first years of life, infant vocalizations change considerably, as infants develop the vocalization skills that enable them to produce speech sounds. Characterizations based on specific acoustic features, proto…

Pediatric Asthma Detection with Googles HeAR Model: An AI-Driven Respiratory Sound Classifier

2025-04-28 · Abul Ehtesham, Saket Kumar, Aditi Singh, Tala Talaei Khoei

Early detection of asthma in children is crucial to prevent long-term respiratory complications and reduce emergency interventions. This work presents an AI-powered diagnostic pipeline that leverages Googles Health Acous…

Diagnostic

Feature Extraction for Machine Learning Based Crackle Detection in Lung Sounds from a Health Survey

2017-05-31 · Morten Grønnesby, Juan Carlos Aviles Solis, Einar Holsbø, Hasse Melbye 외

In recent years, many innovative solutions for recording and viewing sounds from a stethoscope have become available. However, to fully utilize such devices, there is a need for an automated approach for detecting abnorm…

BIG-bench Machine Learning