paper-with-me

홈 › Papers

BatVision with GCC-PHAT Features for Better Sound to Vision Predictions

2020-06-14 · Jesper Haahr Christensen, Sascha Hornauer, Stella Yu

Inspired by sophisticated echolocation abilities found in nature, we train a generative adversarial network to predict plausible depth maps and grayscale layouts from sound. To achieve this, our sound-to-vision model processes binaural echo-returns from chirping sounds. We build upon previous work with BatVision that consists of a sound-to-vision model and a self-collected dataset using our mobile robot and low-cost hardware. We improve on the previous model by introducing several changes to the model, which leads to a better depth and grayscale estimation, and increased perceptual quality. Rather than using raw binaural waveforms as input, we generate generalized cross-correlation (GCC) features and use these as input instead. In addition, we change the model generator and base it on residual learning and use spectral normalization in the discriminator. We compare and present both quantitative and qualitative improvements over our previous BatVision model.

📄 PDF Abstract BibTeX arXiv:2006.07995

Code (1)

SaschaHornauer/Batvision pytorch

Tasks

Generative Adversarial Network

Methods 이 논문이 사용한 방법론

Spectral Normalization Spectral Normalization is a normalization technique used for generative adversarial networks, used to stabilize training of the discriminator. Spectral normalization has the…

Similar Papers 제목 키워드 기반

BatVision: Learning to See 3D Spatial Layout with Two Ears

2019-12-15 · Jesper Haahr Christensen, Sascha Hornauer, Stella Yu

Many species have evolved advanced non-visual perception while artificial systems fall behind. Radar and ultrasound complement camera-based vision but they are often too costly and complex to set up for very limited info…

Robot NavigationVocal Bursts Valence Prediction

Learning Multi-Target TDOA Features for Sound Event Localization and Detection

2024-08-30 · Axel Berg, Johanna Engman, Jens Gulin, Karl Åström 외

Sound event localization and detection (SELD) systems using audio recordings from a microphone array rely on spatial cues for determining the location of sound events. As a consequence, the localization performance of su…

Sound Event Localization and Detection

A study for the effect of the Emphaticness and language and dialect for Voice Onset Time (VOT) in Modern Standard Arabic (MSA)

2013-05-13 · Sulaiman S. AlDahri

The signal sound contains many different features, including Voice Onset Time (VOT), which is a very important feature of stop sounds in many languages. The only application of VOT values is stopping phoneme subsets. Thi…

SVD-PHAT: A Fast Sound Source Localization Method

2019-02-11

This paper introduces a new localization method called SVD-PHAT. The SVD-PHAT method relies on Singular Value Decomposition of the SRP-PHAT projection matrix. A k-d tree is also proposed to speed up the search for the mo…

Sound Source Localization

Robust Sound Source Tracking Using SRP-PHAT and 3D Convolutional Neural Networks

2020-06-16 · David Diaz-Guerra, Antonio Miguel, Jose R. Beltran

In this paper, we present a new single sound source DOA estimation and tracking system based on the well-known SRP-PHAT algorithm and a three-dimensional Convolutional Neural Network. It uses SRP-PHAT power maps as input…