paper-with-me

홈 › Papers

3D-MOV: Audio-Visual LSTM Autoencoder for 3D Reconstruction of Multiple Objects from Video

2021-10-05 · Justin Wilson, Ming C. Lin

3D object reconstructions of transparent and concave structured objects, with inferred material properties, remains an open research problem for robot navigation in unstructured environments. In this paper, we propose a multimodal single- and multi-frame neural network for 3D reconstructions using audio-visual inputs. Our trained reconstruction LSTM autoencoder 3D-MOV accepts multiple inputs to account for a variety of surface types and views. Our neural network produces high-quality 3D reconstructions using voxel representation. Based on Intersection-over-Union (IoU), we evaluate against other baseline methods using synthetic audio-visual datasets ShapeNet and Sound20K with impact sounds and bounding box annotations. To the best of our knowledge, our single- and multi-frame model is the first audio-visual reconstruction neural network for 3D geometry and material representation.

📄 PDF Abstract BibTeX arXiv:2110.02404

Code (0)

등록된 구현이 없습니다.

Tasks

3D geometry3D ReconstructionRobot Navigation

Methods 이 논문이 사용한 방법론

Tanh Activation 설명 없음
Sigmoid Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

End-to-End Audiovisual Fusion with LSTMs

2017-09-12 · Stavros Petridis, Yujiang Wang, Zuwei Li, Maja Pantic

Several end-to-end deep learning approaches have been recently presented which simultaneously extract visual features from the input images and perform visual speech classification. However, research on jointly extractin…

ClassificationGeneral Classificationspeech-recognitionSpeech Recognition

Non-linear prediction with LSTM recurrent neural networks for acoustic novelty detection

2015-10-01 · 2015 International Joint Conference on Neural Networks (IJCNN) 2015 10 · Erik Marchi ; Fabio Vesperini ; Felix Weninger ; Florian Eyben ; Stefano Squartini ; Björn Schuller

Acoustic novelty detection aims at identifying abnormal/novel acoustic signals which differ from the reference/normal data that the system was trained with. In this paper we present a novel approach based on non-linear p…

Acoustic Novelty DetectionDenoisingNovelty Detection

Audio Spectral Enhancement: Leveraging Autoencoders for Low Latency Reconstruction of Long, Lossy Audio Sequences

2021-08-08 · Darshan Deshpande, Harshavardhan Abichandani

With active research in audio compression techniques yielding substantial breakthroughs, spectral reconstruction of low-quality audio waves remains a less indulged topic. In this paper, we propose a novel approach for re…

Audio CompressionQuantizationSpectral Reconstruction

CrossMAE: Cross-Modality Masked Autoencoders for Region-Aware Audio-Visual Pre-Training

2024-01-01 · CVPR 2024 1 · Yuxin Guo, Siyang Sun, Shuailei Ma, Kecheng Zheng 외

Learning joint and coordinated features across modalities is essential for many audio-visual tasks. Existing pre-training methods primarily focus on global information neglecting fine-grained features and positions l…

Audio Word2Vec: Unsupervised Learning of Audio Segment Representations using Sequence-to-sequence Autoencoder

2016-03-03 · Yu-An Chung, Chao-Chung Wu, Chia-Hao Shen, Hung-Yi Lee 외

The vector representations of fixed dimensionality for words (in text) offered by Word2Vec have been shown to be very useful in many application scenarios, in particular due to the semantic information they carry. This p…

DecoderDenoisingDynamic Time Warping