paper-with-me

홈 › Papers

Compression of Higher Order Ambisonics with Multichannel RVQGAN

2024-11-18 · Toni Hirvonen, Mahmoud Namazi

A multichannel extension to the RVQGAN neural coding method is proposed, and realized for data-driven compression of third-order Ambisonics audio. The input- and output layers of the generator and discriminator models are modified to accept multiple (16) channels without increasing the model bitrate. We also propose a loss function for accounting for spatial perception in immersive reproduction, and transfer learning from single-channel models. Listening test results with 7.1.4 immersive playback show that the proposed extension is suitable for coding scene-based, 16-channel Ambisonics content with good quality at 16 kbps when trained and tested on the EigenScape database. The model has potential applications for learning other types of content and multichannel formats.

📄 PDF Abstract BibTeX arXiv:2411.12008

Code (0)

등록된 구현이 없습니다.

Tasks

Transfer Learning

Similar Papers 제목 키워드 기반

Perceptually-motivated Spatial Audio Codec for Higher-Order Ambisonics Compression

2024-01-24 · Christoph Hold, Leo McCormack, Archontis Politis, Ville Pulkki

Scene-based spatial audio formats, such as Ambisonics, are playback system agnostic and may therefore be favoured for delivering immersive audio experiences to a wide range of (potentially unknown) devices. The number of…

Decoder

Dilated U-net based approach for multichannel speech enhancement from First-Order Ambisonics recordings

2020-06-02 · Amélie Bosca, Alexandre Guérin, Lauréline Perotin, Srđan Kitić

We present a CNN architecture for speech enhancement from multichannel first-order Ambisonics mixtures. The data-dependent spatial filters, deduced from a mask-based approach, are used to help an automatic speech recogni…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+1

Direction of Arrival Estimation of Noisy Speech Using Convolutional Recurrent Neural Networks with Higher-Order Ambisonics Signals

2021-02-19 · Nils Poschadel, Robert Hupke, Stephan Preihs, Jürgen Peissig

Training convolutional recurrent neural networks on first-order Ambisonics signals is a well-known approach when estimating the direction of arrival for speech/sound signals. In this work, we investigate whether increasi…

Direction of Arrival Estimation

L3DAS21 Challenge: Machine Learning for 3D Audio Signal Processing

2021-04-12 · Eric Guizzo, Riccardo F. Gramaccioni, Saeid Jamili, Christian Marinoni 외

The L3DAS21 Challenge is aimed at encouraging and fostering collaborative research on machine learning for 3D audio signal processing, with particular focus on 3D speech enhancement (SE) and 3D sound localization and det…

Audio Signal ProcessingBIG-bench Machine LearningSpeech Enhancement

Weakly Guided and Autoregressive Beamformer Parameterization for Generalizable Moving Speaker Extraction in Higher-Order Ambisonics

2026-07-05 · Jakob Kienegger, Tal Peer, Sina Khanagha, Timo Gerkmann arxiv

Linear spatial filters (beamformers) enable robust, generalizable and interpretable speech enhancement with performance guarantees under ideal parameterization. Modern beamformers are often parameterized by deep neural n…

Speech Enhancement