paper-with-me

Papers

Perceptually-motivated Spatial Audio Codec for Higher-Order Ambisonics Compression

2024-01-24 · Christoph Hold, Leo McCormack, Archontis Politis, Ville Pulkki

Scene-based spatial audio formats, such as Ambisonics, are playback system agnostic and may therefore be favoured for delivering immersive audio experiences to a wide range of (potentially unknown) devices. The number of channels required to deliver high spatial resolution Ambisonic audio, however, can be prohibitive for low-bandwidth applications. Therefore, this paper proposes a compression codec, which is based upon the parametric higher-order Directional Audio Coding (HO-DirAC) model. The encoder downmixes the higher-order Ambisonic (HOA) input audio into a reduced number of signals, which are accompanied by perceptually-motivated scene parameters. The downmixed audio is coded using a perceptual audio coder, whereas the parameters are grouped into perceptual bands, quantized, and downsampled. On the decoder side, low Ambisonic orders are fully recovered. Not fully recoverable HOA components are synthesized according to the parameters. The results of a listening test indicate that the proposed parametric spatial audio codec can improve the adopted perceptual audio coder, especially at low to medium-high bitrates, when applied to fifth-order HOA signals.

📄 PDF Abstract BibTeX arXiv:2401.13401

Code (0)

등록된 구현이 없습니다.

Tasks

Decoder

Similar Papers 제목 키워드 기반

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding

2026-06-03 · Eugene Kwek, Feng Liu, Rui Zhang, Wenpeng Yin arxiv

Neural audio codecs are a key component of speech processing pipelines, compressing audio into discrete tokens for downstream modeling. However, existing codecs struggle to balance reconstruction quality with token effic…

Voice Conversion

Analyzing and Mitigating Inconsistency in Discrete Audio Tokens for Neural Codec Language Models

2024-09-28 · Wenrui Liu, Zhifang Guo, Jin Xu, YuanJun Lv 외

Building upon advancements in Large Language Models (LLMs), the field of audio processing has seen increased interest in training audio generation tasks with discrete audio token sequences. However, directly discretizing…

Audio GenerationLanguage ModelingLanguage Modelling

Towards Improved Objective Perceptual Audio Quality Assessment -- Part 1: A Novel Data-Driven Cognitive Model

2024-11-27 · Pablo M. Delgado, Jürgen Herre

Efficient audio quality assessment is vital for streamlining audio codec development. Objective assessment tools have been developed over time to algorithmically predict quality ratings from subjective assessments, the g…

Audio Quality AssessmentPrediction

Psychoacoustic Calibration of Loss Functions for Efficient End-to-End Neural Audio Coding

2020-12-31 · Kai Zhen, Mi Suk Lee, Jongmo Sung, SeungKwon Beack 외

Conventional audio coding technologies commonly leverage human perception of sound, or psychoacoustics, to reduce the bitrate while preserving the perceptual quality of the decoded audio signals. For neural audio codecs,…

Gull: A Generative Multifunctional Audio Codec

2024-04-07 · Yi Luo, Jianwei Yu, Hangting Chen, Rongzhi Gu 외

We introduce Gull, a generative multifunctional audio codec. Gull is a general purpose neural audio compression and decompression model which can be applied to a wide range of tasks and applications such as real-time com…

Audio CompressionAudio Source SeparationAudio Super-ResolutionDecoder+2