paper-with-me

Papers

Point Cloud Audio Processing

2021-05-06 · Krishna Subramani, Paris Smaragdis

Most audio processing pipelines involve transformations that act on fixed-dimensional input representations of audio. For example, when using the Short Time Fourier Transform (STFT) the DFT size specifies a fixed dimension for the input representation. As a consequence, most audio machine learning models are designed to process fixed-size vector inputs which often prohibits the repurposing of learned models on audio with different sampling rates or alternative representations. We note, however, that the intrinsic spectral information in the audio signal is invariant to the choice of the input representation or the sampling rate. Motivated by this, we introduce a novel way of processing audio signals by treating them as a collection of points in feature space, and we use point cloud machine learning models that give us invariance to the choice of representation parameters, such as DFT size or the sampling rate. Additionally, we observe that these methods result in smaller models, and allow us to significantly subsample the input representation with minimal effects to a trained model performance.

📄 PDF Abstract BibTeX arXiv:2105.02469

Code (1)

SubramaniKrishna/point-cloud-audio 공식 구현 pytorch

Tasks

BIG-bench Machine Learning

Similar Papers 제목 키워드 기반

Points2Sound: From mono to binaural audio using 3D point cloud scenes

2021-04-26 · Francesc Lluís, Vasileios Chatziioannou, Alex Hofmann

For immersive applications, the generation of binaural sound that matches its visual counterpart is crucial to bring meaningful experiences to people in a virtual environment. Recent studies have shown the possibility of…

Audio Synthesis

PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis

2024-12-11 · Yifan Xie, Tao Feng, Xin Zhang, Xiangyang Luo 외

Talking head synthesis with arbitrary speech audio is a crucial challenge in the field of digital humans. Recently, methods based on radiance fields have received increasing attention due to their ability to synthesize h…

Music source separation conditioned on 3D point clouds

2021-02-03 · Francesc Lluís, Vasileios Chatziioannou, Alex Hofmann

Recently, significant progress has been made in audio source separation by the application of deep learning techniques. Current methods that combine both audio and visual information use 2D representations such as images…

Audio Source SeparationMusic Source Separation

Point Cloud Resampling Through Hypergraph Signal Processing

2021-02-12 · Qinwen Deng, Songyang Zhang, Zhi Ding

Three-dimensional (3D) point clouds are important data representations in visualization applications. The rapidly growing utility and popularity of point cloud processing strongly motivate a plethora of research activiti…

Object

A Novel Frame Structure for Cloud-Based Audio-Visual Speech Enhancement in Multimodal Hearing-aids

2022-10-24 · Abhijeet Bishnu, Ankit Gupta, Mandar Gogate, Kia Dashtipour 외

In this paper, we design a first of its kind transceiver (PHY layer) prototype for cloud-based audio-visual (AV) speech enhancement (SE) complying with high data rate and low latency requirements of future multimodal hea…

Lip ReadingSpeech Enhancement