Papers Audio Compression
“Audio Compression” 태그가 달린 논문 42편 · 필터 해제
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
In recent years, large language models have achieved significant success in generative tasks related to speech, audio, music, and other signal domains. A crucial element of these models is the discrete acoustic codecs, w…
Audio CompressionAudio GenerationQuantizationSpiking Music: Audio Compression with Event Based Auto-encoders
Neurons in the brain communicate information via punctual events called spikes. The timing of spikes is thought to carry rich information, but it is not clear how to leverage this in digital systems. We demonstrate that …
Audio CompressionMusic CompressionMAGMA: Music Aligned Generative Motion Autodecoder
Mapping music to dance is a challenging problem that requires spatial and temporal coherence along with a continual synchronization with the music's progression. Taking inspiration from large language models, we introduc…
Audio CompressionDecoderMotion GenerationEdge Storage Management Recipe with Zero-Shot Data Compression for Road Anomaly Detection
Recent studies show edge computing-based road anomaly detection systems which may also conduct data collection simultaneously. However, the edge computers will have small data storage but we need to store the collected a…
Anomaly DetectionAudio CompressionAudio Super-ResolutionData Compression+3Siamese SIREN: Audio Compression with Implicit Neural Representations
Implicit Neural Representations (INRs) have emerged as a promising method for representing diverse data modalities, including 3D shapes, images, and audio. While recent research has demonstrated successful applications o…
Audio CompressionQuantifying Spatial Audio Quality Impairment
Spatial audio quality is a highly multifaceted concept, with many interactions between environmental, geometrical, anatomical, psychological, and contextual considerations. Methods for characterization or evaluation of t…
Audio CompressionMusic Source SeparationHigh-Fidelity Audio Compression with Improved RVQGAN
Language models have been successfully used to model natural signals, such as images, speech, and music. A key component of these models is a high quality neural compression model that can compress high-dimensional natur…
Audio CompressionAudio GenerationQuantizationDC CoMix TTS: An End-to-End Expressive TTS with Discrete Code Collaborated with Mixer
Despite the huge successes made in neutral TTS, content-leakage remains a challenge. In this paper, we propose a new input representation and simple architecture to achieve improved prosody modeling. Inspired by the rece…
Audio CompressionCompression with Bayesian Implicit Neural Representations
Many common types of data can be represented as functions that map coordinates to signal values, such as pixel locations to RGB values in the case of an image. Based on this view, data can be compressed by overfitting a …
Audio CompressionQuantizationAn investigation of the reconstruction capacity of stacked convolutional autoencoders for log-mel-spectrograms
In audio processing applications, the generation of expressive sounds based on high-level representations demonstrates a high demand. These representations can be used to manipulate the timbre and influence the synthesis…
Audio CompressionCombining Automatic Speaker Verification and Prosody Analysis for Synthetic Speech Detection
The rapid spread of media content synthesis technology and the potentially damaging impact of audio and video deepfakes on people's lives have raised the need to implement systems able to detect these forgeries automatic…
Audio CompressionFace SwappingRhythmSpeaker Verification+4High Fidelity Neural Audio Compression
We introduce a state-of-the-art real-time, high-fidelity, audio codec leveraging neural networks. It consists in a streaming encoder-decoder architecture with quantized latent space trained in an end-to-end fashion. We s…
Audio CompressionAudio Signal ProcessingDecoderVocal Bursts Intensity PredictionScaling and compressing melodies using geometric similarity measures
Melodic similarity measurement is of key importance in music information retrieval. In this paper, we use geometric matching techniques to measure the similarity between two melodies. We represent music as sets of points…
Audio CompressionGeometric MatchingInformation RetrievalMusic Information Retrieval+1On The Effect Of Coding Artifacts On Acoustic Scene Classification
Previous DCASE challenges contributed to an increase in the performance of acoustic scene classification systems. State-of-the-art classifiers demand significant processing capabilities and memory which is challenging fo…
Acoustic Scene ClassificationAudio CompressionClassificationScene ClassificationAudio Spectral Enhancement: Leveraging Autoencoders for Low Latency Reconstruction of Long, Lossy Audio Sequences
With active research in audio compression techniques yielding substantial breakthroughs, spectral reconstruction of low-quality audio waves remains a less indulged topic. In this paper, we propose a novel approach for re…
Audio CompressionQuantizationSpectral ReconstructionUR Channel-Robust Synthetic Speech Detection System for ASVspoof 2021
In this paper, we present UR-AIR system submission to the logical access (LA) and the speech deepfake (DF) tracks of the ASVspoof 2021 Challenge. The LA and DF tasks focus on synthetic speech detection (SSD), i.e. detect…
Audio CompressionFace SwappingSynthetic Speech Detectiontext-to-speech+2Deep Neural Networks and End-to-End Learning for Audio Compression
Recent achievements in end-to-end deep learning have encouraged the exploration of tasks dealing with highly structured data with unified deep network models. Having such models for compressing audio signals has been cha…
Audio CompressionDecoderDeep LearningClefNet: Recurrent Autoencoders with Dynamic Time Warping for Near-Lossless Music Compression and Minimal-Latency Transmission
The onset of coronavirus disease 2019 (COVID-19), an infectious disease caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), has sparked unprecedented change. Due to the public health guidelines impose…
Audio CompressionDynamic Time WarpingMusic CompressionMP3net: coherent, minute-long music generation from raw audio with a simple convolutional GAN
We present a deep convolutional GAN which leverages techniques from MP3/Vorbis audio compression to produce long, high-quality audio samples with long-range coherence. The model uses a Modified Discrete Cosine Transform …
Audio CompressionMusic GenerationBayesian Reconstruction of Fourier Pairs
In a number of data-driven applications such as detection of arrhythmia, interferometry or audio compression, observations are acquired indistinctly in the time or frequency domains: temporal observations allow us to stu…
AstronomyAudio Compression