paper-with-me

홈 › Papers

Extreme Audio Time Stretching Using Neural Synthesis

2022-11-30 · Leonardo Fierro, Alec Wright, Vesa Välimäki, Matti Hämäläinen

A deep neural network solution for time-scale modification (TSM) focused on large stretching factors is proposed, targeting environmental sounds. Traditional TSM artifacts such as transient smearing, loss of presence, and phasiness are heavily accentuated and cause poor audio quality when the TSM factor is four or larger. The weakness of established TSM methods, often based on a phase vocoder structure, lies in the poor description and scaling of the transient and noise components, or nuances, of a sound. Our novel solution combines a sines-transients-noise decomposition with an independent WaveNet synthesizer to provide a better description of the noise component and an improve sound quality for large stretching factors. Results of a subjective listening test against four other TSM algorithms are reported, showing the proposed method to be often superior. The proposed method is stereo compatible and has a wide range of applications related to the slow motion of media content.

📄 PDF Abstract BibTeX arXiv:2211.16992

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Test 설명 없음
Mixture of Logistic Distributions 설명 없음
Dilated Causal Convolution A Dilated Causal Convolution is a causal convolution where the filter is applied over an area larger than its length by…
WaveNet WaveNet is an audio generative model based on the PixelCNN architecture. In order to deal with long-range temporal dependencies…

Similar Papers 제목 키워드 기반

Noise Morphing for Audio Time Stretching

2023-12-22 · Eloi Moliner, Leonardo Fierro, Alec Wright, Matti Hämäläinen 외

This letter introduces an innovative method to enhance the quality of audio time stretching by precisely decomposing a sound into sines, transients, and noise and by improving the processing of the latter component. Whil…

Resynthesis

Neural Pitch-Shifting and Time-Stretching with Controllable LPCNet

2021-10-05 · Max Morrison, Zeyu Jin, Nicholas J. Bryan, Juan-Pablo Caceres 외

Modifying the pitch and timing of an audio signal are fundamental audio editing operations with applications in speech manipulation, audio-visual synchronization, and singing voice editing and synthesis. Thus far, method…

Audio-Visual Synchronization

PeakNetFP: Peak-based Neural Audio Fingerprinting Robust to Extreme Time Stretching

2025-06-26 · Guillem Cortès-Sebastià, Benjamin Martin, Emilio Molina, Xavier Serra 외

This work introduces PeakNetFP, the first neural audio fingerprinting (AFP) system designed specifically around spectral peaks. This novel system is designed to leverage the sparse spectral coordinates typically computed…

Contrastive Learning

Pretrained Conformers for Audio Fingerprinting and Retrieval

2025-08-15 · Kemal Altwlkany, Elmedin Selmanovic, Sead Delalic arxiv

Conformers have shown great results in speech processing due to their ability to capture both local and global interactions. In this work, we utilize a self-supervised contrastive learning framework to train conformer-ba…

Contrastive Learning

Deep Transform: Time-Domain Audio Error Correction via Probabilistic Re-Synthesis

2015-03-19 · Andrew J. R. Simpson

In the process of recording, storage and transmission of time-domain audio signals, errors may be introduced that are difficult to correct in an unsupervised way. Here, we train a convolutional deep neural network to re-…