Papers Resynthesis
“Resynthesis” 태그가 달린 논문 51편 · 필터 해제
Spoken Language Modeling with Duration-Penalized Self-Supervised Units
Spoken language models (SLMs) operate on acoustic units obtained by discretizing self-supervised speech representations. Although the characteristics of these units directly affect performance, the interaction between co…
Language ModelingLanguage ModellingResynthesisSegmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
Quantization in SSL speech models (e.g., HuBERT) improves compression and performance in tasks like language modeling, resynthesis, and text-to-speech but often discards prosodic and paralinguistic information (e.g., emo…
Language ModelingLanguage ModellingQuantizationResynthesis+2Dynamic Tsetlin Machine Accelerators for On-Chip Training at the Edge using FPGAs
The increased demand for data privacy and security in machine learning (ML) applications has put impetus on effective edge training on Internet-of-Things (IoT) nodes. Edge training aims to leverage speed, energy efficien…
ResynthesisRuntime Tunable Tsetlin Machines for Edge Inference on eFPGAs
Embedded Field-Programmable Gate Arrays (eFPGAs) allow for the design of hardware accelerators of edge Machine Learning (ML) applications at a lower power budget compared with traditional FPGA platforms. However, the lim…
Model CompressionResynthesisFocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks
Large language models have revolutionized natural language processing through self-supervised pretraining on massive datasets. Inspired by this success, researchers have explored adapting these methods to speech by discr…
ResynthesisVoice ConversionAnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder
This article introduces AnCoGen, a novel method that leverages a masked autoencoder to unify the analysis, control, and generation of speech signals within a single model. AnCoGen can analyze speech by estimating key att…
Pitch ClassificationPitch controlResynthesisSpeech Enhancement+1DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
Spoken language models (SLMs) have gained increasing attention with advancements in text-based, decoder-only language models. SLMs process text and speech, enabling simultaneous speech understanding and generation. This …
DecoderResynthesisSpeech TokenizationA Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation
Neural Audio Codecs, initially designed as a compression technique, have gained more attention recently for speech generation. Codec models represent each audio frame as a sequence of tokens, i.e., discrete embeddings. T…
ResynthesisLearning Source Disentanglement in Neural Audio Codec
Neural audio codecs have significantly advanced audio compression by efficiently converting continuous audio signals into discrete tokens. These codecs preserve high-quality sound and enable sophisticated sound generatio…
Audio CompressionAudio GenerationDisentanglementResynthesisAutomatic Voice Identification after Speech Resynthesis using PPG
Speech resynthesis is a generic task for which we want to synthesize audio with another audio as input, which finds applications for media monitors and journalists.Among different tasks addressed by speech resynthesis, v…
ResynthesisSpeaker VerificationVoice ConversionAnalyzing Speech Unit Selection for Textless Speech-to-Speech Translation
Recent advancements in textless speech-to-speech translation systems have been driven by the adoption of self-supervised learning techniques. Although most state-of-the-art systems adopt a similar architecture to transfo…
Automatic Speech RecognitionEmotion Recognitionfeature selectionResynthesis+7On the Parameter Estimation of Sinusoidal Models for Speech and Audio Signals
In this paper, we examine the parameter estimation performance of three well-known sinusoidal models for speech and audio. The first one is the standard Sinusoidal Model (SM), which is based on the Fast Fourier Transform…
parameter estimationResynthesisNoise Morphing for Audio Time Stretching
This letter introduces an innovative method to enhance the quality of audio time stretching by precisely decomposing a sound into sines, transients, and noise and by improving the processing of the latter component. Whil…
ResynthesisEmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
We introduce EmphAssess, a prosodic benchmark designed to evaluate the capability of speech-to-speech models to encode and reproduce prosodic emphasis. We apply this to two tasks: speech resynthesis and speech-to-speech …
ResynthesisSpeech-to-Speech TranslationTranslationAV2Wav: Diffusion-Based Re-synthesis from Continuous Self-supervised Features for Audio-Visual Speech Enhancement
Speech enhancement systems are typically trained using pairs of clean and noisy speech. In audio-visual speech enhancement (AVSE), there is not as much ground-truth clean data available; most audio-visual datasets are co…
ResynthesisSpeech EnhancementEvaluation of the Speech Resynthesis Capabilities of the VoicePrivacy Challenge Baseline B1
Speaker anonymization systems continue to improve their ability to obfuscate the original speaker characteristics in a speech signal, but often create processing artifacts and unnatural sounding voices as a tradeoff. Man…
ResynthesisSpeaker anonymizationEXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis
Recent work has shown that it is possible to resynthesize high-quality speech based, not on text, but on low bitrate discrete units that have been learned in a self-supervised fashion and can therefore capture expressive…
ResynthesisSpeech SynthesisWeakly-supervised Contrastive Learning for Unsupervised Object Discovery
Unsupervised object discovery (UOD) refers to the task of discriminating the whole region of objects from the background within a scene without relying on labeled datasets, which benefits the task of bounding-box-level l…
Contrastive LearningImage ReconstructionObjectObject Discovery+2Learning Multilingual Expressive Speech Representation for Prosody Prediction without Parallel Data
We propose a method for speech-to-speech emotionpreserving translation that operates at the level of discrete speech units. Our approach relies on the use of multilingual emotion embedding that can capture affective info…
Machine TranslationProsody PredictionResynthesisTranslationIn-the-wild Speech Emotion Conversion Using Disentangled Self-Supervised Representations and Neural Vocoder-based Resynthesis
Speech emotion conversion aims to convert the expressed emotion of a spoken utterance to a target emotion while preserving the lexical information and the speaker's identity. In this work, we specifically focus on in-the…
Resynthesis