paper-with-me

Papers Resynthesis

“Resynthesis” 태그가 달린 논문 51편 · 필터 해제

Spoken Language Modeling with Duration-Penalized Self-Supervised Units

2025-05-29 · Nicol Visser, Herman Kamper

Spoken language models (SLMs) operate on acoustic units obtained by discretizing self-supervised speech representations. Although the characteristics of these units directly affect performance, the interaction between co…

Language ModelingLanguage ModellingResynthesis

Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information

2025-05-21 · Nicholas Sanders, Yuanchao Li, Korin Richmond, Simon King

Quantization in SSL speech models (e.g., HuBERT) improves compression and performance in tasks like language modeling, resynthesis, and text-to-speech but often discards prosodic and paralinguistic information (e.g., emo…

Language ModelingLanguage ModellingQuantizationResynthesis+2

Dynamic Tsetlin Machine Accelerators for On-Chip Training at the Edge using FPGAs

2025-04-28 · Gang Mao, Tousif Rahman, Sidharth Maheshwari, Bob Pattison 외

The increased demand for data privacy and security in machine learning (ML) applications has put impetus on effective edge training on Internet-of-Things (IoT) nodes. Edge training aims to leverage speed, energy efficien…

Resynthesis

Runtime Tunable Tsetlin Machines for Edge Inference on eFPGAs

2025-02-10 · Tousif Rahman, Gang Mao, Bob Pattison, Sidharth Maheshwari 외

Embedded Field-Programmable Gate Arrays (eFPGAs) allow for the design of hardware accelerators of edge Machine Learning (ML) applications at a lower power budget compared with traditional FPGA platforms. However, the lim…

Model CompressionResynthesis

FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks

2025-02-06 · Luca Della Libera, Francesco Paissan, Cem Subakan, Mirco Ravanelli

Large language models have revolutionized natural language processing through self-supervised pretraining on massive datasets. Inspired by this success, researchers have explored adapting these methods to speech by discr…

ResynthesisVoice Conversion

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder

2025-01-09 · Samir Sadok, Simon Leglaive, Laurent Girin, Gaël Richard 외

This article introduces AnCoGen, a novel method that leverages a masked autoencoder to unify the analysis, control, and generation of speech signals within a single model. AnCoGen can analyze speech by estimating key att…

Pitch ClassificationPitch controlResynthesisSpeech Enhancement+1

DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models

2024-10-31 · Heng-Jui Chang, Hongyu Gong, Changhan Wang, James Glass 외

Spoken language models (SLMs) have gained increasing attention with advancements in text-based, decoder-only language models. SLMs process text and speech, enabling simultaneous speech understanding and generation. This …

DecoderResynthesisSpeech Tokenization

A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation

2024-10-29 · Alexander H. Liu, Qirui Wang, Yuan Gong, James Glass

Neural Audio Codecs, initially designed as a compression technique, have gained more attention recently for speech generation. Codec models represent each audio frame as a sequence of tokens, i.e., discrete embeddings. T…

Resynthesis

Learning Source Disentanglement in Neural Audio Codec

2024-09-17 · Xiaoyu Bie, Xubo Liu, Gaël Richard

Neural audio codecs have significantly advanced audio compression by efficiently converting continuous audio signals into discrete tokens. These codecs preserve high-quality sound and enable sophisticated sound generatio…

Audio CompressionAudio GenerationDisentanglementResynthesis

Automatic Voice Identification after Speech Resynthesis using PPG

2024-08-05 · Thibault Gaudier, Marie Tahon, Anthony Larcher, Yannick Estève

Speech resynthesis is a generic task for which we want to synthesize audio with another audio as input, which finds applications for media monitors and journalists.Among different tasks addressed by speech resynthesis, v…

ResynthesisSpeaker VerificationVoice Conversion

Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation

2024-07-08 · Jarod Duret, Yannick Estève, Titouan Parcollet

Recent advancements in textless speech-to-speech translation systems have been driven by the adoption of self-supervised learning techniques. Although most state-of-the-art systems adopt a similar architecture to transfo…

Automatic Speech RecognitionEmotion Recognitionfeature selectionResynthesis+7

On the Parameter Estimation of Sinusoidal Models for Speech and Audio Signals

2024-01-02 · George P. Kafentzis

In this paper, we examine the parameter estimation performance of three well-known sinusoidal models for speech and audio. The first one is the standard Sinusoidal Model (SM), which is based on the Fast Fourier Transform…

parameter estimationResynthesis

Noise Morphing for Audio Time Stretching

2023-12-22 · Eloi Moliner, Leonardo Fierro, Alec Wright, Matti Hämäläinen 외

This letter introduces an innovative method to enhance the quality of audio time stretching by precisely decomposing a sound into sines, transients, and noise and by improving the processing of the latter component. Whil…

Resynthesis

EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models

2023-12-21 · Maureen de Seyssel, Antony D'Avirro, Adina Williams, Emmanuel Dupoux

We introduce EmphAssess, a prosodic benchmark designed to evaluate the capability of speech-to-speech models to encode and reproduce prosodic emphasis. We apply this to two tasks: speech resynthesis and speech-to-speech …

ResynthesisSpeech-to-Speech TranslationTranslation

AV2Wav: Diffusion-Based Re-synthesis from Continuous Self-supervised Features for Audio-Visual Speech Enhancement

2023-09-14 · Ju-chieh Chou, Chung-Ming Chien, Karen Livescu

Speech enhancement systems are typically trained using pairs of clean and noisy speech. In audio-visual speech enhancement (AVSE), there is not as much ground-truth clean data available; most audio-visual datasets are co…

ResynthesisSpeech Enhancement

Evaluation of the Speech Resynthesis Capabilities of the VoicePrivacy Challenge Baseline B1

2023-08-22 · Ünal Ege Gaznepoglu, Nils Peters

Speaker anonymization systems continue to improve their ability to obfuscate the original speaker characteristics in a speech signal, but often create processing artifacts and unnatural sounding voices as a tradeoff. Man…

ResynthesisSpeaker anonymization

EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

2023-08-10 · Tu Anh Nguyen, Wei-Ning Hsu, Antony D'Avirro, Bowen Shi 외

Recent work has shown that it is possible to resynthesize high-quality speech based, not on text, but on low bitrate discrete units that have been learned in a self-supervised fashion and can therefore capture expressive…

ResynthesisSpeech Synthesis

Weakly-supervised Contrastive Learning for Unsupervised Object Discovery

2023-07-07 · Yunqiu Lv, Jing Zhang, Nick Barnes, Yuchao Dai

Unsupervised object discovery (UOD) refers to the task of discriminating the whole region of objects from the background within a scene without relying on labeled datasets, which benefits the task of bounding-box-level l…

Contrastive LearningImage ReconstructionObjectObject Discovery+2

Learning Multilingual Expressive Speech Representation for Prosody Prediction without Parallel Data

2023-06-29 · Jarod Duret, Titouan Parcollet, Yannick Estève

We propose a method for speech-to-speech emotionpreserving translation that operates at the level of discrete speech units. Our approach relies on the use of multilingual emotion embedding that can capture affective info…

Machine TranslationProsody PredictionResynthesisTranslation

In-the-wild Speech Emotion Conversion Using Disentangled Self-Supervised Representations and Neural Vocoder-based Resynthesis

2023-06-02 · Navin Raj Prabhu, Nale Lehmann-Willenbrock, Timo Gerkmann

Speech emotion conversion aims to convert the expressed emotion of a spoken utterance to a target emotion while preserving the lexical information and the speaker's identity. In this work, we specifically focus on in-the…

Resynthesis
1–20 / 51 다음 →