paper-with-me

Papers

Leveraging WaveNet for Dynamic Listening Head Modeling from Speech

2024-09-08 · Minh-Duc Nguyen, Hyung-Jeong Yang, Seung-Won Kim, Ji-Eun Shin, Soo-Hyung Kim

The creation of listener facial responses aims to simulate interactive communication feedback from a listener during a face-to-face conversation. Our goal is to generate believable videos of listeners' heads that respond authentically to a single speaker by a sequence-to-sequence model with an combination of WaveNet and Long short-term memory network. Our approach focuses on capturing the subtle nuances of listener feedback, ensuring the preservation of individual listener identity while expressing appropriate attitudes and viewpoints. Experiment results show that our method surpasses the baseline models on ViCo benchmark Dataset.

📄 PDF Abstract BibTeX arXiv:2409.05089

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Mixture of Logistic Distributions 설명 없음
Dilated Causal Convolution A Dilated Causal Convolution is a causal convolution where the filter is applied over an area larger than its length by…
WaveNet WaveNet is an audio generative model based on the PixelCNN architecture. In order to deal with long-range temporal dependencies…

Similar Papers 제목 키워드 기반

Speech intelligibility enhancement based on a non-causal Wavenet-like model

2018-09-02 · Interspeech 2018 9 · Muhammed Shifas PV, Vassilis Tsiaras, Yannis Stylianou

Low speech intelligibility in noisy listening conditions makes more difficult our communication with others. Various strate- gies have been suggested to modify a speech signal before it is presented in a noisy listening …

Speaker-independent raw waveform model for glottal excitation

2018-04-25 · Lauri Juvela, Vassilis Tsiaras, Bajibabu Bollepalli, Manu Airaksinen 외

Recent speech technology research has seen a growing interest in using WaveNets as statistical vocoders, i.e., generating speech waveforms from acoustic features. These models have been shown to improve the generated spe…

modelSpeech Synthesistext-to-speechText to Speech+2

Perceptual audio loss function for deep learning

2017-08-20 · Dan Elbaz, Michael Zibulevsky

PESQ and POLQA , are standards are standards for automated assessment of voice quality of speech as experienced by human beings. The predictions of those objective measures should come as close as possible to subjective …

Deep LearningSpeech Enhancement

Parametric Neural Amp Modeling with Active Learning

2025-09-30 · Florian Grötschla, Longxiang Jiao, Luca A. Lanzendörfer, Roger Wattenhofer arxiv

We introduce Panama, an active learning framework to train parametric guitar amp models end-to-end using a combination of an LSTM model and a WaveNet-like architecture. With \model, one can create a virtual amp by record…

Active Learning

Waveform generation for text-to-speech synthesis using pitch-synchronous multi-scale generative adversarial networks

2018-10-30 · Lauri Juvela, Bajibabu Bollepalli, Junichi Yamagishi, Paavo Alku

The state-of-the-art in text-to-speech synthesis has recently improved considerably due to novel neural waveform generation methods, such as WaveNet. However, these methods suffer from their slow sequential inference pro…

Image GenerationSpeech Synthesistext-to-speechText to Speech+2