paper-with-me

Papers

BERT, can HE predict contrastive focus? Predicting and controlling prominence in neural TTS using a language model

2022-07-04 · Brooke Stephenson, Laurent Besacier, Laurent Girin, Thomas Hueber

Several recent studies have tested the use of transformer language model representations to infer prosodic features for text-to-speech synthesis (TTS). While these studies have explored prosody in general, in this work, we look specifically at the prediction of contrastive focus on personal pronouns. This is a particularly challenging task as it often requires semantic, discursive and/or pragmatic knowledge to predict correctly. We collect a corpus of utterances containing contrastive focus and we evaluate the accuracy of a BERT model, finetuned to predict quantized acoustic prominence features, on these samples. We also investigate how past utterances can provide relevant information for this prediction. Furthermore, we evaluate the controllability of pronoun prominence in a TTS model conditioned on acoustic prominence features.

📄 PDF Abstract BibTeX arXiv:2207.01718

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingSpeech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Adam 설명 없음
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Weight Decay 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…

Similar Papers 제목 키워드 기반

CERT: Contrastive Self-supervised Learning for Language Understanding

2020-05-16 · Hongchao Fang, Sicheng Wang, Meng Zhou, Jiayuan Ding 외

Pretrained language models such as BERT, GPT have shown great effectiveness in language understanding. The auxiliary predictive tasks in existing pretraining approaches are mostly defined on tokens, thus may not be able …

Natural Language UnderstandingSelf-Supervised LearningSentenceTranslation

EnzyCLIP: A Cross-Attention Dual Encoder Framework with Contrastive Learning for Predicting Enzyme Kinetic Constants

2025-11-29 · Anas Aziz Khan, Md Shah Fahad, Priyanka, Ramesh Chandra 외 arxiv

Accurate prediction of enzyme kinetic parameters is crucial for drug discovery, metabolic engineering, and synthetic biology applications. Current computational approaches face limitations in capturing complex enzyme-sub…

Protein Language ModelContrastive LearningDrug Discovery

Predicting Antonyms in Context using BERT

2021-08-01 · INLG (ACL) 2021 8 · Ayana Niwa, Keisuke Nishiguchi, Naoaki Okazaki

We address the task of antonym prediction in a context, which is a fill-in-the-blanks problem. This task setting is unique and practical because it requires contrastiveness to the other word and naturalness as a text in …

Prediction

BERT at SemEval-2020 Task 8: Using BERT to Analyse Meme Emotions

2020-12-01 · SEMEVAL 2020 · Adithya Avvaru, Sanath Vobilisetty

Sentiment analysis, being one of the most sought after research problems within Natural Language Processing (NLP) researchers. The range of problems being addressed by sentiment analysis is increasing. Till now, most of …

Sentiment Analysis

Towards High-fidelity Singing Voice Conversion with Acoustic Reference and Contrastive Predictive Coding

2021-10-10 · Chao Wang, Zhonghao Li, Benlai Tang, Xiang Yin 외

Recently, phonetic posteriorgrams (PPGs) based methods have been quite popular in non-parallel singing voice conversion systems. However, due to the lack of acoustic information in PPGs, style and naturalness of the conv…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1