paper-with-me

Papers

Factor-Conditioned Speaking-Style Captioning

2024-06-27 · Atsushi Ando, Takafumi Moriya, Shota Horiguchi, Ryo Masumura

This paper presents a novel speaking-style captioning method that generates diverse descriptions while accurately predicting speaking-style information. Conventional learning criteria directly use original captions that contain not only speaking-style factor terms but also syntax words, which disturbs learning speaking-style information. To solve this problem, we introduce factor-conditioned captioning (FCC), which first outputs a phrase representing speaking-style factors (e.g., gender, pitch, etc.), and then generates a caption to ensure the model explicitly learns speaking-style factors. We also propose greedy-then-sampling (GtS) decoding, which first predicts speaking-style factors deterministically to guarantee semantic accuracy, and then generates a caption based on factor-conditioned sampling to ensure diversity. Experiments show that FCC outperforms the original caption-based training, and with GtS, it generates more diverse captions while keeping style prediction performance.

📄 PDF Abstract BibTeX arXiv:2406.18910

Code (0)

등록된 구현이 없습니다.

Tasks

Diversity

Similar Papers 제목 키워드 기반

StyleCap: Automatic Speaking-Style Captioning from Speech Based on Speech and Language Self-supervised Learning Models

2023-11-28 · Kazuki Yamauchi, Yusuke Ijima, Yuki Saito

We propose StyleCap, a method to generate natural language descriptions of speaking styles appearing in speech. Although most of conventional techniques for para-/non-linguistic information recognition focus on the categ…

DecoderDiversityLanguage ModelingLanguage Modelling+3

LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning

2024-06-12 · Masaya Kawamura, Ryuichi Yamamoto, Yuma Shirahata, Takuya Hasumi 외

We introduce LibriTTS-P, a new corpus based on LibriTTS-R that includes utterance-level descriptions (i.e., prompts) of speaking style and speaker-level prompts of speaker characteristics. We employ a hybrid approach to …

text-to-speechText to Speech

SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning

2024-08-25 · Chien-yu Huang, Min-Han Shih, Ke-Han Lu, Chi-Yuan Hsiao 외

Instruction-based speech processing is becoming popular. Studies show that training with multiple tasks boosts performance, but collecting diverse, large-scale tasks and datasets is expensive. Thus, it is highly desirabl…

Emotion Recognition

Say Anything with Any Style

2024-03-11 · Shuai Tan, Bin Ji, Yu Ding, Ye Pan

Generating stylized talking head with diverse head motions is crucial for achieving natural-looking videos but still remains challenging. Previous works either adopt a regressive method to capture the speaking style, res…

Voice Conversion for Lombard Speaking Style with Implicit and Explicit Acoustic Feature Conditioning

2025-07-12 · Dominika Woszczyk, Manuel Sam Ribeiro, Thomas Merritt, Daniel Korzekwa arxiv

Text-to-Speech (TTS) systems in Lombard speaking style can improve the overall intelligibility of speech, useful for hearing loss and noisy conditions. However, training those models requires a large amount of data and t…

Voice ConversionStyle Transfer