paper-with-me

Papers

StyleCap: Automatic Speaking-Style Captioning from Speech Based on Speech and Language Self-supervised Learning Models

2023-11-28 · Kazuki Yamauchi, Yusuke Ijima, Yuki Saito

We propose StyleCap, a method to generate natural language descriptions of speaking styles appearing in speech. Although most of conventional techniques for para-/non-linguistic information recognition focus on the category classification or the intensity estimation of pre-defined labels, they cannot provide the reasoning of the recognition result in an interpretable manner. StyleCap is a first step towards an end-to-end method for generating speaking-style prompts from speech, i.e., automatic speaking-style captioning. StyleCap is trained with paired data of speech and natural language descriptions. We train neural networks that convert a speech representation vector into prefix vectors that are fed into a large language model (LLM)-based text decoder. We explore an appropriate text decoder and speech feature representation suitable for this new task. The experimental results demonstrate that our StyleCap leveraging richer LLMs for the text decoder, speech self-supervised learning (SSL) features, and sentence rephrasing augmentation improves the accuracy and diversity of generated speaking-style captions. Samples of speaking-style captions generated by our StyleCap are publicly available.

📄 PDF Abstract BibTeX arXiv:2311.16509

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderDiversityLanguage ModelingLanguage ModellingLarge Language ModelSelf-Supervised LearningSentence

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning

2024-08-25 · Chien-yu Huang, Min-Han Shih, Ke-Han Lu, Chi-Yuan Hsiao 외

Instruction-based speech processing is becoming popular. Studies show that training with multiple tasks boosts performance, but collecting diverse, large-scale tasks and datasets is expensive. Thus, it is highly desirabl…

Emotion Recognition

Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences

2023-07-31 · Dingyi Yang, Hongyu Chen, Xinglin Hou, Tiezheng Ge 외

Stylized visual captioning aims to generate image or video descriptions with specific styles, making them more attractive and emotionally appropriate. One major challenge with this task is the lack of paired stylized cap…

DecoderImage CaptioningLanguage Modelling

LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning

2024-06-12 · Masaya Kawamura, Ryuichi Yamamoto, Yuma Shirahata, Takuya Hasumi 외

We introduce LibriTTS-P, a new corpus based on LibriTTS-R that includes utterance-level descriptions (i.e., prompts) of speaking style and speaker-level prompts of speaker characteristics. We employ a hybrid approach to …

text-to-speechText to Speech

Factor-Conditioned Speaking-Style Captioning

2024-06-27 · Atsushi Ando, Takafumi Moriya, Shota Horiguchi, Ryo Masumura

This paper presents a novel speaking-style captioning method that generates diverse descriptions while accurately predicting speaking-style information. Conventional learning criteria directly use original captions that …

Diversity

Audio-Aware Large Language Models as Judges for Speaking Styles

2025-06-06 · Cheng-Han Chiang, Xiaofei Wang, Chung-Ching Lin, Kevin Lin 외

Audio-aware large language models (ALLMs) can understand the textual and non-textual information in the audio input. In this paper, we explore using ALLMs as an automatic judge to assess the speaking styles of speeches. …

Instruction FollowingPitch control