paper-with-me

홈 › Papers

Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion

2024-04-22 · Yingxuan Li, Ryota Hinami, Kiyoharu Aizawa, Yusuke Matsui

Recognizing characters and predicting speakers of dialogue are critical for comic processing tasks, such as voice generation or translation. However, because characters vary by comic title, supervised learning approaches like training character classifiers which require specific annotations for each comic title are infeasible. This motivates us to propose a novel zero-shot approach, allowing machines to identify characters and predict speaker names based solely on unannotated comic images. In spite of their importance in real-world applications, these task have largely remained unexplored due to challenges in story comprehension and multimodal integration. Recent large language models (LLMs) have shown great capability for text understanding and reasoning, while their application to multimodal content analysis is still an open problem. To address this problem, we propose an iterative multimodal framework, the first to employ multimodal information for both character identification and speaker prediction tasks. Our experiments demonstrate the effectiveness of the proposed framework, establishing a robust baseline for these tasks. Furthermore, since our method requires no training data or annotations, it can be used as-is on any comic series.

📄 PDF Abstract BibTeX arXiv:2404.13993

Code (1)

liyingxuan1012/zeroshot-speaker-prediction 공식 구현 pytorch

Similar Papers 제목 키워드 기반

DeID-VC: Speaker De-identification via Zero-shot Pseudo Voice Conversion

2022-09-09 · Ruibin Yuan, Yuxuan Wu, Jacob Li, Jaxter Kim

The widespread adoption of speech-based online services raises security and privacy concerns regarding the data that they use and share. If the data were compromised, attackers could exploit user speech to bypass speaker…

De-identificationSpeaker VerificationVoice Conversion

Identifying Speakers and Addressees of Quotations in Novels with Prompt Learning

2024-08-18 · Yuchen Yan, Hanjie Zhao, Senbin Zhu, Hongde Liu 외

Quotations in literary works, especially novels, are important to create characters, reflect character relationships, and drive plot development. Current research on quotation extraction in novels primarily focuses on qu…

Prompt Learning

YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone

2021-12-04 · Edresson Casanova, Julian Weber, Christopher Shulby, Arnaldo Candido Junior 외

YourTTS brings the power of a multilingual approach to the task of zero-shot multi-speaker TTS. Our method builds upon the VITS model and adds several novel modifications for zero-shot multi-speaker and multilingual trai…

Speech SynthesisText-To-Speech SynthesisVoice ConversionVoice Similarity+2

AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios

2022-04-01 · Yihan Wu, Xu Tan, Bohan Li, Lei He 외

Adaptive text to speech (TTS) can synthesize new voices in zero-shot scenarios efficiently, by using a well-trained source TTS model without adapting it on the speech data of new speakers. Considering seen and unseen spe…

Speech Synthesistext-to-speechText to Speech

CoLMbo: Speaker Language Model for Descriptive Profiling

2025-06-11 · Massa Baali, Shuo Han, Syed Abdul Hannan, Purusottam Samal 외

Speaker recognition systems are often limited to classification tasks and struggle to generate detailed speaker characteristics or provide context-rich descriptions. These models primarily extract embeddings for speaker …

DescriptiveLanguage ModelingLanguage Modellingmodel+3