paper-with-me

홈 › Papers

Emotion-Aware Speech Generation with Character-Specific Voices for Comics

2025-09-18 · Zhiwen Qian, Jinhua Liang, Huan Zhang arxiv

This paper presents an end-to-end pipeline for generating character-specific, emotion-aware speech from comics. The proposed system takes full comic volumes as input and produces speech aligned with each character's dialogue and emotional state. An image processing module performs character detection, text recognition, and emotion intensity recognition. A large language model performs dialogue attribution and emotion analysis by integrating visual information with the evolving plot context. Speech is synthesized through a text-to-speech model with distinct voice profiles tailored to each character and emotion. This work enables automated voiceover generation for comics, offering a step toward interactive and immersive comic reading experience.

📄 PDF Abstract BibTeX arXiv:2509.15253

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Emotion-Controllable Generalized Talking Face Generation

2022-05-02 · Sanjana Sinha, Sandika Biswas, Ravindra Yadav, Brojeshwar Bhowmick

Despite the significant progress in recent years, very few of the AI-based talking face generation methods attempt to render natural emotions. Moreover, the scope of the methods is majorly limited to the characteristics …

Face GenerationOptical Flow EstimationTalking Face GenerationTexture Synthesis

Large Language Models Meet Contrastive Learning: Zero-Shot Emotion Recognition Across Languages

2025-03-25 · Heqing Zou, Fengmao Lv, Desheng Zheng, Eng Siong Chng 외

Multilingual speech emotion recognition aims to estimate a speaker's emotional state using a contactless method across different languages. However, variability in voice characteristics and linguistic diversity poses sig…

Contrastive LearningDiversityEmotion RecognitionSpeech Emotion Recognition

Characteristic-Specific Partial Fine-Tuning for Efficient Emotion and Speaker Adaptation in Codec Language Text-to-Speech Models

2025-01-24 · Tianrui Wang, Meng Ge, Cheng Gong, Chunyu Qiang 외

Recently, emotional speech generation and speaker cloning have garnered significant interest in text-to-speech (TTS). With the open-sourcing of codec language TTS models trained on massive datasets with large-scale param…

Emotion ClassificationSpeaker Identificationtext-to-speechText to Speech

EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis

2026-01-30 · Li Zhou, Hao Jiang, Junjie Li, Tianrui Wang 외 arxiv

Achieving precise and controllable emotional expression is crucial for producing natural and context-appropriate speech in text-to-speech (TTS) synthesis. However, many emotion-aware TTS systems, including large language…

Speech Synthesis

Emotional Face-to-Speech

2025-02-03 · Jiaxin Ye, Boyuan Cao, Hongming Shan

How much can we infer about an emotional voice solely from an expressive face? This intriguing question holds great potential for applications such as virtual character dubbing and aiding individuals with expressive lang…