paper-with-me

홈 › Papers

Unispeaker: A Unified Approach for Multimodality-driven Speaker Generation

2025-01-11 · Zhengyan Sheng, Zhihao Du, Heng Lu, Shiliang Zhang, Zhen-Hua Ling

Recent advancements in personalized speech generation have brought synthetic speech increasingly close to the realism of target speakers' recordings, yet multimodal speaker generation remains on the rise. This paper introduces UniSpeaker, a unified approach for multimodality-driven speaker generation. Specifically, we propose a unified voice aggregator based on KV-Former, applying soft contrastive loss to map diverse voice description modalities into a shared voice space, ensuring that the generated voice aligns more closely with the input descriptions. To evaluate multimodality-driven voice control, we build the first multimodality-based voice control (MVC) benchmark, focusing on voice suitability, voice diversity, and speech quality. UniSpeaker is evaluated across five tasks using the MVC benchmark, and the experimental results demonstrate that UniSpeaker outperforms previous modality-specific models. Speech samples are available at \url{https://UniSpeaker.github.io}.

📄 PDF Abstract BibTeX arXiv:2501.06394

Code (0)

등록된 구현이 없습니다.

Tasks

Diversity

Similar Papers 제목 키워드 기반

FashionEngine: Interactive 3D Human Generation and Editing via Multimodal Controls

2024-04-02 · Tao Hu, Fangzhou Hong, Zhaoxi Chen, Ziwei Liu

We present FashionEngine, an interactive 3D human generation and editing system that creates 3D digital humans via user-friendly multimodal controls such as natural languages, visual perceptions, and hand-drawing sketche…

Virtual Try-on

Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness

2024-01-07 · Sicheng Yang, Zunnan Xu, Haiwei Xue, Yongkang Cheng 외

Current talking avatars mostly generate co-speech gestures based on audio and text of the utterance, without considering the non-speaking motion of the speaker. Furthermore, previous works on co-speech gesture generation…

Gesture GenerationMotion Generation

UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking

2025-12-10 · Xuangeng Chu, Ruicong Liu, Yifei Huang, Yun Liu 외 arxiv

Generating lifelike conversational avatars requires modeling not just isolated speakers, but the dynamic, reciprocal interaction of speaking and listening. However, modeling the listener is exceptionally challenging: dir…

UniFLG: Unified Facial Landmark Generator from Text or Speech

2023-02-28 · Kentaro Mitsui, Yukiya Hono, Kei Sawada

Talking face generation has been extensively investigated owing to its wide applicability. The two primary frameworks used for talking face generation comprise a text-driven framework, which generates synchronized speech…

DecoderFace GenerationSpeech SynthesisTalking Face Generation+2

Audio-Driven Co-Speech Gesture Video Generation

2022-12-05 · Xian Liu, Qianyi Wu, Hang Zhou, Yuanqi Du 외

Co-speech gesture is crucial for human-machine interaction and digital entertainment. While previous works mostly map speech audio to human skeletons (e.g., 2D keypoints), directly generating speakers' gestures in the im…

Video Generation