paper-with-me

Papers

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing

2025-05-22 · Junjie Zheng, Zihao Chen, Chaofan Ding, Yunming Liang, Yihan Fan, Huan Yang, Lei Xie, Xinhan Di

Current movie dubbing technology can produce the desired speech using a reference voice and input video, maintaining perfect synchronization with the visuals while effectively conveying the intended emotions. However, crucial aspects of movie dubbing, including adaptation to various dubbing styles, effective handling of dialogue, narration, and monologues, as well as consideration of subtle details such as speaker age and gender, remain insufficiently explored. To tackle these challenges, we introduce a multi-modal generative framework. First, it utilizes a multi-modal large vision-language model (VLM) to analyze visual inputs, enabling the recognition of dubbing types and fine-grained attributes. Second, it produces high-quality dubbing using large speech generation models, guided by multi-modal inputs. Additionally, a movie dubbing dataset with annotations for dubbing types and subtle details is constructed to enhance movie understanding and improve dubbing quality for the proposed multi-modal framework. Experimental results across multiple benchmark datasets show superior performance compared to state-of-the-art (SOTA) methods. In details, the LSE-D, SPK-SIM, EMO-SIM, and MCD exhibit improvements of up to 1.09%, 8.80%, 19.08%, and 18.74%, respectively.

📄 PDF Abstract BibTeX arXiv:2505.16279

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Long-range Multimodal Pretraining for Movie Understanding

2023-08-18 · ICCV 2023 1 · Dawit Mureja Argaw, Joon-Young Lee, Markus Woodson, In So Kweon 외

Learning computer vision models from (and for) movies has a long-standing history. While great progress has been attained, there is still a need for a pretrained multimodal model that can perform well in the ever-growing…

A Case Study of Deep Learning Based Multi-Modal Methods for Predicting the Age-Suitability Rating of Movie Trailers

2021-01-26 · Mahsa Shafaei, Christos Smailis, Ioannis A. Kakadiaris, Thamar Solorio

In this work, we explore different approaches to combine modalities for the problem of automated age-suitability rating of movie trailers. First, we introduce a new dataset containing videos of movie trailers in English …

A Case Study of Deep Learning-Based Multi-Modal Methods for Labeling the Presence of Questionable Content in Movie Trailers

2021-09-01 · RANLP 2021 9 · Mahsa Shafaei, Christos Smailis, Ioannis Kakadiaris, Thamar Solorio

In this work, we explore different approaches to combine modalities for the problem of automated age-suitability rating of movie trailers. First, we introduce a new dataset containing videos of movie trailers in English …

Movie Recommendation with Poster Attention via Multi-modal Transformer Feature Fusion

2024-07-12 · Linhan Xia, Yicheng Yang, Ziou Chen, Zheng Yang 외

Pre-trained models learn general representations from large datsets which can be fine-turned for specific tasks to significantly reduce training time. Pre-trained models like generative pretrained transformers (GPT), bid…

Movie Recommendation

Does a Technique for Building Multimodal Representation Matter? -- Comparative Analysis

2022-06-09 · Maciej Pawłowski, Anna Wróblewska, Sylwia Sysko-Romańczuk

Creating a meaningful representation by fusing single modalities (e.g., text, images, or audio) is the core concept of multimodal learning. Although several techniques for building multimodal representations have been pr…