paper-with-me

Papers

HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

2025-05-07 · Teng Hu, Zhentao Yu, Zhengguang Zhou, Sen Liang, Yuan Zhou, Qin Lin, Qinglin Lu

Customized video generation aims to produce videos featuring specific subjects under flexible user-defined conditions, yet existing methods often struggle with identity consistency and limited input modalities. In this paper, we propose HunyuanCustom, a multi-modal customized video generation framework that emphasizes subject consistency while supporting image, audio, video, and text conditions. Built upon HunyuanVideo, our model first addresses the image-text conditioned generation task by introducing a text-image fusion module based on LLaVA for enhanced multi-modal understanding, along with an image ID enhancement module that leverages temporal concatenation to reinforce identity features across frames. To enable audio- and video-conditioned generation, we further propose modality-specific condition injection mechanisms: an AudioNet module that achieves hierarchical alignment via spatial cross-attention, and a video-driven injection module that integrates latent-compressed conditional video through a patchify-based feature-alignment network. Extensive experiments on single- and multi-subject scenarios demonstrate that HunyuanCustom significantly outperforms state-of-the-art open- and closed-source methods in terms of ID consistency, realism, and text-video alignment. Moreover, we validate its robustness across downstream tasks, including audio and video-driven customized video generation. Our results highlight the effectiveness of multi-modal conditioning and identity-preserving strategies in advancing controllable video generation. All the code and models are available at https://hunyuancustom.github.io.

📄 PDF Abstract BibTeX arXiv:2505.04512

Code (1)

tencent-hunyuan/hunyuancustom pytorch

Tasks

Human-Domain Subject-to-VideoSingle-Domain Subject-to-VideoVideo AlignmentVideo Generation

Similar Papers 제목 키워드 기반

MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models

2024-12-02 · Xiaomin Li, Xu Jia, Qinghe Wang, Haiwen Diao 외

Existing pretrained text-to-video (T2V) models have demonstrated impressive abilities in generating realistic videos with basic motion or camera movement. However, these models exhibit significant limitations when genera…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+1

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation

2026-05-17 · Yuheng Chen, Qingdong He, Teng Hu, Yuji Wang 외 arxiv

The landscape of joint audio and video generation has been fundamentally transformed by the advent of powerful foundation models. Despite these strides, achieving cohesive multimodal customization for the simultaneous pr…

Video Generation

A Comprehensive Ecosystem for Open-Domain Customized Video Generation

2026-06-10 · Jingxu Zhang, Yuqian Hong, Daneul Kim, Kai Qiu 외 arxiv

Recent progress in video generation has shown impressive visual synthesis capabilities. However, open-domain customized video generation remains limited by the lack of large-scale, annotated datasets capturing diverse id…

Video Generation

Deep CNN with late fusion for realtime multimodal emotion recognition

2024-04-15 · Expert Systems with Applications 2024 4 · Chhavi Dixita, Shashank Mouli Satapathy

Emotion recognition is a fundamental aspect of human communication and plays a crucial role in various domains. This project aims at developing an efficient model for real-time multimodal emotion recognition in video…

Emotion RecognitionMultimodal Emotion Recognition

GroupVideo: Multi-Identity Customized Text-to-Video Generation

2026-07-23 · Xinyang Song, Libin Wang, Jianxin Sun, Qi Li 외 arxiv

Current identity customized video generation methodologies are predominantly limited to single-identity scenarios, as the lack of explicit identity separation mechanisms often leads to identity confusion in multi-identit…

Text-to-Video Generation