paper-with-me

홈 › Papers

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

2025-07-28 · Yuying Ge, Yixiao Ge, Chen Li, Teng Wang, Junfu Pu, Yizhuo Li, Lu Qiu, Jin Ma, Lisheng Duan, Xinyu Zuo, Jinwen Luo, Weibo Gu, Zexuan Li, Xiaojing Zhang, Yangyu Tao, Han Hu, Di Wang, Ying Shan arxiv

Real-world user-generated short videos, especially those distributed on platforms such as WeChat Channel and TikTok, dominate the mobile internet. However, current large multimodal models lack essential temporally-structured, detailed, and in-depth video comprehension capabilities, which are the cornerstone of effective video search and recommendation, as well as emerging video applications. Understanding real-world shorts is actually challenging due to their complex visual elements, high information density in both visuals and audio, and fast pacing that focuses on emotional expression and viewpoint delivery. This requires advanced reasoning to effectively integrate multimodal information, including visual, audio, and text. In this work, we introduce ARC-Hunyuan-Video, a multimodal model that processes visual, audio, and textual signals from raw video inputs end-to-end for structured comprehension. The model is capable of multi-granularity timestamped video captioning and summarization, open-ended video question answering, temporal video grounding, and video reasoning. Leveraging high-quality data from an automated annotation pipeline, our compact 7B-parameter model is trained through a comprehensive regimen: pre-training, instruction fine-tuning, cold start, reinforcement learning (RL) post-training, and final instruction fine-tuning. Quantitative evaluations on our introduced benchmark ShortVid-Bench and qualitative comparisons demonstrate its strong performance in real-world video comprehension, and it supports zero-shot or fine-tuning with a few samples for diverse downstream applications. The real-world production deployment of our model has yielded tangible and measurable improvements in user engagement and satisfaction, a success supported by its remarkable efficiency, with stress tests indicating an inference time of just 10 seconds for a one-minute video on H20 GPU.

📄 PDF Abstract BibTeX arXiv:2507.20939

Code (0)

등록된 구현이 없습니다.

Tasks

Video Question AnsweringReinforcement LearningVideo CaptioningVideo Grounding

Similar Papers 제목 키워드 기반

HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

2025-05-07 · Teng Hu, Zhentao Yu, Zhengguang Zhou, Sen Liang 외

Customized video generation aims to produce videos featuring specific subjects under flexible user-defined conditions, yet existing methods often struggle with identity consistency and limited input modalities. In this p…

Human-Domain Subject-to-VideoSingle-Domain Subject-to-VideoVideo AlignmentVideo Generation

HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

2025-08-23 · Sizhe Shan, Qiulin Li, Yutao Cui, Miles Yang 외 arxiv

Recent advances in video generation produce visually realistic content, yet the absence of synchronized audio severely compromises immersion. To address key challenges in video-to-audio generation, including multimodal d…

Audio GenerationVideo Generation

HunyuanVideo 1.5 Technical Report

2025-11-24 · Bing Wu, Chang Zou, Changlin Li, Duojun Huang 외 arxiv

We present HunyuanVideo 1.5, a lightweight yet powerful open-source video generation model that achieves state-of-the-art visual quality and motion coherence with only 8.3 billion parameters, enabling efficient inference…

Video Super-ResolutionVideo Generation

HunyuanVideo: A Systematic Framework For Large Video Generative Models

2024-12-03 · Weijie Kong, Qi Tian, Zijian Zhang, Rox Min 외

Recent advancements in video generation have significantly impacted daily life for both individuals and industries. However, the leading video generation models remain closed-source, resulting in a notable performance ga…

Video AlignmentVideo Generation

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

2025-05-26 · Yi Chen, Sen Liang, Zixiang Zhou, Ziyao Huang 외

Recent years have witnessed significant progress in audio-driven human animation. However, critical challenges remain in (i) generating highly dynamic videos while preserving character consistency, (ii) achieving precise…

Human Animation