paper-with-me

홈 › Papers

OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation

2026-04-20 · Lei Zhu, Xing Cai, Yingjie Chen, Yiheng Li, Binxin Yang, Hao Liu, Jie Chen, Chen Li, Jing LYu arxiv

Recent advancements in audio-video joint generation models have demonstrated impressive capabilities in content creation. However, generating high-fidelity human-centric videos in complex, real-world physical scenes remains a significant challenge. We identify that the root cause lies in the structural deficiencies of existing datasets across three dimensions: limited global scene and camera diversity, sparse interaction modeling (both person-person and person-object), and insufficient individual attribute alignment. To bridge these gaps, we present OmniHuman, a large-scale, multi-scene dataset designed for fine-grained human modeling. OmniHuman provides a hierarchical annotation covering video-level scenes, frame-level interactions, and individual-level attributes. To facilitate this, we develop a fully automated pipeline for high-quality data collection and multi-modal annotation. Complementary to the dataset, we establish the OmniHuman Benchmark (OHBench), a three-level evaluation system that provides a scientific diagnosis for human-centric audio-video synthesis. Crucially, OHBench introduces metrics that are highly consistent with human perception, filling the gaps in existing benchmarks by providing a comprehensive diagnosis across global scenes, relational interactions, and individual attributes.

📄 PDF Abstract BibTeX arXiv:2604.18326

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

2025-02-03 · Gaojie Lin, Jianwen Jiang, Jiaqi Yang, Zerong Zheng 외

End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generatio…

Human AnimationHuman-Object Interaction DetectionMotion GenerationVideo Generation

OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation

2026-05-12 · Yiren Song, Xiyao Deng, Pei Yang, Yihan Wang 외 arxiv

Cross-embodiment video generation aims to transfer motions across different humanoid embodiments, such as human-to-robot and robot-to-robot, enabling scalable data generation for embodied intelligence. A major challenge …

Video Generation

OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation

2025-08-26 · Jianwen Jiang, Weihong Zeng, Zerong Zheng, Jiaqi Yang 외 arxiv

Existing video avatar models can produce fluid human animations, yet they struggle to move beyond mere physical likeness to capture a character's authentic essence. Their motions typically synchronize with low-level cues…

LongCat-Video-Avatar 1.5 Technical Report

2026-05-26 · Meituan LongCat Team, Xunliang Cai, Meng Cheng, Feng Gao 외 arxiv

Despite advances in audio-driven video generation, achieving commercial-grade stability remains challenging. We present LongCat-Video-Avatar 1.5, an upgraded open-source framework prioritizing systematic engineering and …

Video Generation

Wan-S2V: Audio-Driven Cinematic Video Generation

2025-08-26 · Xin Gao, Li Hu, Siqi Hu, Mingyang Huang 외 arxiv

Current state-of-the-art (SOTA) methods for audio-driven character animation demonstrate promising performance for scenarios primarily involving speech and singing. However, they often fall short in more complex film and…

Video Generation