paper-with-me

홈 › Papers

Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation

2025-08-14 · Harold Haodong Chen, Haojian Huang, Qifeng Chen, Harry Yang, Ser-Nam Lim arxiv

Recent advancements in video generation have enabled the creation of high-quality, visually compelling videos. However, generating videos that adhere to the laws of physics remains a critical challenge for applications requiring realism and accuracy. In this work, we propose PhysHPO, a novel framework for Hierarchical Cross-Modal Direct Preference Optimization, to tackle this challenge by enabling fine-grained preference alignment for physically plausible video generation. PhysHPO optimizes video alignment across four hierarchical granularities: a) Instance Level, aligning the overall video content with the input prompt; b) State Level, ensuring temporal consistency using boundary frames as anchors; c) Motion Level, modeling motion trajectories for realistic dynamics; and d) Semantic Level, maintaining logical consistency between narrative and visuals. Recognizing that real-world videos are the best reflections of physical phenomena, we further introduce an automated data selection pipeline to efficiently identify and utilize "good data" from existing large-scale text-video datasets, thereby eliminating the need for costly and time-intensive dataset construction. Extensive experiments on both physics-focused and general capability benchmarks demonstrate that PhysHPO significantly improves physical plausibility and overall video generation quality of advanced models. To the best of our knowledge, this is the first work to explore fine-grained preference alignment and data selection for video generation, paving the way for more realistic and human-preferred video generation paradigms.

📄 PDF Abstract BibTeX arXiv:2508.10858

Code (0)

등록된 구현이 없습니다.

Tasks

Video GenerationVideo Alignment

Similar Papers 제목 키워드 기반

Beyond Binary Preference: Aligning Diffusion Models to Fine-grained Criteria by Decoupling Attributes

2026-01-07 · Chenye Meng, Zejian Li, Zhongni Liu, Yize Li 외 arxiv

Post-training alignment of diffusion models relies on simplified signals, such as scalar rewards or binary preferences. This limits alignment with complex human expertise, which is hierarchical and fine-grained. To addre…

Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM Agents

2025-09-26 · Heyang Gao, Zexu Sun, Erxue Min, Hengyi Cai 외 arxiv

Large Language Models (LLMs) as autonomous agents are increasingly tasked with solving complex, long-horizon problems. Aligning these agents via preference-based offline methods like Direct Preference Optimization (DPO) …

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

2025-04-17 · Haojian Huang, Haodong Chen, Shengqiong Wu, Meng Luo 외

Large Video Models (LVMs) built upon Large Language Models (LLMs) have shown promise in video understanding but often suffer from misalignment with human intuition and video hallucination issues. To address these challen…

HallucinationVideo Understanding

HLG: Comprehensive 3D Room Construction via Hierarchical Layout Generation

2025-08-25 · Xiping Wang, Yuxi Wang, Mengqi Zhou, Junsong Fan 외 arxiv

Realistic 3D indoor scene generation is crucial for virtual reality, interior design, embodied intelligence, and scene understanding. While existing methods have made progress in coarse-scale furniture arrangement, they …

Scene UnderstandingScene Generation

Hierarchical Latent Reasoning for LLM-based Recommendation

2026-07-30 · Peiyu Hu, Siying Gu, Weihai Lu, Zhuodong Liu 외 arxiv

Large Language Models (LLMs) have shown strong potential for recommendation by leveraging their semantic understanding and contextual modeling capabilities. Recent studies further introduce reasoning mechanisms to improv…

Representation Learning