paper-with-me

Papers

VideoRepair: Improving Text-to-Video Generation via Misalignment Evaluation and Localized Refinement

2024-11-22 · Daeun Lee, Jaehong Yoon, Jaemin Cho, Mohit Bansal

Recent text-to-video (T2V) diffusion models have demonstrated impressive generation capabilities across various domains. However, these models often generate videos that have misalignments with text prompts, especially when the prompts describe complex scenes with multiple objects and attributes. To address this, we introduce VideoRepair, a novel model-agnostic, training-free video refinement framework that automatically identifies fine-grained text-video misalignments and generates explicit spatial and textual feedback, enabling a T2V diffusion model to perform targeted, localized refinements. VideoRepair consists of four stages: In (1) video evaluation, we detect misalignments by generating fine-grained evaluation questions and answering those questions with MLLM. In (2) refinement planning, we identify accurately generated objects and then create localized prompts to refine other areas in the video. Next, in (3) region decomposition, we segment the correctly generated area using a combined grounding module. We regenerate the video by adjusting the misaligned regions while preserving the correct regions in (4) localized refinement. On two popular video generation benchmarks (EvalCrafter and T2V-CompBench), VideoRepair substantially outperforms recent baselines across various text-video alignment metrics. We provide a comprehensive analysis of VideoRepair components and qualitative examples.

📄 PDF Abstract BibTeX arXiv:2411.15115

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video GenerationVideo AlignmentVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

VideoGen-Eval: Agent-based System for Video Generation Evaluation

2025-03-30 · Yuhang Yang, Ke Fan, Shangkun Sun, Hongxiang Li 외

The rapid advancement of video generation has rendered existing evaluation systems inadequate for assessing state-of-the-art models, primarily due to simple prompts that cannot showcase the model's capabilities, fixed ev…

DiversityVideo Generation

MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation

2025-02-03 · Haibo Tong, Zhaoyang Wang, Zhaorun Chen, Haonian Ji 외

Recent advancements in video generation have significantly improved the ability to synthesize videos from text instructions. However, existing models still struggle with key challenges such as instruction misalignment, c…

BenchmarkingFairnessHallucinationMixture-of-Experts+1

DeepSound-V1: Start to Think Step-by-Step in the Audio Generation from Videos

2025-03-28 · Yunming Liang, Zihao Chen, Chaofan Ding, Xinhan Di

Currently, high-quality, synchronized audio is synthesized from video and optional text inputs using various multi-modal joint learning frameworks. However, the precise alignment between the visual and generated audio do…

Audio GenerationLarge Language Model

MM-Sonate: Multimodal Controllable Audio-Video Generation with Zero-Shot Voice Cloning

2026-01-04 · Chunyu Qiang, Jun Wang, Xiaopeng Wang, Kang Yin 외 arxiv

Joint audio-video generation aims to synthesize synchronized multisensory content, yet current unified models struggle with fine-grained acoustic control, particularly for identity-preserving speech. Existing approaches …

Video Generation

Aligning Anime Video Generation with Human Feedback

2025-04-14 · Bingwen Zhu, Yudong Jiang, Baohan Xu, Siqian Yang 외

Anime video generation faces significant challenges due to the scarcity of anime data and unusual motion patterns, leading to issues such as motion distortion and flickering artifacts, which result in misalignment with h…

Video Generation