paper-with-me

홈 › Papers

VOR-Bench: A Human Perception-Driven Benchmark for Video Object Removal

2026-09-15 · Haonan Huang, Tianrui Qiu, Xianghao Zang, Yinan Du, Zhixiang He, Chi Zhang, Hao Sun, Zhongjiang He, Tianwei Cao, Xuchong Zhang, Hongbin Sun, Kongming Liang, Zhanyu Ma arxiv

Despite its crucial role in video object removal (VOR), existing evaluation paradigms face two critical limitations: questionable references and a misalignment between tradi- tional metrics and human preference. To address these challenges, we introduce VOR- Bench, which advances VOR evaluation through three integrated components. First, we present the VOR Dataset (VORD), the first benchmark dataset providing both paired edited videos and graffiti masks. Its unique strength lies in a diverse data spectrum, which encompasses model-generated, tool-rendered, and camera-captured data, ensuring robust assessment across real-world scenarios. Second, we develop rMPAF, a realistic Motion- capable Paired-video Acquisition Framework. By combining the strengths of image- based object removal and fine-tuned video generation models, rMPAF automatically generates realistic, motion-coherent paired videos. Finally, we propose three evaluation dimensions and introduce VOR-MDSM, the first perception-driven VLM-based scoring model specifically designed for mask-guided VOR. It bridges the gap between arithmetic metrics and human perception by covering the essential visual attributes and matching nuanced human judgment. Extensive experiments demonstrate that VOR-Bench yields evaluation results that align closely with human perception, achieving a remarkable cor- relation (\r{ho} > 0.9) with subjective assessments. We will release VOR-Bench along with its documentation to ensure full reproducibility.

📄 PDF Abstract BibTeX arXiv:2609.16878

Code (2)

iszhanjiawei/video-to-audio-arxiv-daily
liutaocode/Video-Generation-arxiv-daily ★ 20

Tasks

Video Generation

Similar Papers 제목 키워드 기반

VMBench: A Benchmark for Perception-Aligned Video Motion Generation

2025-03-13 · Xinrang Ling, Chen Zhu, Meiqi Wu, Hangyu Li 외

Video generation has advanced rapidly, improving evaluation methods, yet assessing video's motion remains a major challenge. Specifically, there are two key issues: 1) current motion metrics do not fully align with human…

Motion GenerationVideo Generation

VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment

2024-08-21 · Shangkun Sun, Xiaoyu Liang, Songlin Fan, Wenxu Gao 외

Text-driven video editing has recently experienced rapid development. Despite this, evaluating edited videos remains a considerable challenge. Current metrics tend to fail to align with human perceptions, and effective q…

Video AlignmentVideo EditingVideo Quality AssessmentVisual Question Answering (VQA)

HVD: Human Vision-Driven Video Representation Learning for Text-Video Retrieval

2026-01-22 · Zequn Xie, Xin Liu, Boyun Zhang, Yuxiao Lin 외 arxiv

The success of CLIP has driven substantial progress in text-video retrieval. However, current methods often suffer from "blind" feature interaction, where the model struggles to discern key visual information from backgr…

Representation LearningVideo Retrieval

Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars

2026-02-02 · Youliang Zhang, Zhengguang Zhou, Zhentao Yu, Ziyao Huang 외 arxiv

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (G…

Video Generation

PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning

2026-03-27 · Shaoxuan Li, Zhixuan Zhao, Hanze Deng, Zirun Ma 외 arxiv

We introduce PerceptionComp, a manually annotated benchmark for complex, long-horizon, perception-centric video reasoning. PerceptionComp is designed so that no single moment is sufficient: answering each question requir…

Spatial Reasoning