paper-with-me

Papers

VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment

2024-08-21 · Shangkun Sun, Xiaoyu Liang, Songlin Fan, Wenxu Gao, Wei Gao

Text-driven video editing has recently experienced rapid development. Despite this, evaluating edited videos remains a considerable challenge. Current metrics tend to fail to align with human perceptions, and effective quantitative metrics for video editing are still notably absent. To address this, we introduce VE-Bench, a benchmark suite tailored to the assessment of text-driven video editing. This suite includes VE-Bench DB, a video quality assessment (VQA) database for video editing. VE-Bench DB encompasses a diverse set of source videos featuring various motions and subjects, along with multiple distinct editing prompts, editing results from 8 different models, and the corresponding Mean Opinion Scores (MOS) from 24 human annotators. Based on VE-Bench DB, we further propose VE-Bench QA, a quantitative human-aligned measurement for the text-driven video editing task. In addition to the aesthetic, distortion, and other visual quality indicators that traditional VQA methods emphasize, VE-Bench QA focuses on the text-video alignment and the relevance modeling between source and edited videos. It proposes a new assessment network for video editing that attains superior performance in alignment with human preferences. To the best of our knowledge, VE-Bench introduces the first quality assessment dataset for video editing and an effective subjective-aligned quantitative metric for this domain. All data and code will be publicly available at https://github.com/littlespray/VE-Bench.

📄 PDF Abstract BibTeX arXiv:2408.11481

Code (2)

littlespray/e-bench 공식 구현 pytorch
littlespray/ve-bench 공식 구현 pytorch

Tasks

Video AlignmentVideo EditingVideo Quality AssessmentVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Towards Scalable Human-aligned Benchmark for Text-guided Image Editing

2025-05-01 · CVPR 2025 1 · Suho Ryu, Kihyun Kim, Eugene Baek, Dongsoo Shin 외

A variety of text-guided image editing models have been proposed recently. However, there is no widely-accepted standard evaluation method mainly due to the subjective nature of the task, letting researchers rely on manu…

text-guided-image-editing

IE-Bench: Advancing the Measurement of Text-Driven Image Editing for Human Perception Alignment

2025-01-17 · Shangkun Sun, Bowen Qu, Xiaoyu Liang, Songlin Fan 외

Recent advances in text-driven image editing have been significant, yet the task of accurately evaluating these edited images continues to pose a considerable challenge. Different from the assessment of text-driven image…

Image Generation

SubData: Bridging Heterogeneous Datasets to Enable Theory-Driven Evaluation of Political and Demographic Perspectives in LLMs

2024-12-21 · Leon Fröhling, Pietro Bernardelle, Gianluca Demartini

As increasingly capable large language models (LLMs) emerge, researchers have begun exploring their potential for subjective tasks. While recent work demonstrates that LLMs can be aligned with diverse human perspectives,…

Hate Speech Detection

IE-Critic-R1: Advancing the Explanatory Measurement of Text-Driven Image Editing for Human Perception Alignment

2025-11-22 · Bowen Qu, Shangkun Sun, Xiaoyu Liang, Wei Gao arxiv

Recent advances in text-driven image editing have been significant, yet the task of accurately evaluating these edited images continues to pose a considerable challenge. Different from the assessment of text-driven image…

Reinforcement LearningImage GenerationImage Editing

HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding

2026-05-19 · Mengqi Shi, Haopeng Zhang arxiv

While Multimodal Large Language Models (MLLMs) exhibit strong performance on standard video tasks, their ability to faithfully summarize and reason over complex narratives remains poorly evaluated. Existing summarization…