paper-with-me

Papers

Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model

2026-01-07 · Yuan Wang, Borui Liao, Huijuan Huang, Jinda Lu, Ouxiang Li, Kuien Liu, Meng Wang, Xiang Wang arxiv

Recent advances in video reward models and post-training strategies have improved text-to-video (T2V) generation. While these models typically assess visual quality, motion quality, and text alignment, they often overlook key structural distortions, such as abnormal object appearances and interactions, which can degrade the overall quality of the generative video. To address this gap, we introduce REACT, a frame-level reward model designed specifically for structural distortions evaluation in generative videos. REACT assigns point-wise scores and attribution labels by reasoning over video frames, focusing on recognizing distortions. To support this, we construct a large-scale human preference dataset, annotated based on our proposed taxonomy of structural distortions, and generate additional data using a efficient Chain-of-Thought (CoT) synthesis pipeline. REACT is trained with a two-stage framework: (1) supervised fine-tuning with masked loss for domain knowledge injection, followed by (2) reinforcement learning with Group Relative Policy Optimization (GRPO) and pairwise rewards to enhance reasoning capability and align output scores with human preferences. During inference, a dynamic sampling mechanism is introduced to focus on frames most likely to exhibit distortion. We also present REACT-Bench, a benchmark for generative video distortion evaluation. Experimental results demonstrate that REACT complements existing reward models in assessing structutal distortion, achieving both accurate quantitative evaluations and interpretable attribution analysis.

📄 PDF Abstract BibTeX arXiv:2601.04033

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors

2024-12-12 · Yabo Chen, Chen Yang, Jiemin Fang, Xiaopeng Zhang 외

Single-image 3D reconstruction remains a fundamental challenge in computer vision due to inherent geometric ambiguities and limited viewpoint information. Recent advances in Latent Video Diffusion Models (LVDMs) offer pr…

3D ReconstructionImage to 3DVideo Generation

Generative Latent Video Compression

2025-10-11 · Zongyu Guo, Zhaoyang Jia, Jiahao Li, Xiaoyi Zhang 외 arxiv

Perceptual optimization is widely recognized as essential for neural compression, yet balancing the rate-distortion-perception tradeoff remains challenging. This difficulty is especially pronounced in video compression, …

Deep Iterative Frame Interpolation for Full-frame Video Stabilization

2019-09-05 · Jinsoo Choi, In So Kweon

Video stabilization is a fundamental and important technique for higher quality videos. Prior works have extensively explored video stabilization, but most of them involve cropping of the frame boundaries and introduce m…

Video Stabilization

Enhancing Video Inpainting with Aligned Frame Interval Guidance

2025-10-24 · Ming Xie, Junqiu Yu, Qiaole Dong, Xiangyang Xue 외 arxiv

Recent image-to-video (I2V) based video inpainting methods have made significant strides by leveraging single-image priors and modeling temporal consistency across masked frames. Nevertheless, these methods suffer from s…

Image InpaintingVideo Inpainting

Direct Motion Models for Assessing Generated Videos

2025-04-30 · Kelsey Allen, Carl Doersch, Guangyao Zhou, Mohammed Suhail 외

A current limitation of video generative video models is that they generate plausible looking frames, but poor motion -- an issue that is not well captured by FVD and other popular methods for evaluating generated videos…

Action Recognition