paper-with-me

홈 › Papers

Refining Multidimensional Video Reward Models via Disentangled Influence Functions

2026-05-27 · Muyao Wang, Zeke Xie, Hideki Nakayama arxiv

As Text-to-Video (T2V) generation models continue to evolve, the complexity of video evaluation necessitates a fine-grained assessment across various axes. To address this, recent works have focused on developing Multidimensional Video Reward Models (MVRMs), which decompose the evaluation process to better align with the multifaceted nature of human visual perception. However, training effective MVRMs is fundamentally challenged by the complex nature of video data. In this work, we identify a critical phenomenon termed Dimensional Heterogeneity: the reliability of a training sample can vary substantially across evaluation dimensions, meaning that a sample may provide reliable supervision for one objective while inducing high supervision risk for another. Consequently, prevailing data-centric methods that filter based on global scalar metrics are ill-posed for T2V tasks. To address this, we propose a disentangled influence framework that that efficiently estimates dimension-specific supervision risk. Leveraging this framework, we introduce two dimension-disentangled refinement strategies: Dimension-Disentangled Pruning, which removes extreme high-risk samples, and Dimension-Disentangled Reweighting, which softly down-weights high-risk supervision. Extensive experiments demonstrate that our disentangled strategies significantly outperform global filtering baselines, yielding reward models with superior alignment to ground truth.

📄 PDF Abstract BibTeX arXiv:2605.28203

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Disentangled Multidimensional Metric Learning for Music Similarity

2020-08-09 · Jongpil Lee, Nicholas J. Bryan, Justin Salamon, Zeyu Jin 외

Music similarity search is useful for a variety of creative tasks such as replacing one music recording with another recording with a similar "feel", a common task in video editing. For this task, it is typically necessa…

Metric LearningSpecificityVideo Editing

On the Expressivity of Multidimensional Markov Reward

2023-07-22 · Shuwa Miura

We consider the expressivity of Markov rewards in sequential decision making under uncertainty. We view reward functions in Markov Decision Processes (MDPs) as a means to characterize desired behaviors of agents. Assumin…

Decision MakingDecision Making Under UncertaintySequential Decision Making

VIRAL: Vision-grounded Integration for Reward design And Learning

2025-05-28 · Valentin Cuzin-Rambaud, Emilien Komlenovic, Alexandre Faure, Bruno Yun

The alignment between humans and machines is a critical challenge in artificial intelligence today. Reinforcement learning, which aims to maximize a reward function, is particularly vulnerable to the risks associated wit…

DISENTANGLED STATE SPACE MODELS: UNSUPERVISED LEARNING OF DYNAMICS ACROSS HETEROGENEOUS ENVIRONMENTS

2019-03-27 · ICLR Workshop DeepGenStruct 2019 · Ðorđe Miladinović, Waleed Gondal, Bernhard Schölkopf, Joachim M. Buhmann 외

Sequential data often originates from diverse environments. Across them exist both shared regularities and environment specifics. To learn robust cross-environment descriptions of sequences we introduce disentangled stat…

PredictionState Space Models

Disentangling Influence: Using Disentangled Representations to Audit Model Predictions

2019-06-20 · NeurIPS 2019 12 · Charles T. Marx, Richard Lanas Phillips, Sorelle A. Friedler, Carlos Scheidegger 외

Motivated by the need to audit complex and black box models, there has been extensive research on quantifying how data features influence model predictions. Feature influence can be direct (a direct influence on model ou…