paper-with-me

Papers

Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis

2026-01-31 · Zicheng Kong, Dehua Ma, Zhenbo Xu, Alven Yang, Yiwei Ru, Haoran Wang, Zixuan Zhou, Fuqing Bie, Liuyu Xiang, Huijia Wu, Jian Zhao, Zhaofeng He arxiv

Multimodal large language models (MLLMs) struggle with alignment due to the limitations of existing reward models (RMs), which are predominantly vision-centric, dependent on costly human labels, and provide opaque scalar scores that fail to capture nuanced reasoning, leading to brittle alignment. We present Omni-RRM, an \textbf{Omni}-modal \textbf{R}ubric-grounded \textbf{R}eward \textbf{M}odel that generates multi-dimensional reward signals across text, image, video, and audio. To overcome the high cost and inherent inconsistency of human-centric evaluation in multi-dimensional reasoning, we introduce \textbf{Omni-Preference}, a high-quality dataset constructed via automatic rubric-grounded preference synthesis. In this pipeline, teacher models reconcile raw preferences into explicit justifications, ensuring that the synthesized supervision is both high-fidelity and interpretable. Omni-RRM is trained using a progressive SFT + GRPO regimen, specifically optimized to sharpen reward discrimination on low-margin, hard preference pairs. It achieves state-of-the-art accuracy on video (80.2\% on ShareGPT-Video) and audio benchmarks (66.8\% on Audio-HH-RLHF and 65.0\% on TA2T), yielding a five-benchmark Overall accuracy of 70.4\% and a +17.0\% relative gain over its backbone. Furthermore, Omni-RRM effectively guides Best-of-$N$ selection and exhibits robust transfer to text-only alignment. All resources, including the dataset, training and inference code, and model checkpoints are available at https://tmfk418.github.io/Omni-RRM.

📄 PDF Abstract BibTeX arXiv:2602.00846

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniQuality-R: Advancing Reward Models Through All-Encompassing Quality Assessment

2025-10-12 · Yiting Lu, Fengbin Guan, Yixin Gao, Yan Zhong 외 arxiv

Current visual evaluation approaches are typically constrained to a single task. To address this, we propose OmniQuality-R, a unified reward modeling framework that transforms multi-task quality reasoning into continuous…

Reinforcement Learning

Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences

2025-10-27 · Zhuoran Jin, Hongbang Yuan, Kejian Zhu, Jiachun Li 외 arxiv

Reward models (RMs) play a critical role in aligning AI behaviors with human preferences, yet they face two fundamental challenges: (1) Modality Imbalance, where most RMs are mainly focused on text and image modalities, …

OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling

2025-09-15 · Yang Zhou, Yifan Wang, Jianjun Zhou, Wenzheng Chang 외 arxiv

The field of 4D world modeling - aiming to jointly capture spatial geometry and temporal dynamics - has witnessed remarkable progress in recent years, driven by advances in large-scale generative models and multimodal le…

Video Generation

OmniNWM: Omniscient Driving Navigation World Models

2025-10-21 · Bohan Li, Zhuang Ma, Dalong Du, Baorui Peng 외 arxiv

Autonomous driving world models are expected to work effectively across three core dimensions: state, action, and reward. However, existing methods are typically restricted to fragmented modality modeling, short-horizon …

Autonomous Driving

M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive Performance

2025-02-26 · Qingpei Guo, Kaiyou Song, Zipeng Feng, Ziping Ma 외

We present M2-omni, a cutting-edge, open-source omni-MLLM that achieves competitive performance to GPT-4o. M2-omni employs a unified multimodal sequence modeling framework, which empowers Large Language Models(LLMs) to a…