paper-with-me

Papers

Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences

2025-10-27 · Zhuoran Jin, Hongbang Yuan, Kejian Zhu, Jiachun Li, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao arxiv

Reward models (RMs) play a critical role in aligning AI behaviors with human preferences, yet they face two fundamental challenges: (1) Modality Imbalance, where most RMs are mainly focused on text and image modalities, offering limited support for video, audio, and other modalities; and (2) Preference Rigidity, where training on fixed binary preference pairs fails to capture the complexity and diversity of personalized preferences. To address the above challenges, we propose Omni-Reward, a step toward generalist omni-modal reward modeling with support for free-form preferences, consisting of: (1) Evaluation: We introduce Omni-RewardBench, the first omni-modal RM benchmark with free-form preferences, covering nine tasks across five modalities including text, image, video, audio, and 3D; (2) Data: We construct Omni-RewardData, a multimodal preference dataset comprising 248K general preference pairs and 69K instruction-tuning pairs for training generalist omni-modal RMs; (3) Model: We propose Omni-RewardModel, which includes both discriminative and generative RMs, and achieves strong performance on Omni-RewardBench as well as other widely used reward modeling benchmarks.

📄 PDF Abstract BibTeX arXiv:2510.23451

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration

2026-05-27 · Xinchen Zhang, Bowei Liu, Jiale Liu, Chufan Shi 외 arxiv

Visual outcomes are increasingly central to multimodal large language models, making reliable and fine-grained verification essential for scaling generalist foundation models. In this work, we investigate multimodal meta…

Reinforcement Learning

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning

2026-06-24 · Zhiyuan Han, Beier Zhu, Wenwen Tong, Pengyang Shao 외 arxiv

We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their reasoning trajectories and (ii) exhibit unfaithful behavior, often hallucinating …

Reinforcement Learning

Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis

2026-01-31 · Zicheng Kong, Dehua Ma, Zhenbo Xu, Alven Yang 외 arxiv

Multimodal large language models (MLLMs) struggle with alignment due to the limitations of existing reward models (RMs), which are predominantly vision-centric, dependent on costly human labels, and provide opaque scalar…

Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method

2025-05-20 · Xinshen Zhang, Zhen Ye, Xu Zheng

Omnidirectional images (ODIs), with their 360{\deg} field of view, provide unparalleled spatial awareness for immersive applications like augmented reality and embodied AI. However, the capability of existing multi-modal…

HallucinationObject LocalizationQuestion AnsweringVisual Question Answering

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization

2026-05-11 · Yeongtak Oh, Dongwook Lee, Sangkwon Park, Heeseung Kim 외 arxiv

While multimodal large language models have advanced across text, image, and audio, personalization research has remained primarily vision-language, with unified omnimodal benchmarking that jointly covers text, image, an…

Visual Grounding