paper-with-me

홈 › Papers

Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation

2026-06-29 · Xin Zou, Haolin Deng, Yibo Yan, Shuliang Liu, Zhiwei Jin, Chen Chen, Haonan Lu, Xuming Hu arxiv

Multimodal Large Language Models (MLLMs) are prone to hallucination as their generation preferences are insufficiently calibrated to visual evidence, causing them to fall back on linguistic priors, rather than faithful grounding. In this work, we start from an empirical observation: when query-relevant visual evidence is explicitly strengthened using the model's own attention, generation becomes more accurate, suggesting that many failures do not arise solely from missing perception, but from an insufficient tendency to trust the evidence the model has already attended to. Motivated by this finding, we propose Oriented Pickup Preference Optimization (\texttt{OPPO}), an evidence-aware alignment objective that learns preferences over the strength of visual evidence, rather than only response quality. Concretely, \texttt{OPPO} contrasts the same faithful response under stronger, anchored, weaker-evidence views, turning naive visual preference into ordered visual-evidence alignment. We further combine this objective with fine-grained span-level and token-level regularization to stabilize the training. Besides, we provide a theoretical analysis showing that ordered evidence margins induce a positive lower bound on local visual sensitivity. Extensive evaluations across hallucination and general-purpose benchmarks demonstrate that \texttt{OPPO} consistently outperforms baseline methods.

📄 PDF Abstract BibTeX arXiv:2606.29805

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Standby-Based Deadlock Avoidance Method for Multi-Agent Pickup and Delivery Tasks

2022-01-16 · Tomoki Yamauchi, Yuki Miyashita, Toshiharu Sugawara

The multi-agent pickup and delivery (MAPD) problem, in which multiple agents iteratively carry materials without collisions, has received significant attention. However, many conventional MAPD algorithms assume a specifi…

Motif-Video 2B: Technical Report

2026-04-14 · Junghwan Lim, Wai Ting Cheung, Minsu Ha, Beomgyu Kim 외 arxiv

Training strong video generation models usually requires massive datasets, large parameter counts, and substantial compute. In this work, we ask whether strong text-to-video quality is possible at a much smaller budget: …

Representation LearningVideo Generation

Understanding the Dynamics of Drivers' Locations for Passengers Pickup Performance: A Case Study

2020-09-09 · Punit Rathore, Ali Zonoozi, Omid Geramifard, Tan Kian Lee

With the emergence of e-hailing taxi services, a growing number of scholars have attempted to analyze the taxi trips data to gain insights from drivers' and passengers' flow patterns and understand different dynamics of …

Clustering

Understanding and Visualizing the District of Columbia Capital Bikeshare System Using Data Analysis for Balancing Purposes

2017-08-14 · Kiana Roshan Zamir, Ali Shafahi, Ali Haghani

Bike sharing systems' popularity has consistently been rising during the past years. Managing and maintaining these emerging systems are indispensable parts of these systems. Visualizing the current operations can assist…

Management

Guitar Pickups I: Analysis of the Effect of Winding and Wire Gauge on Single Coil Electric Guitar Pickups

2024-09-29 · Charles Batchelor, Jack Gooding, William Marriott, Nikola Chalashkanov 외

Guitar Pickups have been in production for nearly 100 years, and the question of how exactly one pickup is tonally superior to another is still subject to a high level of debate. This paper is the first in a set demystif…