V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning
Multimodal large language models (MLLMs) have achieved remarkable success in general perception, yet complex multi-step visual reasoning remains a persistent challenge. Although recent agentic approaches incorporate tool use, they often neglect critical execution feedback. Consequently, they suffer from the imagination-action-observer (IAO) bias, a misalignment between prior imagination and observer feedback that undermines reasoning stability and optimality. To bridge this gap, we introduce V-ABS, an action-observer driven beam search framework that enables deliberate reasoning through thinker-actor-observer iterations. We also propose an entropy-based adaptive weighting algorithm to mitigate the IAO bias by dynamically balancing the confidence scores between the policy priors and the observational feedback. Moreover, we construct a large-scale supervised fine-tuning (SFT) dataset comprising over 80k samples to guide the model to assign higher prior confidence to correct action paths. Extensive experiments across eight diverse benchmarks show that V-ABS achieves state-of-the-art performance, delivering an average improvement of 19.7% on the Qwen3-VL-8B baseline and consistent gains across both open-source and proprietary models.
Code (0)
등록된 구현이 없습니다.
Tasks
Visual ReasoningSimilar Papers 제목 키워드 기반
Data-Driven and Stealthy Deactivation of Safety Filters
Safety filters ensure that control actions that are executed are always safe, no matter the controller in question. Previous work has proposed a simple and stealthy false-data injection attack for deactivating such safet…
Data-Driven Nonlinear State Observation using Video Measurements
State observation is necessary for feedback control but often challenging for nonlinear systems. While Kazantzis-Kravaris/Luenberger (KKL) observer gives a generic design, its model-based numerical solution is difficult.…
Cognitive-Driven Optimization of Sparse Array Transceiver for MIMO Radar Beamforming
Cognitive multiple-input multiple-output (MIMO) radar is capable of adjusting system parameters adaptively by sensing and learning in complex dynamic environment. Beamforming performance of MIMO radar is guided by both b…
BEAM: Brainwave Empathy Assessment Model for Early Childhood
Empathy in young children is crucial for their social and emotional development, yet predicting it remains challenging. Traditional methods often only rely on self-reports or observer-based labeling, which are susceptibl…
Contrastive LearningInfinite-dimensional observers for high order boundary-controlled port-Hamiltonian systems
This letter investigates the design of a class of infinite-dimensional observers for one dimensional (1D) boundary controlled port-Hamiltonian systems (BC-PHS) defined by differential operators of order $N \geq 1$. The c…