paper-with-me

홈 › Papers

Beyond Dominant Patches: Spatial Credit Redistribution For Grounded Vision-Language Models

2026-02-25 · Niamul Hassan Samin, Md Arifur Rahman, Abdullah Ibne Hanif Arean, Juena Ahmed Noshin, Md Ashikur Rahman arxiv

Vision-Language Models (VLMs) often hallucinate objects that are not present in the input image. We identify a contributing cause of this behavior, which we term spatial credit collapse: in early transformer layers, hidden-state activation concentrates on a small number of visual patches, suppressing surrounding contextual evidence and increasing reliance on language priors. Across seven models we observe a strong correlation between visual attention entropy and hallucination rate (r = -0.65, p < 0.001), suggesting that reduced spatial credit diversity contributes to hallucination. To address this issue we propose Spatial Credit Redistribution (SCR), a training-free inference-time method. SCR uses a lightweight two-pass procedure. A diagnostic pass identifies the top-K high-attention source patches and their spatial neighbors. A redistribution pass then scales each source by 1/lambda (~0.91) and injects a (lambda - 1) weighted copy of its hidden state into neighboring patches, restoring suppressed visual context without modifying model weights. Because the diagnostic pass is performed once per image and reused across the output sequence, the added latency is negligible (<0.5 ms per token for 100-token responses). We evaluate SCR across seven model configurations from four VLM families (Chameleon, LLaVA-1.5, Qwen-VL/Qwen2-VL, and InternVL2) on five benchmarks: POPE, CHAIR, MME, HallusionBench, and AMBER. SCR reduces POPE-Adversarial hallucination by 4.6-6.0 percentage points and CHAIR-s by 41-51 percent while preserving caption quality (CIDEr drop <=0.8). Compared with prior inference-time methods including OPERA, VCD, OA-VCD, DoLa, VLI, SID, and CRoPS, SCR achieves a better trade-off between hallucination reduction, generation quality, and latency.

📄 PDF Abstract BibTeX arXiv:2602.22469

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Global redistribution and local migration in semi-discrete host-parasitoid population dynamic models

2019-12-30

Host-parasitoid population dynamics is often probed using a semi-discrete/hybrid modeling framework. Here, the update functions in the discrete-time model connecting year-to-year changes in the population densities are o…

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO

2026-08-04 · Zhe Cao, Miaowen Wen, Fangjiong Chen arxiv

Reinforcement learning with verifiable rewards (RLVR) com- monly optimizes each correct completion as an independent learning signal. In GRPO, this completion-level uniformity creates structure-level skew: recurring corr…

Reinforcement Learning

RREDCoT: Segment-Level Reward Redistribution for Reasoning Models

2026-06-04 · Mykyta Ielanskyi, Kajetan Schweighofer, Lukas Aichberger, Sepp Hochreiter arxiv

Recent advancements in reasoning language models have been driven by Reinforcement Learning (RL) fine-tuning. Most often, these rely on the Group Relative Policy Optimization (GRPO) algorithm or modifications thereof to …

Reinforcement Learning

Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning

2024-12-15 · Yun Qu, Yuhang Jiang, Boyuan Wang, Yixiu Mao 외

Reinforcement learning (RL) often encounters delayed and sparse feedback in real-world applications, even with only episodic rewards. Previous approaches have made some progress in reward redistribution for credit assign…

Decision MakingLarge Language Modelreinforcement-learningReinforcement Learning+1

$TAR^2$: Temporal-Agent Reward Redistribution for Optimal Policy Preservation in Multi-Agent Reinforcement Learning

2025-02-07 · Aditya Kapoor, Kale-ab Tessera, Mayank Baranwal, Harshad Khadilkar 외

In cooperative multi-agent reinforcement learning (MARL), learning effective policies is challenging when global rewards are sparse and delayed. This difficulty arises from the need to assign credit across both agents an…

Multi-agent Reinforcement LearningTAR