paper-with-me

홈 › Papers

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning

2026-05-10 · Xuan Gong, Hanbo Huang, Hao Zheng, Yiran Zhang, Wenbin Dai, Weishu Zhao, Shiyu Liang arxiv

Long chain-of-thought (CoT) reasoning improves large vision--language models, but visual information often fades during generation, limiting long-horizon multimodal reasoning. Existing methods either re-inject vision at inference or train policies for stronger grounding, but where to intervene relies on perception heuristics rather than principled gain analysis, and how local visual influence propagates remains implicit. We study this problem from an information-theoretic standpoint and derive a lower bound on the downstream visual gain of a one-step intervention, which suggests two factors: local branching room (token entropy) and downstream visual propagation potential (suffix divergence from a vision-marginalized reference). Guided by this analysis, we propose reflection-anchor policy optimization (RAPO), a GRPO-based policy optimization method that selects high-entropy reflection anchors and optimizes a chain-masked finite-window KL surrogate for downstream visual dependence. Experiments on reasoning-intensive and general-domain benchmarks show that RAPO delivers substantial gains over strong baselines across multiple LVLM backbones. Mechanism analyses further indicate that reflection anchors are enriched for visually sensitive decision points and that RAPO increases contrastive visual-dependence signals along generated trajectories.

📄 PDF Abstract BibTeX arXiv:2605.09614

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

Stake the Points: Structure-Faithful Instance Unlearning

2026-03-13 · Kiseong Hong, JungKyoo Shin, Eunwoo Kim arxiv

Machine unlearning (MU) addresses privacy risks in pretrained models. The main goal of MU is to remove the influence of designated data while preserving the utility of retained knowledge. Achieving this goal requires pre…

Image ClassificationFace Recognition

Soft Multipath Information-Based UWB Tracking in Cluttered Scenarios: Preliminaries and Validations

2024-05-28 · Chenglong Li, Zukun Lu, Long Huang, Shaojie Ni 외

In this paper, we investigate ultra-wideband (UWB) localization and tracking in cluttered environments. Instead of mitigating the multipath, we exploit the specular reflections to enhance the localizability and improve t…

RbFT-Net: Rectify-Before-Fuse Temporal Radar Anchors for 4D Radar-Camera Depth Completion

2026-08-13 · Wentao Zhao, Shouxuan Wu, Yongtao Cen, Tianchen Deng 외 arxiv

Dense metric depth prediction from cameras and millimeter-wave radar offers a cost-effective sensing solution for autonomous systems. However, radar measurements are inherently sparse and susceptible to clutter, multipat…

Depth Completion

Indoor Positioning for Public Safety: Role of UAVs, LEOs, and Propagation-Aware Techniques

2025-03-15 · Gaurav Duggal, Harish K. Dureppagari, Harpreet S. Dhillon, Jeffrey H. Reed 외

Effective indoor positioning is critical for public safety, enabling first responders to locate at-risk individuals accurately during emergency scenarios. However, traditional Global Navigation Satellite Systems (GNSS) o…

Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs

2026-04-14 · Jianhao Chen, Haoyang Chen, Hanjie Zhao, Haozhe Liang 외 arxiv

Vision-Language Models (VLMs) expand the attack surface of safety-aligned systems by coupling visual perception with text generation. Existing multimodal jailbreak attacks primarily rely on crafted visual content, advers…

Adversarial Attack