paper-with-me

Papers

ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model

2025-10-28 · Juntian Zhang, Song Jin, Chuanqi Cheng, Yuhan Liu, Yankai Lin, Xun Zhang, Yufei Zhang, Fei Jiang, Guojun Yin, Wei Lin, Rui Yan arxiv

The limited capacity for fine-grained visual perception presents a critical bottleneck for Vision-Language Models (VLMs) in real-world applications. Addressing this is challenging due to the scarcity of high-quality data and the limitations of existing methods: supervised fine-tuning (SFT) often compromises general capabilities, while reinforcement fine-tuning (RFT) prioritizes textual reasoning over visual perception. To bridge this gap, we propose a novel two-stage task that structures visual perception learning as a coarse-to-fine progressive process. Based on this task formulation, we develop ViPER, a self-bootstrapping framework specifically designed to enable iterative evolution through self-critiquing and self-prediction. By synergistically integrating image-level and instance-level reconstruction with a two-stage reinforcement learning strategy, ViPER establishes a closed-loop training paradigm, where internally synthesized data directly fuel the enhancement of perceptual ability. Applied to the Qwen2.5-VL family, ViPER produces the Qwen-Viper series. With an average gain of 1.7% on seven comprehensive benchmarks spanning various tasks and up to 6.0% on fine-grained perception, Qwen-Viper consistently demonstrates superior performance across different vision-language scenarios while maintaining generalizability. Beyond enabling self-improvement in perceptual capabilities, ViPER provides concrete evidence for the reciprocal relationship between generation and understanding, a breakthrough to developing more autonomous and capable VLMs.

📄 PDF Abstract BibTeX arXiv:2510.24285

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

VIPER: Visual Perception and Explainable Reasoning for Sequential Decision-Making

2025-03-19 · Mohamed Salim Aissi, Clemence Grislain, Mohamed Chetouani, Olivier Sigaud 외

While Large Language Models (LLMs) excel at reasoning on text and Vision-Language Models (VLMs) are highly effective for visual perception, applying those models for visual instruction-based planning remains a widely ope…

Decision MakingSequential Decision Making

VIPER Strike: Defeating Visual Reasoning CAPTCHAs via Structured Vision-Language Inference

2026-01-10 · Minfeng Qi, Dongyang He, Qin Wang, Lefeng Zhang arxiv

Visual Reasoning CAPTCHAs (VRCs) combine visual scenes with natural-language queries that demand compositional inference over objects, attributes, and spatial relations. They are increasingly deployed as a primary defens…

Visual Reasoning

Comparing and Combining Approximate Computing Frameworks

2021-02-16 · Saeid Barati, Gordon Kindlmann, Hank Hoffmann

Approximate computing frameworks configure applications so they can operate at a range of points in an accuracy-performance trade-off space. Prior work has introduced many frameworks to create approximate programs. As ap…

Mimicking the Physicist's Eye:A VLM-centric Approach for Physics Formula Discovery

2025-08-24 · Jiaqi Liu, Songning Lai, Pengze Li, Di Yu 외 arxiv

Automated discovery of physical laws from observational data in the real world is a grand challenge in AI. Current methods, relying on symbolic regression or LLMs, are limited to uni-modal data and overlook the rich, vis…

Reinforcement Learning

Exposing Functional Fusion: A New Class of Strategic Backdoor in Dynamic Prompt Architectures

2026-05-19 · Zeyao Liu, Zhendong Zhao, Xiaojun Chen, Xin Zhao 외 arxiv

Existing ViT backdoor attacks based on backbone-overwriting full-tuning are computationally expensive and inflict performance degradation. This has forced adversaries towards the Visual Parameter-Efficient Fine-Tuning (P…

parameter-efficient fine-tuning