paper-with-me

홈 › Papers

Viper-F1: Fast and Fine-Grained Multimodal Understanding with Cross-Modal State-Space Modulation

2025-11-14 · Quoc-Huy Trinh arxiv

Recent advances in multimodal large language models (MLLMs) have enabled impressive progress in vision-language understanding, yet their high computational cost limits deployment in resource-constrained scenarios such as robotic manipulation, personal assistants, and smart cameras. Most existing methods rely on Transformer-based cross-attention, whose quadratic complexity hinders efficiency. Moreover, small vision-language models often struggle to precisely capture fine-grained, task-relevant visual regions, leading to degraded performance on fine-grained reasoning tasks that limit their effectiveness in the real world. To address these issues, we introduce Viper-F1, a Hybrid State-Space Vision-Language Model that replaces attention with efficient Liquid State-Space Dynamics. To further enhance visual grounding, we propose a Token-Grid Correlation Module, which computes lightweight correlations between text tokens and image patches and modulates the state-space dynamics via FiLM conditioning. This enables the model to selectively emphasize visual regions relevant to the textual prompt while maintaining linear-time inference. Experimental results across multiple benchmarks demonstrate that Viper-F1 achieves accurate, fine-grained understanding with significantly improved efficiency.

📄 PDF Abstract BibTeX arXiv:2511.11177

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model

2025-10-28 · Juntian Zhang, Song Jin, Chuanqi Cheng, Yuhan Liu 외 arxiv

The limited capacity for fine-grained visual perception presents a critical bottleneck for Vision-Language Models (VLMs) in real-world applications. Addressing this is challenging due to the scarcity of high-quality data…

Reinforcement Learning

VIPER: Visual Perception and Explainable Reasoning for Sequential Decision-Making

2025-03-19 · Mohamed Salim Aissi, Clemence Grislain, Mohamed Chetouani, Olivier Sigaud 외

While Large Language Models (LLMs) excel at reasoning on text and Vision-Language Models (VLMs) are highly effective for visual perception, applying those models for visual instruction-based planning remains a widely ope…

Decision MakingSequential Decision Making

TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding

2025-11-20 · Boshen Xu, Zihan Xiao, Jiaze Li, Jianzhong Ju 외 arxiv

We introduce TimeViper, a hybrid vision-language model designed to tackle challenges of long video understanding. Processing long videos demands both an efficient model architecture and an effective mechanism for handlin…

Mimicking the Physicist's Eye:A VLM-centric Approach for Physics Formula Discovery

2025-08-24 · Jiaqi Liu, Songning Lai, Pengze Li, Di Yu 외 arxiv

Automated discovery of physical laws from observational data in the real world is a grand challenge in AI. Current methods, relying on symbolic regression or LLMs, are limited to uni-modal data and overlook the rich, vis…

Reinforcement Learning

Instant Preference Alignment for Text-to-Image Diffusion Models

2025-08-25 · Yang Li, Songlin Yang, Xiaoxuan Han, Wei Wang 외 arxiv

Text-to-image (T2I) generation has greatly enhanced creative expression, yet achieving preference-aligned generation in a real-time and training-free manner remains challenging. Previous methods often rely on static, pre…

Image Generation