paper-with-me

Papers

Zoomer: Adaptive Image Focus Optimization for Black-box MLLM

2025-04-30 · Jiaxu Qian, Chendong Wang, Yifan Yang, Chaoyun Zhang, Huiqiang Jiang, Xufang Luo, Yu Kang, QIngwei Lin, Anlan Zhang, Shiqi Jiang, Ting Cao, Tianjun Mao, Suman Banerjee, Guyue Liu, Saravan Rajmohan, Dongmei Zhang, Yuqing Yang, Qi Zhang, Lili Qiu

Recent advancements in multimodal large language models (MLLMs) have broadened the scope of vision-language tasks, excelling in applications like image captioning and interactive question-answering. However, these models struggle with accurately processing visual data, particularly in tasks requiring precise object recognition and fine visual details. Stringent token limits often result in the omission of critical information, hampering performance. To address these limitations, we introduce \SysName, a novel visual prompting mechanism designed to enhance MLLM performance while preserving essential visual details within token limits. \SysName features three key innovations: a prompt-aware strategy that dynamically highlights relevant image regions, a spatial-preserving orchestration schema that maintains object integrity, and a budget-aware prompting method that balances global context with crucial visual details. Comprehensive evaluations across multiple datasets demonstrate that \SysName consistently outperforms baseline methods, achieving up to a $26.9\%$ improvement in accuracy while significantly reducing token consumption.

📄 PDF Abstract BibTeX arXiv:2505.00742

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningObject RecognitionQuestion AnsweringVisual Prompting

Similar Papers 제목 키워드 기반

ZOOMER: Boosting Retrieval on Web-scale Graphs by Regions of Interest

2022-03-20 · Yuezihan Jiang, Yu Cheng, Hanyu Zhao, Wentao Zhang 외

We introduce ZOOMER, a system deployed at Taobao, the largest e-commerce platform in China, for training and serving GNN-based recommendations over web-scale graphs. ZOOMER is designed for tackling two challenges present…

Retrieval

Visually Grounded Follow-up Questions: a Dataset of Spatial Questions Which Require Dialogue History

2021-08-01 · ACL (splurobonlp) 2021 8 · Tianai Dong, Alberto Testoni, Luciana Benotti, Raffaella Bernardi

In this paper, we define and evaluate a methodology for extracting history-dependent spatial questions from visual dialogues. We say that a question is history-dependent if it requires (parts of) its dialogue history to …

VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning

2025-12-26 · Yang Ding, Yizhen Zhang, Xin Lai, Ruihang Chu 외 arxiv

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language tasks yet remain limited in long video understanding due to the limited context window. Consequently, prevailing approaches te…

Reinforcement Learning

UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding

2026-04-15 · Fei Tang, Bofan Chen, Zhengxi Lu, Tongbo Chen 외 arxiv

GUI grounding, which localizes interface elements from screenshots given natural language queries, remains challenging for small icons and dense layouts. Test-time zoom-in methods improve localization by cropping and re-…

Natural Language Queries

OmniZoomer: Learning to Move and Zoom in on Sphere at High-Resolution

2023-08-16 · ICCV 2023 1 · Zidong Cao, Hao Ai, Yan-Pei Cao, Ying Shan 외

Omnidirectional images (ODIs) have become increasingly popular, as their large field-of-view (FoV) can offer viewers the chance to freely choose the view directions in immersive environments such as virtual reality. The …