paper-with-me

Papers

CARES: Context-Aware Resolution Selector for VLMs

2025-10-22 · Moshe Kimhi, Nimrod Shabtay, Raja Giryes, Chaim Baskin, Eli Schwartz arxiv

Large vision-language models (VLMs) commonly process images at native or high resolution to remain effective across tasks. This inflates visual tokens ofter to 97-99% of total tokens, resulting in high compute and latency, even when low-resolution images would suffice. We introduce \emph{CARES}-a \textbf{C}ontext-\textbf{A}ware \textbf{R}esolution \textbf{S}elector, a lightweight preprocessing module that, given an image-query pair, predicts the \emph{minimal} sufficient input resolution. CARES uses a compact VLM (350M) to extract features and predict when a target pretrained VLM's response converges to its peak ability to answer correctly. Though trained as a discrete classifier over a set of optional resolutions, CARES interpolates continuous resolutions at inference for fine-grained control. Across five multimodal benchmarks spanning documents and natural images, as well as diverse target VLMs, CARES preserves task performance while reducing compute by up to 80%.

📄 PDF Abstract BibTeX arXiv:2510.19496

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models

2024-06-10 · Peng Xia, Ze Chen, Juanxi Tian, Yangrui Gong 외

Artificial intelligence has significantly impacted medical applications, particularly with the advent of Medical Large Vision Language Models (Med-LVLMs), sparking optimism for the future of automated and personalized he…

Fairness

Context-aware Session-based Recommendation with Graph Neural Networks

2023-10-14 · Zhihui Zhang, Jianxiang Yu, Xiang Li

Session-based recommendation (SBR) is a task that aims to predict items based on anonymous sequences of user behaviors in a session. While there are methods that leverage rich context information in sessions for SBR, mos…

Session-Based Recommendations

StackTok: Accelerating VLMs Inference with Budget-Adaptive Visual Token Selection

2026-09-15 · Zhenbin Wang, Lei Zhang, Lituan Wang, Wei Huang 외 arxiv

Increasing image resolution produces ever-longer visual-token sequences in vision-language models (VLMs), substantially raising their inference cost. To reduce this overhead without retraining, existing methods select co…

CABM: Content-Aware Bit Mapping for Single Image Super-Resolution Network with Large Input

2023-04-13 · CVPR 2023 1 · Senmao Tian, Ming Lu, Jiaming Liu, Yandong Guo 외

With the development of high-definition display devices, the practical scenario of Super-Resolution (SR) usually needs to super-resolve large input like 2K to higher resolution (4K/8K). To reduce the computational and me…

2k4k8kImage Super-Resolution+2

The CARESSES EU-Japan project: making assistive robots culturally competent

2017-08-21 · Barbara Bruno, Nak Young Chong, Hiroko Kamide, Sanjeev Kanoria 외

The nursing literature shows that cultural competence is an important requirement for effective healthcare. We claim that personal assistive robots should likewise be culturally competent, that is, they should be aware o…