paper-with-me

홈 › Papers

LookThere! Sparse Vision by Reinforced Selection

2026-09-04 · Sreehari Rammohan, Yousef Yassin, Anthony Fuller, Junfeng Wen, Carl Vondrick, Evan Shelhamer arxiv

Vision transformers typically treat every image token as equally important, yet for most tasks in computer vision only a fraction are needed. Adaptive computation methods accelerate inference by choosing which tokens to process, but existing methods struggle at extreme sparsity and require heuristics that may not generalize like token diversity and attention scores. We address these limitations with LookThere, achieving a new pareto frontier in performance-compute trade-offs through an end-to-end reinforcement learning framework that jointly trains a shallow input selector and a deep representation extractor. The selector learns where to look and the extractor learns what to see, together saving computation by selecting only what is worth processing for a given task without relying on auxiliary signals. We show that LookThere only selects the task-specific input, excelling at sparse recognition in high-resolution settings (traffic signs, billiards), and maintaining accuracy with as little as 0.2% of the input. It generalizes across tasks and models, including global recognition (ImageNet classification), local recognition (ADE20K segmentation), zero-shot classification (by distillation), and regression (counting). Across all settings, LookThere surpasses state-of-the-art selection to provide a general and scalable framework for specialized and efficient adaptive computation.

📄 PDF Abstract BibTeX arXiv:2609.04698

Code (2)

Tavish9/awesome-daily-AI-arxiv ★ 114
arxivsub/arXivSub_daily_arxiv ★ 4

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Learning to Label: A Reinforced Self-Evolving Framework for Semi-supervised Referring Expression Segmentation

2026-05-27 · Runlong Cao, Ying Zang, Chuanwei Zhou, Tianrun Chen 외 arxiv

Semi-supervised referring expression segmentation (SS-RES) aims to achieve precise pixel-level language grounding under limited annotation, yet suffers from limited supervision and unreliable pseudo-labels when exploitin…

Referring Expression Segmentation

ReaSon: Reinforced Causal Search with Information Bottleneck for Video Understanding

2025-11-16 · Yuan Zhou, Litao Hua, Shilong Jin, Wentao Huang 외 arxiv

Keyframe selection has become essential for video understanding with vision-language models (VLMs) due to limited input tokens and the temporal sparsity of relevant information across video frames. Video understanding of…

Reinforcement Learning

Simplifying Reinforced Feature Selection via Restructured Choice Strategy of Single Agent

2020-09-19 · Xiaosa Zhao, Kunpeng Liu, Wei Fan, Lu Jiang 외

Feature selection aims to select a subset of features to optimize the performances of downstream predictive tasks. Recently, multi-agent reinforced feature selection (MARFS) has been introduced to automate feature select…

feature selection

AutoFS: Automated Feature Selection via Diversity-aware Interactive Reinforcement Learning

2020-08-27 · Wei Fan, Kunpeng Liu, Hao liu, Pengyang Wang 외

In this paper, we study the problem of balancing effectiveness and efficiency in automated feature selection. Feature selection is a fundamental intelligence for machine learning and predictive analysis. After exploring …

Diversityfeature selectionNavigatereinforcement-learning+2

Reinforced Self-Attention Network: a Hybrid of Hard and Soft Attention for Sequence Modeling

2018-01-31 · Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang 외

Many natural language processing tasks solely rely on sparse dependencies between a few tokens in a sentence. Soft attention mechanisms show promising performance in modeling local/global dependencies by soft probabiliti…

Hard AttentionNatural Language InferenceSentence