paper-with-me

홈 › Papers

TaskTok: Delving into Task Tokens for Task-driven Image Restoration

2026-06-25 · Hongjae Lee, Sojung Kang, Jaeseong Yu, Seung-Won Jung arxiv

While traditional image restoration focuses on perceptual quality, Task-Driven Image Restoration (TDIR) aims to maximize the performance of downstream high-level vision tasks. Recent approaches leveraging generative priors have shown promise for TDIR; however, they typically suffer from computational inefficiency and potential semantic alteration by indiscriminately updating all latent tokens. In this paper, we posit that not all visual information is equally important for machine perception. Through an analysis of the latent token space, we observe that task-relevant cues are unevenly distributed across the token sequence, exhibiting index-wise specialization. This suggests that selectively refining a subset of tokens can be sufficient for task-driven objectives. Leveraging this insight, we propose TaskTok, a novel framework that selectively restores only task-relevant tokens via a learnable token switch and a lightweight token refinement module. Extensive experiments across image classification, semantic segmentation, and object detection demonstrate that TaskTok significantly enhances task performance with high computational efficiency. The source code is available at https://github.com/jimmy9704/TaskTok

📄 PDF Abstract BibTeX arXiv:2606.26615

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencySemantic SegmentationImage ClassificationImage Restoration

Similar Papers 제목 키워드 기반

Learning Spectral-Decomposed Tokens for Domain Generalized Semantic Segmentation

2024-07-26 · Jingjun Yi, Qi Bi, Hao Zheng, Haolan Zhan 외

The rapid development of Vision Foundation Model (VFM) brings inherent out-domain generalization for a variety of down-stream tasks. Among them, domain generalized semantic segmentation (DGSS) holds unique challenges as …

Domain GeneralizationSemantic Segmentation

Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding

2024-07-19 · Renshan Zhang, Yibo Lyu, Rui Shao, Gongwei Chen 외

Cropping high-resolution document images into multiple sub-images is the most widely used approach for current Multimodal Large Language Models (MLLMs) to do document understanding. Most of current document understanding…

document understandingInformativeness

PR-MIM: Delving Deeper into Partial Reconstruction in Masked Image Modeling

2024-11-24 · Zhong-Yu Li, Yunheng Li, Deng-Ping Fan, Ming-Ming Cheng

Masked image modeling has achieved great success in learning representations but is limited by the huge computational costs. One cost-saving strategy makes the decoder reconstruct only a subset of masked tokens and throw…

Decoder

Hubness and Pollution: Delving into Cross-Space Mapping for Zero-Shot Learning

2015-07-01 · IJCNLP 2015 7 · Angeliki Lazaridou, Georgiana Dinu, Marco Baroni
Semantic Textual SimilarityZero-Shot Learning

Axiom Pinpointing

2020-03-18 · Rafael Peñaloza

Axiom pinpointing refers to the task of finding the specific axioms in an ontology which are responsible for a consequence to follow. This task has been studied, under different names, in many research areas, leading to …