paper-with-me

홈 › Papers

Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice

2026-05-11 · Xiusheng Huang, Xin Jiang, Jun Zhao, Kang Liu, Yequan Wang arxiv

Accurate and effective discrete image tokenization is crucial for long image sequence processing. However, current methods rigidly compress all content at a fixed rate, ignoring the variable information density of images and leading to either redundancy or information loss. Inspired by information entropy, we propose TaTok, a Theoretically grounded adaptive image Tokenization framework. We rigorously identify two key drawbacks in existing methods: information insufficiency when reconstructing images with patch tokens alone, and information redundancy among patch tokens. To address these, we introduce global tokens that model mutual information across patch tokens, and a Dynamic Token Filtering (DTF) algorithm based on cumulative conditional entropy to eliminate redundancy. Experiments confirm TaTok's state-of-the-art performance, delivering a 1.3x gFID improvement and 8.7x inference speedup. By allocating tokens according to information richness, TaTok enables more compressed yet accurate image tokenization, offering valuable insights for future research.

📄 PDF Abstract BibTeX arXiv:2605.16384

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Nighttime Hazy Image Enhancement via Progressively and Mutually Reinforcing Night-Haze Priors

2026-01-05 · Chen Zhu, Huiwen Zhang, Mu He, Yujie Li 외 arxiv

Enhancing the visibility of nighttime hazy images is challenging due to the complex degradation distributions. Existing methods mainly address a single type of degradation (e.g., haze or low-light) at a time, ignoring th…

Image EnhancementImage Restoration

TridentSE: Guiding Speech Enhancement with 32 Global Tokens

2022-10-24 · Dacheng Yin, Zhiyuan Zhao, Chuanxin Tang, Zhiwei Xiong 외

In this paper, we present TridentSE, a novel architecture for speech enhancement, which is capable of efficiently capturing both global information and local details. TridentSE maintains T-F bin level representation to c…

Speech Enhancement

Other Tokens Matter: Exploring Global and Local Features of Vision Transformers for Object Re-Identification

2024-04-23 · Yingquan Wang, Pingping Zhang, Dong Wang, Huchuan Lu

Object Re-Identification (Re-ID) aims to identify and retrieve specific objects from images captured at different places and times. Recently, object Re-ID has achieved great success with the advances of Vision Transforme…

Object

PPformer: Using pixel-wise and patch-wise cross-attention for low-light image enhancement

2024-01-15 · Computer Vision and Image Understanding 2024 1 · J Dang, Y Zhong, X Qin

Recently, transformer-based methods have shown strong competition compared to CNN-based methods on the low-light image enhancement task, by employing the self-attention for feature extraction. Transformer-based methods p…

Image EnhancementLow-Light Image Enhancement

ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching

2025-12-19 · Qi Zhang, Yuxu Chen, Lei Deng, Lili Shen arxiv

Contrastive Language-Image Pretraining (CLIP) has achieved remarkable performance in various multimodal tasks. However, it still struggles with compositional image-text matching, particularly in accurately associating ob…

Image-text matching