Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice
Accurate and effective discrete image tokenization is crucial for long image sequence processing. However, current methods rigidly compress all content at a fixed rate, ignoring the variable information density of images and leading to either redundancy or information loss. Inspired by information entropy, we propose TaTok, a Theoretically grounded adaptive image Tokenization framework. We rigorously identify two key drawbacks in existing methods: information insufficiency when reconstructing images with patch tokens alone, and information redundancy among patch tokens. To address these, we introduce global tokens that model mutual information across patch tokens, and a Dynamic Token Filtering (DTF) algorithm based on cumulative conditional entropy to eliminate redundancy. Experiments confirm TaTok's state-of-the-art performance, delivering a 1.3x gFID improvement and 8.7x inference speedup. By allocating tokens according to information richness, TaTok enables more compressed yet accurate image tokenization, offering valuable insights for future research.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Nighttime Hazy Image Enhancement via Progressively and Mutually Reinforcing Night-Haze Priors
Enhancing the visibility of nighttime hazy images is challenging due to the complex degradation distributions. Existing methods mainly address a single type of degradation (e.g., haze or low-light) at a time, ignoring th…
Image EnhancementImage RestorationTridentSE: Guiding Speech Enhancement with 32 Global Tokens
In this paper, we present TridentSE, a novel architecture for speech enhancement, which is capable of efficiently capturing both global information and local details. TridentSE maintains T-F bin level representation to c…
Speech EnhancementOther Tokens Matter: Exploring Global and Local Features of Vision Transformers for Object Re-Identification
Object Re-Identification (Re-ID) aims to identify and retrieve specific objects from images captured at different places and times. Recently, object Re-ID has achieved great success with the advances of Vision Transforme…
ObjectPPformer: Using pixel-wise and patch-wise cross-attention for low-light image enhancement
Recently, transformer-based methods have shown strong competition compared to CNN-based methods on the low-light image enhancement task, by employing the self-attention for feature extraction. Transformer-based methods p…
Image EnhancementLow-Light Image EnhancementABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching
Contrastive Language-Image Pretraining (CLIP) has achieved remarkable performance in various multimodal tasks. However, it still struggles with compositional image-text matching, particularly in accurately associating ob…
Image-text matching