paper-with-me

홈 › Papers

SweetTokenizer: Semantic-Aware Spatial-Temporal Tokenizer for Compact Visual Discretization

2024-12-11 · Zhentao Tan, Ben Xue, Jian Jia, Junhao Wang, Wencai Ye, Shaoyun Shi, MingJie Sun, Wenjin Wu, Quan Chen, Peng Jiang

This paper presents the \textbf{S}emantic-a\textbf{W}ar\textbf{E} spatial-t\textbf{E}mporal \textbf{T}okenizer (SweetTokenizer), a compact yet effective discretization approach for vision data. Our goal is to boost tokenizers' compression ratio while maintaining reconstruction fidelity in the VQ-VAE paradigm. Firstly, to obtain compact latent representations, we decouple images or videos into spatial-temporal dimensions, translating visual information into learnable querying spatial and temporal tokens through a \textbf{C}ross-attention \textbf{Q}uery \textbf{A}uto\textbf{E}ncoder (CQAE). Secondly, to complement visual information during compression, we quantize these tokens via a specialized codebook derived from off-the-shelf LLM embeddings to leverage the rich semantics from language modality. Finally, to enhance training stability and convergence, we also introduce a curriculum learning strategy, which proves critical for effective discrete visual representation learning. SweetTokenizer achieves comparable video reconstruction fidelity with only \textbf{25\%} of the tokens used in previous state-of-the-art video tokenizers, and boost video generation results by \textbf{32.9\%} w.r.t gFVD. When using the same token number, we significantly improves video and image reconstruction results by \textbf{57.1\%} w.r.t rFVD on UCF-101 and \textbf{37.2\%} w.r.t rFID on ImageNet-1K. Additionally, the compressed tokens are imbued with semantic information, enabling few-shot recognition capabilities powered by LLMs in downstream applications.

📄 PDF Abstract BibTeX arXiv:2412.10443

Code (0)

등록된 구현이 없습니다.

Tasks

Image ReconstructionRepresentation LearningVideo GenerationVideo Reconstruction

Methods 이 논문이 사용한 방법론

VQ-VAE VQ-VAE is a type of variational autoencoder that uses vector quantisation to obtain a discrete latent representation. It differs from…

Similar Papers 제목 키워드 기반

How Can Large Language Models Understand Spatial-Temporal Data?

2024-01-25 · Lei Liu, Shuo Yu, Runze Wang, Zhenxun Ma 외

While Large Language Models (LLMs) dominate tasks like natural language processing and computer vision, harnessing their power for spatial-temporal forecasting remains challenging. The disparity between sequential text a…

Natural Language Understanding

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

2026-06-11 · Guozhen Zhang, Xuerui Qiu, Yutao Cui, Tianhui Song 외 arxiv

Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space. In this paper, we present HYDRA-X, the first UMM that unifies image an…

OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

2024-06-13 · Junke Wang, Yi Jiang, Zehuan Yuan, Binyue Peng 외

Tokenizer, serving as a translator to map the intricate visual data into a compact latent space, lies at the core of visual generative models. Based on the finding that existing tokenizers are tailored to image or video …

Video GenerationVideo Prediction

HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language Models

2025-03-14 · Ziqin Zhou, Yifan Yang, Yuqing Yang, Tianyu He 외

Text-to-video generation poses significant challenges due to the inherent complexity of video data, which spans both temporal and spatial dimensions. It introduces additional redundancy, abrupt variations, and a domain g…

Text-to-Video GenerationVideo Generation

DeRA: Decoupled Representation Alignment for Video Tokenization

2025-12-04 · Pengbo Guo, Junke Wang, Zhen Xing, Chengxu Liu 외 arxiv

This paper presents DeRA, a novel 1D video tokenizer that decouples the spatial-temporal representation learning in video tokenization to achieve better training efficiency and performance. Specifically, DeRA maintains a…

Representation LearningVideo Generation