paper-with-me

홈 › Papers

FOUNDv2: Learning Unified User Quantized Tokenizers for User Representation

2025-08-01 · Chuan He, Yang Chen, Bin Dou, Wuliang Huang, Baokun Wang, Yongchao Liu, Xing Fu, Yu Cheng, Chuntao Hong, Weiqiang Wang, Zhongle Xie, Jiajun Zheng, Xin-Wei Yao arxiv

User representation learning serves as a fundamental pillar for personalized services on large-scale web platforms. Despite its importance, conventional continuous embedding methods face significant challenges, including the lack of a unified paradigm for multi-source data integration, prohibitive storage overhead due to low information density, and the lack of multi-scale modeling granularity. To overcome these limitations, we introduce FOUNDv2, a comprehensive user representation scheme centered on the Unified User Quantized Tokenizer U2QT) framework. FOUNDv2 transforms heterogeneous user data into a standardized discrete token space through a robust two-stage architecture. Specifically, the framework first extracts compact feature representations and subsequently employs a multi-view RQ-VAE to discretize them into storage-efficient tokens using shared and source-specific codebooks. To empower these representations with predictive intelligence, we further design multi-scale alignment objectives to capture both fine-grained behavioral dependencies and macro-temporal periodicity. Extensive experiments on various benchmarks demonstrate that FOUNDv2 consistently outperforms task-specific baselines while achieving substantial reductions in storage and computational costs. Finally, the large-scale deployment of FOUNDv2 on Alipay validates its practical scalability and efficiency across diverse industrial scenarios. The main code is available at: https://github.com/chuanhe1999/FOUNDv2.

📄 PDF Abstract BibTeX arXiv:2508.00956

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Rethinking the Objectives of Vector-Quantized Tokenizers for Image Synthesis

2022-12-06 · CVPR 2024 1 · YuChao Gu, Xintao Wang, Yixiao Ge, Ying Shan 외

Vector-Quantized (VQ-based) generative models usually consist of two basic components, i.e., VQ tokenizers and generative transformers. Prior research focuses on improving the reconstruction fidelity of VQ tokenizers but…

Conditional Image GenerationDecoderImage GenerationSemantic Compression

GQ-VAE: A gated quantized VAE for learning variable length tokens

2025-12-26 · Theo Datta, Kayla Huang, Sham Kakade, David Brandfonbrener arxiv

While most frontier models still use deterministic frequency-based tokenization algorithms such as byte-pair encoding (BPE), there has been significant recent work to design learned neural tokenizers. However, these sche…

Learning Graph Quantized Tokenizers

2024-10-17 · Limei Wang, Kaveh Hassani, Si Zhang, Dongqi Fu 외

Transformers serve as the backbone architectures of Foundational Models, where domain-specific tokenizers allow them to adapt to various domains. Graph Transformers (GTs) have recently emerged as leading models in geomet…

Graph LearningQuantizationSelf-Supervised Learning

TokBench: Evaluating Your Visual Tokenizer before Visual Generation

2025-05-23 · Junfeng Wu, Dongliang Luo, Weizhi Zhao, Zhihao Xie 외

In this work, we reveal the limitations of visual tokenizers and VAEs in preserving fine-grained features, and propose a benchmark to evaluate reconstruction performance for two challenging visual contents: text and face…

Face RecognitionFace ReconstructionImage CompressionOptical Character Recognition (OCR)

Sampling from Your Language Model One Byte at a Time

2025-06-17 · Jonathan Hayase, Alisa Liu, Noah A. Smith, Sewoong Oh

Tokenization is used almost universally by modern language models, enabling efficient text representation using multi-byte or multi-character tokens. However, prior work has shown that tokenization can introduce distorti…

Code GenerationLanguage ModelingLanguage Modelling