paper-with-me

홈 › Papers

SA-RSQ: A Versatile Sparse Representation Framework for Multi-modal Recommender Systems

2026-08-24 · Xiang Wang, Shigang Quan, Tingzhen Chang, Kang Yang, Sitong Chen, Yabo Fan, Xingxing Wang, Zhaodian He arxiv

Deploying high-dimensional multimodal features in industrial recommender systems incurs substantial storage and latency overhead. Hard quantization is compact but introduces boundary distortion, whereas dense soft quantization couples representation quality to the limited storage budget. We propose Sparse Activation-based Residual Soft Quantization (SA-RSQ), which uses Top-K sparse routing and softmax weights to store compact (Index, Probability) tuples. The stored tuples decouple per-item storage from codebook dimensionality; for a fixed selected support, gradients propagate through the routing weights and weighted reconstruction without relying on a straight-through estimator. Experiments on a proprietary food-delivery advertising dataset show favorable reconstruction-performance and CTR trade-offs across storage budgets of 8-48 bytes per item. A preliminary Next-Distribution Prediction study and a one-week online A/B test further demonstrate the practical potential of SA-RSQ, with relative lifts of +2.51% in CTR and +3.66% in CPM.

📄 PDF Abstract BibTeX arXiv:2608.22979

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Versatile Multi-Modal Pre-Training for Human-Centric Perception

2022-03-25 · CVPR 2022 1 · Fangzhou Hong, Liang Pan, Zhongang Cai, Ziwei Liu

Human-centric perception plays a vital role in vision and graphics. But their data annotations are prohibitively expensive. Therefore, it is desirable to have a versatile pre-train model that serves as a foundation for d…

Contrastive LearningHuman ParsingRepresentation Learning

M3imic: Learning a Versatile Whole-Body Controller for Multimodal Motion Mimicking

2026-06-03 · Zuxing Lu, Ziang Zheng, Yao Lyu, Jingyu Liu 외 arxiv

Building a general-purpose whole-body controller is essential for enabling diverse motion capabilities in humanoid robots across a wide range of downstream tasks, including locomotion and loco-manipulation. Different tas…

Reinforcement Learning

Large Multi-modal Models Can Interpret Features in Large Multi-modal Models

2024-11-22 · Kaichen Zhang, Yifei Shen, Bo Li, Ziwei Liu

Recent advances in Large Multimodal Models (LMMs) lead to significant breakthroughs in both academia and industry. One question that arises is how we, as humans, can understand their internal neural representations. This…

Multimodal sparse representation learning and applications

2015-11-19 · Miriam Cha, Youngjune Gwon, H. T. Kung

Unsupervised methods have proven effective for discriminative tasks in a single-modality scenario. In this paper, we present a multimodal framework for learning sparse representations that can capture semantic correlatio…

ClassificationDenoisingDictionary LearningEvent Detection+6

MMQ: Multimodal Mixture-of-Quantization Tokenization for Semantic ID Generation and User Behavioral Adaptation

2025-08-21 · Yi Xu, Moyu Zhang, Chenxuan Li, Zhihao Liao 외 arxiv

Recommender systems traditionally represent items using unique identifiers (ItemIDs), but this approach struggles with large, dynamic item corpora and sparse long-tail data, limiting scalability and generalization. Seman…