paper-with-me

Papers

Wukong: Towards a Scaling Law for Large-Scale Recommendation

2024-03-04 · Buyun Zhang, Liang Luo, Yuxin Chen, Jade Nie, Xi Liu, Daifeng Guo, Yanli Zhao, Shen Li, Yuchen Hao, Yantao Yao, Guna Lakshminarayanan, Ellie Dingqiao Wen, Jongsoo Park, Maxim Naumov, Wenlin Chen

Scaling laws play an instrumental role in the sustainable improvement in model quality. Unfortunately, recommendation models to date do not exhibit such laws similar to those observed in the domain of large language models, due to the inefficiencies of their upscaling mechanisms. This limitation poses significant challenges in adapting these models to increasingly more complex real-world datasets. In this paper, we propose an effective network architecture based purely on stacked factorization machines, and a synergistic upscaling strategy, collectively dubbed Wukong, to establish a scaling law in the domain of recommendation. Wukong's unique design makes it possible to capture diverse, any-order of interactions simply through taller and wider layers. We conducted extensive evaluations on six public datasets, and our results demonstrate that Wukong consistently outperforms state-of-the-art models quality-wise. Further, we assessed Wukong's scalability on an internal, large-scale dataset. The results show that Wukong retains its superiority in quality over state-of-the-art models, while holding the scaling law across two orders of magnitude in model complexity, extending beyond 100 GFLOP/example, where prior arts fall short.

📄 PDF Abstract BibTeX arXiv:2403.02545

Code (2)

clabrugere/wukong-recommendation pytorch
reczoo/FuxiCTR/tree/main/model_zoo/WuKong pytorch

Tasks

Language ModellingLarge Language Model

Similar Papers 제목 키워드 기반

WHALE: A Scalable Unified Model for Recommendation with Wukong-HSTU Architecture

2026-07-19 · Renqin Cai, Dawei Sun, Yuanjun Yao, Zhiyong Wang 외 arxiv

As scalability becomes increasingly important in recommendation modeling, recent architectures have advanced the modeling of two broad sources of ranking signals along separate paths: non-sequence features, including use…

Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training Benchmark

2022-02-14 · Jiaxi Gu, Xiaojun Meng, Guansong Lu, Lu Hou 외

Vision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datasets. However, the lack of large-scale dat…

BenchmarkingContrastive Learningimage-classificationImage Classification+6

MTmixAtt: Integrating Mixture-of-Experts with Multi-Mix Attention for Large-Scale Recommendation

2025-10-17 · Xianyang Qi, Yuan Tian, Zhaoyu Hu, Zhirui Kuai 외 arxiv

Industrial recommender systems critically depend on high-quality ranking models. However, traditional pipelines still rely on manual feature engineering and scenario-specific architectures, which hinder cross-scenario tr…

Feature Engineering

VoiceWukong: Benchmarking Deepfake Voice Detection

2024-09-10 · Ziwei Yan, Yanjie Zhao, Haoyu Wang

With the rapid advancement of technologies like text-to-speech (TTS) and voice conversion (VC), detecting deepfake voices has become increasingly crucial. However, both academia and industry lack a comprehensive and intu…

BenchmarkingFace SwappingLarge Language Modeltext-to-speech+2

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling

2024-10-08 · Xudong Xie, Hao Yan, Liang Yin, Yang Liu 외

Multimodal document understanding is a challenging task to process and comprehend large amounts of textual and visual information. Recent advances in Large Language Models (LLMs) have significantly improved the performan…

document understandingLanguage ModelingLanguage ModellingLarge Language Model+2