paper-with-me

Papers

RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation

2025-09-03 · Sashuai Zhou, Weinan Gan, Qijiong Liu, Ke Lei, Jieming Zhu, Hai Huang, Yan Xia, Ruiming Tang, Zhenhua Dong, Zhou Zhao arxiv

Recent advances in LLM-based recommendation have shown promise, yet their cross-domain generalization is hindered by a fundamental mismatch between language-centric pretraining and the recommendation task. Existing methods, relying on language-level knowledge, fail to capture dynamic, item-level user interests across domains. To bridge this gap, we propose RecBase, a domain-agnostic foundational model pretrained with a recommendation-oriented objective. RecBase leverages a large-scale, heterogeneous, cross-domain corpus with unified textual representations and feature mappings to enhance cross-domain generalization. To further align item semantics across domains, we introduce a unified item tokenizer that encodes items into hierarchical concept identifiers, enabling structured representation and efficient vocabulary sharing. The model is trained using an autoregressive objective to capture complex item-level sequential patterns. On eight real-world datasets, our 1.5B-parameter model matches or surpasses the performance of LLM baselines up to 7B parameters in zero-shot and cross-domain recommendation tasks.

📄 PDF Abstract BibTeX arXiv:2509.03131

Code (0)

등록된 구현이 없습니다.

Tasks

Domain Generalization

Similar Papers 제목 키워드 기반

Zero-shot Medical Event Prediction Using a Generative Pre-trained Transformer on Electronic Health Records

2025-03-07 · Ekaterina Redekop, Zichen Wang, Rushikesh Kulkarni, Mara Pleasure 외

Longitudinal data in electronic health records (EHRs) represent an individual`s clinical history through a sequence of codified concepts, including diagnoses, procedures, medications, and laboratory tests. Foundational m…

Diagnostic

VideoPoet: A Large Language Model for Zero-Shot Video Generation

2023-12-21 · Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama 외

We present VideoPoet, a language model capable of synthesizing high-quality video, with matching audio, from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that process…

DecoderLanguage ModelingLanguage ModellingLarge Language Model+2

GraphCLIP: Enhancing Transferability in Graph Foundation Models for Text-Attributed Graphs

2024-10-14 · Yun Zhu, Haizhou Shi, Xiaotang Wang, Yongchao Liu 외

Recently, research on Text-Attributed Graphs (TAGs) has gained significant attention due to the prevalence of free-text node features in real-world applications and the advancements in Large Language Models (LLMs) that b…

Few-Shot LearningTAG

The effectiveness of MAE pre-pretraining for billion-scale pretraining

2023-03-23 · ICCV 2023 1 · Mannat Singh, Quentin Duval, Kalyan Vasudev Alwala, Haoqi Fan 외

This paper revisits the standard pretrain-then-finetune paradigm used in computer vision for visual recognition tasks. Typically, state-of-the-art foundation models are pretrained using large scale (weakly) supervised da…

Action ClassificationAction RecognitionFew-Shot Image Classificationimage-classification+6

Alethia: A Foundational Encoder for Voice Deepfakes

2026-04-30 · Yi Zhu, Brahmi Dwivedi, Jayaram Raghuram, Surya Koppisetti arxiv

Existing voice deepfake detection and localization models rely heavily on representations extracted from speech foundation models (SFMs). However, downstream finetuning has now reached a state of diminishing returns. In …

Zero-shot GeneralizationDeepFake Detection