paper-with-me

Papers

RBoard: A Unified Platform for Reproducible and Reusable Recommender System Benchmarks

2024-09-09 · Xinyang Shao, Edoardo D'Amico, Gabor Fodor, Tri Kurniawan Wijaya

Recommender systems research lacks standardized benchmarks for reproducibility and algorithm comparisons. We introduce RBoard, a novel framework addressing these challenges by providing a comprehensive platform for benchmarking diverse recommendation tasks, including CTR prediction, Top-N recommendation, and others. RBoard's primary objective is to enable fully reproducible and reusable experiments across these scenarios. The framework evaluates algorithms across multiple datasets within each task, aggregating results for a holistic performance assessment. It implements standardized evaluation protocols, ensuring consistency and comparability. To facilitate reproducibility, all user-provided code can be easily downloaded and executed, allowing researchers to reliably replicate studies and build upon previous work. By offering a unified platform for rigorous, reproducible evaluation across various recommendation scenarios, RBoard aims to accelerate progress in the field and establish a new standard for recommender systems benchmarking in both academia and industry. The platform is available at https://rboard.org and the demo video can be found at https://bit.ly/rboard-demo.

📄 PDF Abstract BibTeX arXiv:2409.05526

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingClick-Through Rate PredictionRecommendation Systems

Similar Papers 제목 키워드 기반

ORBIT -- Open Recommendation Benchmark for Reproducible Research with Hidden Tests

2025-10-30 · Jingyuan He, Jiongnan Liu, Vishan Vishesh Oberoi, Bolin Wu 외 arxiv

Recommender systems are among the most impactful AI applications, interacting with billions of users every day, guiding them to relevant products, services, or information tailored to their preferences. However, the rese…

The Collective Knowledge project: making ML models more portable and reproducible with open APIs, reusable best practices and MLOps

2020-06-12 · Grigori Fursin

This article provides an overview of the Collective Knowledge technology (CK or cKnowledge). CK attempts to make it easier to reproduce ML&systems research, deploy ML models in production, and adapt them to continuously …

Benchmarkingobject-detectionObject Detection

RecAD: Towards A Unified Library for Recommender Attack and Defense

2023-09-09 · Changsheng Wang, Jianbai Ye, Wenjie Wang, Chongming Gao 외

In recent years, recommender systems have become a ubiquitous part of our daily lives, while they suffer from a high risk of being attacked due to the growing commercial and social values. Despite significant research pr…

BenchmarkingRecommendation Systems

AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems

2025-05-26 · Yu Shang, Peijie Liu, Yuwei Yan, Zijing Wu 외

The emergence of agentic recommender systems powered by Large Language Models (LLMs) represents a paradigm shift in personalized recommendations, leveraging LLMs' advanced reasoning and role-playing capabilities to enabl…

BenchmarkingRecommendation Systems

Designing UNICORN: a Unified Benchmark for Imaging in Computational Pathology, Radiology, and Natural Language

2026-03-03 · Michelle Stegeman, Lena Philipp, Fennie van der Graaf, Marina D'Amato 외 arxiv

Medical foundation models show promise to learn broadly generalizable features from large, diverse datasets. This could be the base for reliable cross-modality generalization and rapid adaptation to new, task-specific go…