paper-with-me

Papers

Towards Sharing Task Environments to Support Reproducible Evaluations of Interactive Recommender Systems

2019-09-13 · Andrea Barraza-Urbina, Mathieu d'Aquin

Beyond sharing datasets or simulations, we believe the Recommender Systems (RS) community should share Task Environments. In this work, we propose a high-level logical architecture that will help to reason about the core components of a RS Task Environment, identify the differences between Environments, datasets and simulations; and most importantly, understand what needs to be shared about Environments to achieve reproducible experiments. The work presents itself as valuable initial groundwork, open to discussion and extensions.

📄 PDF Abstract BibTeX arXiv:1909.06133

Code (0)

등록된 구현이 없습니다.

Tasks

Recommendation Systems

Similar Papers 제목 키워드 기반

Reproducible Subjective Evaluation

2022-03-08 · Max Morrison, Brian Tang, Gefei Tan, Bryan Pardo

Human perceptual studies are the gold standard for the evaluation of many research tasks in machine learning, linguistics, and psychology. However, these studies require significant time and cost to perform. As a result,…

Gym-Ignition: Reproducible Robotic Simulations for Reinforcement Learning

2019-11-05 · Diego Ferigo, Silvio Traversaro, Giorgio Metta, Daniele Pucci

This paper presents Gym-Ignition, a new framework to create reproducible robotic environments for reinforcement learning research. It interfaces with the new generation of Gazebo, part of the Ignition Robotics suite, whi…

OpenAI Gymreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Radiuma: A Unified Zero-Code Executable Graphical Workflow Generator for Reproducible and Shareable Medical Image Analysis and Machine Learning

2026-05-22 · Mohammad Salmanpour, Mehrdad Oveisi, Isaac Shiri, Arman Rahmim arxiv

Medical image computing software is essential for identifying imaging biomarkers that can support diagnosis, prognosis, treatment planning, and clinical research. However, the lack of standardized, user-friendly, and rep…

Tonic: A Deep Reinforcement Learning Library for Fast Prototyping and Benchmarking

2020-11-15 · Fabio Pardo

Deep reinforcement learning has been one of the fastest growing fields of machine learning over the past years and numerous libraries have been open sourced to support research. However, most codebases have a steep learn…

Benchmarkingcontinuous-controlContinuous ControlDeep Reinforcement Learning+3

REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

2025-04-15 · Divyansh Garg, Shaun VanWeelden, Diego Caples, Andis Draguns 외

We introduce REAL, a benchmark and framework for multi-turn agent evaluations on deterministic simulations of real-world websites. REAL comprises high-fidelity, deterministic replicas of 11 widely-used websites across do…

Autonomous Web NavigationBenchmarkingInformation RetrievalRetrieval