paper-with-me

홈 › Papers

ConvApparel: A Benchmark Dataset and Validation Framework for User Simulators in Conversational Recommenders

2026-02-18 · Ofer Meshi, Krisztian Balog, Sally Goldman, Avi Caciularu, Guy Tennenholtz, Jihwan Jeong, Amir Globerson, Craig Boutilier arxiv

The promise of LLM-based user simulators to improve conversational AI is hindered by a critical "realism gap," leading to systems that are optimized for simulated interactions, but may fail to perform well in the real world. We introduce ConvApparel, a new dataset of human-AI conversations designed to address this gap. Its unique dual-agent data collection protocol -- using both "good" and "bad" recommenders -- enables counterfactual validation by capturing a wide spectrum of user experiences, enriched with first-person annotations of user satisfaction. We propose a comprehensive validation framework that combines statistical alignment, a human-likeness score, and counterfactual validation to test for generalization. Our experiments reveal a significant realism gap across all simulators. However, the framework also shows that data-driven simulators outperform a prompted baseline, particularly in counterfactual validation where they adapt more realistically to unseen behaviors, suggesting they embody more robust, if imperfect, user models.

📄 PDF Abstract BibTeX arXiv:2602.16938

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GBV-SQL: Guided Generation and SQL2Text Back-Translation Validation for Multi-Agent Text2SQL

2025-09-16 · Daojun Chen, Xi Wang, Shenyuan Ren, Qingzhi Ma 외 arxiv

While Large Language Models have significantly advanced Text2SQL generation, a critical semantic gap persists where syntactically valid queries often misinterpret user intent. To mitigate this challenge, we propose GBV-S…

BIRL: Benchmark on Image Registration methods with Landmark validation

2019-12-31 · Jiri Borovec

This report presents a generic image registration benchmark with automatic evaluation using landmark annotations. The key features of the BIRL framework are: easily extendable, performance evaluation, parallel experiment…

BIRLImage RegistrationMedical Image Registration

Deep Adversarial Learning with Activity-Based User Discrimination Task for Human Activity Recognition

2024-10-01 · Francisco M. Calatrava-Nicolás, Shoko Miyauchi, Oscar Martinez Mozos

We present a new adversarial deep learning framework for the problem of human activity recognition (HAR) using inertial sensors worn by people. Our framework incorporates a novel adversarial activity-based discrimination…

Activity RecognitionHuman Activity Recognition

Ask, Fail, Repeat: Meeseeks, an Iterative Feedback Benchmark for LLMs' Multi-turn Instruction-Following Ability

2025-04-30 · JiaMing Wang, Yunke Zhao, Peng Ding, Jun Kuang 외

The ability to follow instructions accurately is fundamental for Large Language Models (LLMs) to serve as reliable agents in real-world applications. For complex instructions, LLMs often struggle to fulfill all requireme…

Instruction FollowingIntent Recognition

Multifaceted User Modeling in Recommendation: A Federated Foundation Models Approach

2024-12-22 · Chunxu Zhang, Guodong Long, Hongkuan Guo, Zhaojie Liu 외

Multifaceted user modeling aims to uncover fine-grained patterns and learn representations from user data, revealing their diverse interests and characteristics, such as profile, preference, and personality. Recent studi…