paper-with-me

홈 › Papers

ARTA: Collection and Classification of Ambiguous Requests and Thoughtful Actions

2021-06-15 · SIGDIAL (ACL) 2021 7 · Shohei Tanaka, Koichiro Yoshino, Katsuhito Sudoh, Satoshi Nakamura

Human-assisting systems such as dialogue systems must take thoughtful, appropriate actions not only for clear and unambiguous user requests, but also for ambiguous user requests, even if the users themselves are not aware of their potential requirements. To construct such a dialogue agent, we collected a corpus and developed a model that classifies ambiguous user requests into corresponding system actions. In order to collect a high-quality corpus, we asked workers to input antecedent user requests whose pre-defined actions could be regarded as thoughtful. Although multiple actions could be identified as thoughtful for a single user request, annotating all combinations of user requests and system actions is impractical. For this reason, we fully annotated only the test data and left the annotation of the training data incomplete. In order to train the classification model on such training data, we applied the positive/unlabeled (PU) learning method, which assumes that only a part of the data is labeled with positive examples. The experimental results show that the PU learning method achieved better performance than the general positive/negative (PN) learning method to classify thoughtful actions given an ambiguous user request.

📄 PDF Abstract BibTeX arXiv:2106.07999

Code (1)

ahclab/arta_corpus 공식 구현

Tasks

Classification

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

Private, Augmentation-Robust and Task-Agnostic Data Valuation Approach for Data Marketplace

2024-11-01 · Tayyebeh Jahani-Nezhad, Parsa Moradi, Mohammad Ali Maddah-Ali, Giuseppe Caire

Evaluating datasets in data marketplaces, where the buyer aim to purchase valuable data, is a critical challenge. In this paper, we introduce an innovative task-agnostic data valuation method called PriArTa which is an a…

Data ValuationPrivacy Preserving

HumanMCP: A Human-Like Query Dataset for Evaluating MCP Tool Retrieval Performance

2025-12-18 · Shubh Laddha, Lucas Changbencharoen, Win Kuptivej, Surya Shringla 외 arxiv

Model Context Protocol (MCP) servers contain a collection of thousands of open-source standardized tools, linking LLMs to external systems; however, existing datasets and benchmarks lack realistic, human-like user querie…

TartanGround: A Large-Scale Dataset for Ground Robot Perception and Navigation

2025-05-15 · Manthan Patel, Fan Yang, Yuheng Qiu, Cesar Cadena 외

We present TartanGround, a large-scale, multi-modal dataset to advance the perception and autonomy of ground robots operating in diverse environments. This dataset, collected in various photorealistic simulation environm…

Optical Flow Estimation

Data Collection for Dialogue System: A Startup Perspective

2018-06-01 · NAACL 2018 6 · Yiping Kang, Yunqi Zhang, Jonathan K. Kummerfeld, Lingjia Tang 외

Industrial dialogue systems such as Apple Siri and Google Now rely on large scale diverse and robust training data to enable their sophisticated conversation capability. Crowdsourcing provides a scalable and inexpensive …

General Classificationintent-classificationIntent ClassificationText Classification

Reasoning about Intent for Ambiguous Requests

2025-11-13 · Irina Saparina, Mirella Lapata arxiv

Large language models often respond to ambiguous requests by implicitly committing to one interpretation, frustrating users and creating safety risks when that interpretation is wrong. We propose generating a single stru…

Conversational Question AnsweringReinforcement LearningSemantic Parsing