paper-with-me

홈 › Papers

Hey AI, Can You Solve Complex Tasks by Talking to Agents?

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Training giant models from scratch for each complex task is resource- and data-inefficient. To help develop models that can leverage existing systems, we propose a new challenge: Learning to solve complex tasks by communicating with existing agents (or models) in natural language. We design a synthetic benchmark, CommaQA, with three complex reasoning tasks (explicit, implicit, numeric) designed to be solved by communicating with existing QA agents. For instance, using text and table QA agents to answer questions such as "Who had the longest javelin throw from USA?". We show that black-box models struggle to learn this task from scratch (accuracy under 50\%) even with access to each agent's knowledge and gold facts supervision. In contrast, models that learn to communicate with agents outperform black-box models, reaching scores of 100\% when given gold decomposition supervision. However, we show that the challenge of learning to solve complex tasks by communicating with existing agents \emph{without relying on any auxiliary supervision or data} still remains highly elusive. We will release CommaQA, along with a compositional generalization test split, to advance research in this direction.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Hey AI, Can You Solve Complex Tasks by Talking to Agents?

2021-10-16 · Findings (ACL) 2022 5 · Tushar Khot, Kyle Richardson, Daniel Khashabi, Ashish Sabharwal

Training giant models from scratch for each complex task is resource- and data-inefficient. To help develop models that can leverage existing systems, we propose a new challenge: Learning to solve complex tasks by commun…

GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes

2026-06-26 · Amit Parekh, Sabrina McCallum, Kareem Al-Hasan, Malvina Nikandrou 외 arxiv

Multimodal models are increasingly deployed to solve tasks collaboratively with humans or other artificial agents. While existing benchmarks show that they possess the fundamental capabilities, the various conditions tha…

Interactive Conversational Head Generation

2023-07-05 · Mohan Zhou, Yalong Bai, Wei zhang, Ting Yao 외

We introduce a new conversation head generation benchmark for synthesizing behaviors of a single interlocutor in a face-to-face conversation. The capability to automatically synthesize interlocutors which can participate…

SentenceTalking Head Generation

AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

2023-04-25 · Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi 외

Large language models (LLMs) have exhibited remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. Despite the recent success, current LLMs are not capable…

Talking with the Theorem Prover to Interactively Solve Natural Language Inference

2021-11-01 · PACLIC 2021 11 · Atsushi Sumita, Yusuke Miyao, Koji Mineshima
Natural Language Inference