paper-with-me

홈 › Papers

What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation

2026-02-12 · Zhenlong Yuan, Yue Wang, Dapeng Zhang, Kejin Cui, Rui Chen, Jing Tang, Lei Sun, Hongwei Yu, Chengxuan Qian, Xiangxiang Chu, Shuo Li, Yuyin Zhou arxiv

Multimodal Large Language Models have shown promising capabilities in bridging visual and textual reasoning, yet their reasoning capabilities in Open-Vocabulary Human-Object Interaction (OV-HOI) are limited by cross-modal hallucinations and limited viewpoints of images. To address this, we propose ImagineAgent, an agentic framework that integrates cognitive mapping, tool-augmented reinforcement learning (RL), and generative world modeling for robust OV-HOI understanding. Specifically, we first propose an innovative CoT dataset named hicodet-6K for supervised fine-tuning (SFT), which effectively bridges the perception-to-cognition gap by structuring perceived entities into interaction pairs for comprehensive predictions. Subsequently, we develop a multimodal tool library integrating online retrieval, image cropping, and generative modeling, enabling the agent to dynamically augment reasoning with domain-specific tools to resolve visual-semantic ambiguities and hallucinations during inference. Moreover, we incorporate a generative model to reconstruct alternative viewpoints, enabling the agent to 'imagine' under limited viewpoints. Finally, we propose a composite reward mechanism to jointly optimize prediction accuracy and tool efficiency. Evaluations on both SWIG-HOI and HICO-DET datasets demonstrate that our method achieves state-of-the-art performance while requiring merely 36.7% of the training data compared to existing methods, validating our robustness, empirical effectiveness and efficiency.

📄 PDF Abstract BibTeX arXiv:2602.11499

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningImage Cropping

Similar Papers 제목 키워드 기반

Language as a Cognitive Tool to Imagine Goals in Curiosity-Driven Exploration

2020-02-21 · Cédric Colas, Tristan Karch, Nicolas Lair, Jean-Michel Dussoux 외

Developmental machine learning studies how artificial agents can model the way children learn open-ended repertoires of skills. Such agents need to create and represent goals, select which ones to pursue and learn to ach…

Deep Reinforcement LearningLanguage AcquisitionLanguage ModellingReinforcement Learning

Language as a Cognitive Tool to Imagine Goals in Curiosity Driven Exploration

2020-12-01 · NeurIPS 2020 12 · Cédric Colas, Tristan Karch, Nicolas Lair, Jean-Michel Dussoux 외

Developmental machine learning studies how artificial agents can model the way children learn open-ended repertoires of skills. Such agents need to create and represent goals, select which ones to pursue and learn to ach…

Deep Reinforcement Learning

Language-Goal Imagination to Foster Creative Exploration in Deep RL

2020-06-12 · ICML Workshop LaReL 2020 7 · Tristan Karch, Nicolas Lair, Cédric Colas, Jean-Michel Dussoux 외

Developmental machine learning studies how artificial agents can model the way children learn open-ended repertoires of skills. Children are known to use language and its compositionality as a tool to imagine description…

A New Paradigm for Counterfactual Reasoning in Fairness and Recourse

2024-01-25 · Lucius E. J. Bynum, Joshua R. Loftus, Julia Stoyanovich

Counterfactuals and counterfactual reasoning underpin numerous techniques for auditing and understanding artificial intelligence (AI) systems. The traditional paradigm for counterfactual reasoning in this literature is t…

counterfactualCounterfactual ReasoningFairness

Imagined Autocurricula

2025-09-11 · Ahmet H. Güzel, Matthew Thomas Jackson, Jarek Luca Liesen, Tim Rocktäschel 외 arxiv

Training agents to act in embodied environments typically requires vast training data or access to accurate simulation, neither of which exists for many cases in the real world. Instead, world models are emerging as an a…