paper-with-me

Papers

Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition

2026-01-22 · Geo Ahn, Inwoong Lee, Taeoh Kim, Minho Shim, Dongyoon Wee, Jinwoo Choi arxiv

Zero-Shot Compositional Action Recognition (ZS-CAR) requires recognizing novel verb-object combinations composed of previously observed primitives. In this work, we tackle a key failure mode: models predict verbs via object-driven shortcuts (i.e., relying on the labeled object class) rather than temporal evidence. We argue that sparse compositional supervision and verb-object learning asymmetry can promote object-driven shortcut learning. Our analysis with proposed diagnostic metrics shows that existing methods overfit to training co-occurrence patterns and underuse temporal verb cues, resulting in weak generalization to unseen compositions. To address object-driven shortcuts, we propose Robust COmpositional REpresentations (RCORE) with two components. Co-occurrence Prior Regularization (CPR) adds explicit supervision for unseen compositions and regularizes the model against frequent co-occurrence priors by treating them as hard negatives. Temporal Order Regularization for Composition (TORC) enforces temporal-order sensitivity to learn temporally grounded verb representations. Across Sth-com and EK100-com, RCORE reduces shortcut diagnostics and consequently improves compositional generalization.

📄 PDF Abstract BibTeX arXiv:2601.16211

Code (0)

등록된 구현이 없습니다.

Tasks

Action Recognition

Similar Papers 제목 키워드 기반

Learning to Solve Tasks with Exploring Prior Behaviours

2023-07-06 · Ruiqi Zhu, Siyuan Li, Tianhong Dai, Chongjie Zhang 외

Demonstrations are widely used in Deep Reinforcement Learning (DRL) for facilitating solving tasks with sparse rewards. However, the tasks in real-world scenarios can often have varied initial conditions from the demonst…

Deep Reinforcement Learning

Spot-Compose: A Framework for Open-Vocabulary Object Retrieval and Drawer Manipulation in Point Clouds

2024-04-18 · Oliver Lemke, Zuria Bauer, René Zurbrügg, Marc Pollefeys 외

In recent years, modern techniques in deep learning and large-scale datasets have led to impressive progress in 3D instance segmentation, grasp pose estimation, and robotics. This allows for accurate detection directly i…

3D Instance SegmentationInstance SegmentationPose EstimationRetrieval+2

OpenD: A Benchmark for Language-Driven Door and Drawer Opening

2022-12-10 · Yizhou Zhao, Qiaozi Gao, Liang Qiu, Govind Thattai 외

We introduce OPEND, a benchmark for learning how to use a hand to open cabinet doors or drawers in a photo-realistic and physics-reliable simulation environment driven by language instruction. To solve the task, we propo…

Spatial Reasoning

Online Estimation and Manipulation of Articulated Objects

2026-01-04 · Russell Buchanan, Adrian Röfer, João Moura, Abhinav Valada 외 arxiv

From refrigerators to kitchen drawers, humans interact with articulated objects effortlessly every day while completing household chores. For automating these tasks, service robots must be capable of manipulating arbitra…

CoDraw: Collaborative Drawing as a Testbed for Grounded Goal-driven Communication

2017-12-15 · ACL 2019 7 · Jin-Hwa Kim, Nikita Kitaev, Xinlei Chen, Marcus Rohrbach 외

In this work, we propose a goal-driven collaborative task that combines language, perception, and action. Specifically, we develop a Collaborative image-Drawing game between two agents, called CoDraw. Our game is grounde…

Imitation Learning