paper-with-me

Papers

Simultaneous Detection and Interaction Reasoning for Object-Centric Action Recognition

2024-04-18 · Xunsong Li, Pengzhan Sun, Yangcen Liu, Lixin Duan, Wen Li

The interactions between human and objects are important for recognizing object-centric actions. Existing methods usually adopt a two-stage pipeline, where object proposals are first detected using a pretrained detector, and then are fed to an action recognition model for extracting video features and learning the object relations for action recognition. However, since the action prior is unknown in the object detection stage, important objects could be easily overlooked, leading to inferior action recognition performance. In this paper, we propose an end-to-end object-centric action recognition framework that simultaneously performs Detection And Interaction Reasoning in one stage. Particularly, after extracting video features with a base network, we create three modules for concurrent object detection and interaction reasoning. First, a Patch-based Object Decoder generates proposals from video patch tokens. Then, an Interactive Object Refining and Aggregation identifies important objects for action recognition, adjusts proposal scores based on position and appearance, and aggregates object-level info into a global video representation. Lastly, an Object Relation Modeling module encodes object relations. These three modules together with the video feature extractor can be trained jointly in an end-to-end fashion, thus avoiding the heavy reliance on an off-the-shelf object detector, and reducing the multi-stage training burden. We conduct experiments on two datasets, Something-Else and Ikea-Assembly, to evaluate the performance of our proposed approach on conventional, compositional, and few-shot action recognition tasks. Through in-depth experimental analysis, we show the crucial role of interactive objects in learning for action recognition, and we can outperform state-of-the-art methods on both datasets.

📄 PDF Abstract BibTeX arXiv:2404.11903

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionObjectobject-detectionObject Detection

Methods 이 논문이 사용한 방법론

INFO This study presents the analysis and principle of an innovative optimizer named weIghted meaN oF vectOrs (INFO) to optimize different problems. INFO is a modified weight mean…
BASE 설명 없음

Similar Papers 제목 키워드 기반

Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection

2026-04-02 · Soo Won Seo, KyungChae Lee, Hyungchan Cho, Taein Son 외 arxiv

Human-Object Interaction (HOI) detection aims to localize human-object pairs and classify their interactions from a single image, a task that demands strong visual understanding and nuanced contextual reasoning. Recent a…

Human-Object Interaction Detection

Spatial-Temporal Human-Object Interaction Detection

2025-08-24 · Xu Sun, Yunqing He, Tongwei Ren, Gangshan Wu arxiv

In this paper, we propose a new instance-level human-object interaction detection task on videos called ST-HOID, which aims to distinguish fine-grained human-object interactions (HOIs) and the trajectories of subjects an…

Human-Object Interaction DetectionVideo Visual Relation Detection

MECCANO: A Multimodal Egocentric Dataset for Humans Behavior Understanding in the Industrial-like Domain

2022-09-19 · Francesco Ragusa, Antonino Furnari, Giovanni Maria Farinella

Wearable cameras allow to acquire images and videos from the user's perspective. These data can be processed to understand humans behavior. Despite human behavior analysis has been thoroughly investigated in third person…

Action AnticipationAction RecognitionHuman-Object Interaction Detection

The MECCANO Dataset: Understanding Human-Object Interactions from Egocentric Videos in an Industrial-like Domain

2020-10-12 · Francesco Ragusa, Antonino Furnari, Salvatore Livatino, Giovanni Maria Farinella

Wearable cameras allow to collect images and videos of humans interacting with the world. While human-object interactions have been thoroughly investigated in third person vision, the problem has been understudied in ego…

Action RecognitionActive Object DetectionHuman-Object Interaction DetectionObject+3

Generating Human-Centric Visual Cues for Human-Object Interaction Detection via Large Vision-Language Models

2023-11-26 · Yu-Wei Zhan, Fan Liu, Xin Luo, Liqiang Nie 외

Human-object interaction (HOI) detection aims at detecting human-object pairs and predicting their interactions. However, the complexity of human behavior and the diverse contexts in which these interactions occur make i…

Human-Object Interaction Detection