DecAug: Augmenting HOI Detection via Decomposition
Human-object interaction (HOI) detection requires a large amount of annotated data. Current algorithms suffer from insufficient training samples and category imbalance within datasets. To increase data efficiency, in this paper, we propose an efficient and effective data augmentation method called DecAug for HOI detection. Based on our proposed object state similarity metric, object patterns across different HOIs are shared to augment local object appearance features without changing their state. Further, we shift spatial correlation between humans and objects to other feasible configurations with the aid of a pose-guided Gaussian Mixture Model while preserving their interactions. Experiments show that our method brings up to 3.3 mAP and 1.6 mAP improvements on V-COCO and HICODET dataset for two advanced models. Specifically, interactions with fewer samples enjoy more notable improvement. Our method can be easily integrated into various HOI detection models with negligible extra computational consumption. Our code will be made publicly available.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationDomain GeneralizationHuman-Object Interaction DetectionObjectSimilar Papers 제목 키워드 기반
DecAug: Out-of-Distribution Generalization via Decomposed Feature Representation and Semantic Augmentation
While deep learning demonstrates its strong ability to handle independent and identically distributed (IID) data, it often suffers from out-of-distribution (OoD) generalization, where the test data come from another dist…
Domain GeneralizationImage ClassificationOut-of-Distribution GeneralizationPointAugmenting: Cross-Modal Augmentation for 3D Object Detection
Camera and LiDAR are two complementary sensors for 3D object detection in the autonomous driving context. Camera provides rich texture and color cues while LiDAR specializes in relative distance sensing. The challeng…
3D Object DetectionAutonomous DrivingData AugmentationObject+3Algorithmic progress in computer vision
We investigate algorithmic progress in image classification on ImageNet, perhaps the most well-known test bed for computer vision. We estimate a model, informed by work on neural scaling laws, and infer a decomposition o…
Attributeimage-classificationImage ClassificationAGILE: Approach-based Grasp Inference Learned from Element Decomposition
Humans, this species expert in grasp detection, can grasp objects by taking into account hand-object positioning information. This work proposes a method to enable a robot manipulator to learn the same, grasping objects …
Domain AdaptationRobotic GraspingEnd-to-End Neural Context Reconstruction in Chinese Dialogue
We tackle the problem of context reconstruction in Chinese dialogue, where the task is to replace pronouns, zero pronouns, and other referring expressions with their referent nouns so that sentences can be processed in i…
coreference-resolutionCoreference ResolutionPOSPosition+1