Multimodal Graph Network Modeling for Human-Object Interaction Detection with PDE Graph Diffusion
Existing GNN-based Human-Object Interaction (HOI) detection methods rely on simple MLPs to fuse instance features and propagate information. However, this mechanism is largely empirical and lack of targeted information propagation process. To address this problem, we propose Multimodal Graph Network Modeling (MGNM) for HOI detection with Partial Differential Equation (PDE) graph diffusion. Specifically, we first design a multimodal graph network framework that explicitly models the HOI detection task within a four-stage graph structure. Next, we propose a novel PDE diffusion mechanism to facilitate information propagation within this graph. This mechanism leverages multimodal features to propaganda information via a white-box PDE diffusion equation. Furthermore, we design a variational information squeezing (VIS) mechanism to further refine the multimodal features extracted from CLIP, thereby mitigating the impact of noise inherent in pretrained Vision-Language Models. Extensive experiments demonstrate that our MGNM achieves state-of-the-art performance on two widely used benchmarks: HICO-DET and V-COCO. Moreover, when integrated with a more advanced object detector, our method yields significant performance gains while maintaining an effective balance between rare and non-rare categories.
Code (0)
등록된 구현이 없습니다.
Tasks
Human-Object Interaction DetectionSimilar Papers 제목 키워드 기반
Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI
Egocentric video captures activities from the wearer's perspective, providing a direct view of human attention, hand--object interaction, and goal-directed behavior. This perspective is increasingly important for wearabl…
Representation LearningDomain GeneralizationDecision MakingGeometric Visual Fusion Graph Neural Networks for Multi-Person Human-Object Interaction Recognition in Videos
Human-Object Interaction (HOI) recognition in videos requires understanding both visual patterns and geometric relationships as they evolve over time. Visual and geometric features offer complementary strengths. Visual f…
Graph LearningGraph Neural NetworkHuman-Object Interaction DetectionMultimodal Priors-Augmented Text-Driven 3D Human-Object Interaction Generation
We address the challenging task of text-driven 3D human-object interaction (HOI) motion generation. Existing methods primarily rely on a direct text-to-HOI mapping, which suffers from three key limitations due to the sig…
Spatio-Temporal Interaction Graph Parsing Networks for Human-Object Interaction Recognition
For a given video-based Human-Object Interaction scene, modeling the spatio-temporal relationship between humans and objects are the important cue to understand the contextual information presented in the video. With the…
Human-Object Interaction DetectionObjectAn Abstract Specification of VoxML as an Annotation Language
VoxML is a modeling language used to map natural language expressions into real-time visualizations using commonsense semantic knowledge of objects and events. Its utility has been demonstrated in embodied simulation env…
Human Agent CollaborationHuman-Object Interaction DetectionObject