Local-Global Information Interaction Debiasing for Dynamic Scene Graph Generation
The task of dynamic scene graph generation (DynSGG) aims to generate scene graphs for given videos, which involves modeling the spatial-temporal information in the video. However, due to the long-tailed distribution of samples in the dataset, previous DynSGG models fail to predict the tail predicates. We argue that this phenomenon is due to previous methods that only pay attention to the local spatial-temporal information and neglect the consistency of multiple frames. To solve this problem, we propose a novel DynSGG model based on multi-task learning, DynSGG-MTL, which introduces the local interaction information and global human-action interaction information. The interaction between objects and frame features makes the model more fully understand the visual context of the single image. Long-temporal human actions supervise the model to generate multiple scene graphs that conform to the global constraints and avoid the model being unable to learn the tail predicates. Extensive experiments on Action Genome dataset demonstrate the efficacy of our proposed framework, which not only improves the dynamic scene graph generation but also alleviates the long-tail problem.
Code (0)
등록된 구현이 없습니다.
Tasks
Graph GenerationMulti-Task LearningScene Graph GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
It Is Different When Items Are Older: Debiasing Recommendations When Selection Bias and User Preferences Are Dynamic
User interactions with recommender systems (RSs) are affected by user selection bias, e.g., users are more likely to rate popular items (popularity bias) or items that they expect to enjoy beforehand (positivity bias). M…
Recommendation SystemsSelection biasFreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency Debiasing
Deepfake detectors often struggle to generalize to novel forgery types due to biases learned from limited training data. In this paper, we identify a new type of model bias in the frequency domain, termed spectral bias, …
Representation LearningDomain GeneralizationDeepFake DetectionFreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency Debiasing
Deepfake detectors often struggle to generalize to novel forgery types due to biases learned from limited training data. In this paper, we identify a new type of model bias in the frequency domain, termed spectral bi…
DeepFake DetectionDomain GeneralizationFace SwappingRepresentation LearningGlobal-local Motion Transformer for Unsupervised Skeleton-based Action Learning
We propose a new transformer model for the task of unsupervised learning of skeleton motion sequences. The existing transformer model utilized for unsupervised skeleton-based action learning is learned the instantaneous …
Use Symmetry to Elucidate the Roles of Global Shape and Local Interactions in Protein Dynamics and Cooperativity
Shape had been intuitively recognized to play a dominant role in determining the global motion patterns of bio-molecular assemblies. However, it is not clear exactly how shape determines the motion patterns. What about t…