Linguistic Structures as Weak Supervision for Visual Scene Graph Generation
Prior work in scene graph generation requires categorical supervision at the level of triplets - subjects and objects, and predicates that relate them, either with or without bounding box information. However, scene graph generation is a holistic task: thus holistic, contextual supervision should intuitively improve performance. In this work, we explore how linguistic structures in captions can benefit scene graph generation. Our method captures the information provided in captions about relations between individual triplets, and context for subjects and objects (e.g. visual properties are mentioned). Captions are a weaker type of supervision than triplets since the alignment between the exhaustive list of human-annotated subjects and objects in triplets, and the nouns in captions, is weak. However, given the large and diverse sources of multimodal data on the web (e.g. blog posts with images and captions), linguistic supervision is more scalable than crowdsourced triplets. We show extensive experimental comparisons against prior methods which leverage instance- and image-level supervision, and ablate our method to show the impact of leveraging phrasal and sequential context, and techniques to improve localization of subjects and objects.
Code (1)
Tasks
Graph GenerationScene Graph GenerationSimilar Papers 제목 키워드 기반
WeaQA: Weak Supervision via Captions for Visual Question Answering
Methodologies for training visual question answering (VQA) models assume the availability of datasets with human-annotated \textit{Image-Question-Answer} (I-Q-A) triplets. This has led to heavy reliance on datasets and a…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)VirPro: Visual-referred Probabilistic Prompt Learning for Weakly-Supervised Monocular 3D Detection
Monocular 3D object detection typically relies on pseudo-labeling techniques to reduce dependency on real-world annotations. Recent advances demonstrate that deterministic linguistic cues can serve as effective auxiliary…
Monocular 3D Object DetectionWeakly-Supervised 3D Scene Graph Generation via Visual-Linguistic Assisted Pseudo-labeling
Learning to build 3D scene graphs is essential for real-world perception in a structured and rich fashion. However, previous 3D scene graph generation methods utilize a fully supervised learning manner and require a larg…
3d scene graph generationGraph GenerationGraph Neural NetworkScene Graph GenerationST-LDM: A Universal Framework for Text-Grounded Object Generation in Real Images
We present a novel image editing scenario termed Text-grounded Object Generation (TOG), defined as generating a new object in the real image spatially conditioned by textual descriptions. Existing diffusion models exhibi…
DenoisingLinguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition
Text images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and appearance-based features, while the linguistic dimension incorporates con…
Contrastive LearningScene Text RecognitionSelf-Supervised Learningself-supervised scene text recognition