Visual Relationship Detection
5개 벤치마크 · 논문 82편 · 이 태스크의 논문 보기 →
Benchmarks
VRD Relationship Detection
VRD Phrase Detection
VRD Predicate Detection
VRD
Visual Genome
Most implemented
Exploring Long Tail Visual Relationship Recognition with Large Vocabulary
Graphical Contrastive Losses for Scene Graph Parsing
Representing Prior Knowledge Using Randomly, Weighted Feature Networks for Visual Relationship Detection
Spatial-Temporal Transformer for Dynamic Scene Graph Generation
Papers
METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection
Open-vocabulary video visual relationship detection aims to detect objects and their relationships in videos without being restricted by predefined object or relationship categories. Existing methods leverage the rich se…
Objectobject-detectionObject DetectionRelationship Detection+1End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting
Open-vocabulary video visual relationship detection aims to expand video visual relationship detection beyond annotated categories by detecting unseen relationships between both seen and unseen objects in videos. Existin…
DecoderObjectobject-detectionObject Detection+5Groupwise Query Specialization and Quality-Aware Multi-Assignment for Transformer-based Visual Relationship Detection
Visual Relationship Detection (VRD) has seen significant advancements with Transformer-based architectures recently. However, we identify two key limitations in a conventional label assignment for training Transformer-ba…
RelationRelationship DetectionScene Graph GenerationVisual Relationship DetectionScene-Graph ViT: End-to-End Open-Vocabulary Visual Relationship Detection
Visual relationship detection aims to identify objects and their relationships in images. Prior methods approach this task by adding separate relationship modules or decoders to existing object detection architectures. T…
DecoderObjectobject-detectionObject Detection+2Video Relationship Detection Using Mixture of Experts
Machine comprehension of visual information from images and videos by neural networks faces two primary challenges. Firstly, there exists a computational and inference gap in connecting vision and language, making it dif…
Action RecognitionMixture-of-ExpertsObjectReading Comprehension+3RelVAE: Generative Pretraining for few-shot Visual Relationship Detection
Visual relations are complex, multimodal concepts that play an important role in the way humans perceive the world. As a result of their complexity, high-quality, diverse and large scale datasets for visual relations are…
Predicate ClassificationRelationship DetectionVisual Relationship Detection