Generalized Referring Expression Comprehension
1개 벤치마크 · 논문 10편 · 이 태스크의 논문 보기 →
Benchmarks
gRefCOCO
Most implemented
MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding
Multi-task Collaborative Network for Joint Referring Expression Comprehension and Segmentation
GREC: Generalized Referring Expression Comprehension
Universal Instance Perception as Object Discovery and Retrieval
Vision-Language Transformer and Query Generation for Referring Segmentation
Papers
Making Dialogue Grounding Data Rich: A Three-Tier Data Synthesis Framework for Generalized Referring Expression Comprehension
Dialogue-Based Generalized Referring Expression Comprehension (GREC) requires models to ground the expression and unlimited targets in complex visual scenes while resolving coreference across a long dialogue context. How…
Generalized Referring Expression ComprehensionLIHE: Linguistic Instance-Split Hyperbolic-Euclidean Framework for Generalized Weakly-Supervised Referring Expression Comprehension
Existing Weakly-Supervised Referring Expression Comprehension (WREC) methods, while effective, are fundamentally limited by a one-to-one mapping assumption, hindering their ability to handle expressions corresponding to …
Generalized Referring Expression ComprehensionImproving Generalized Visual Grounding with Instance-aware Joint Learning
Generalized visual grounding tasks, including Generalized Referring Expression Comprehension (GREC) and Segmentation (GRES), extend the classical visual grounding paradigm by accommodating multi-target and non-target sce…
Generalized Referring Expression ComprehensionSemantic SegmentationVisual GroundingHierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension
In this work, we address the challenging task of Generalized Referring Expression Comprehension (GREC). Compared to the classic Referring Expression Comprehension (REC) that focuses on single-target expressions, GREC ext…
Generalized Referring Expression ComprehensionGeneralized Referring Expression SegmentationObject CountingPhrase Grounding+3SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion
Visual grounding is a common vision task that involves grounding descriptive sentences to the corresponding regions of an image. Most existing methods use independent image-text encoding and apply complex hand-crafted mo…
DescriptiveGeneralized Referring Expression ComprehensionReferring Expression ComprehensionVisual GroundingGREC: Generalized Referring Expression Comprehension
The objective of Classic Referring Expression Comprehension (REC) is to produce a bounding box corresponding to the object mentioned in a given textual description. Commonly, existing datasets and techniques in classic R…
Generalized Referring Expression ComprehensionReferring ExpressionReferring Expression Comprehension