Referring Expression
1개 벤치마크 · 논문 425편 · 이 태스크의 논문 보기 →
Benchmarks
SQA3D
Most implemented
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
UNITER: UNiversal Image-TExt Representation Learning
Image Segmentation Using Text and Image Prompts
MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding
Towards Visual Grounding: A Survey
Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Papers
A-PAIR: A Benchmark and Identity-Consistent Grounding Framework for Air-Ground Cross-View Referring Person Detection
Air-ground cross-view referring person detection is a necessary component in the language-to-perception-to-control chain of collective embodied intelligence, grounding a language command into the same physical target bef…
Referring ExpressionWhen Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs
Visual grounding is typically evaluated as a one-shot mapping from an informative referring expression to a visual target. This formulation misses a central property of real-world reference: initial referring expressions…
Referring ExpressionVisual GroundingGroupForward: Building Referable 3D Scenes via Instance-Grouped Feed-Forward Gaussian Splatting
Simultaneously reconstructing and understanding 3D environments is essential for embodied agents. Toward this goal, feed-forward semantic 3D Gaussian Splatting (3DGS) efficiently constructs semantic scene representations…
Referring ExpressionScaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding
Referring Expression Comprehension (REC) is commonly studied under dataset-specific fine-tuning, resulting in specialist models with limited cross-dataset generalization. In this work, we revisit REC from the perspective…
Referring ExpressionVisual GroundingVision-Language Grounding as Bidirectional Concept Correspondence
Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the cor…
Referring ExpressionImage SegmentationPhrase GroundingCRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA
Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions. K…
Visual Question AnsweringReferring ExpressionObject RecognitionVisual Grounding