paper-with-me

Referring Expression

1개 벤치마크 · 논문 425편 · 이 태스크의 논문 보기 →

Benchmarks

SQA3D

결과 1개

Most implemented

Towards Visual Grounding: A Survey

2024-12-28 · 구현 4개

Papers

A-PAIR: A Benchmark and Identity-Consistent Grounding Framework for Air-Ground Cross-View Referring Person Detection

2026-08-28 · Zhoupeng Guo, Xinjie Yao, Yunqi Zhu, Zhihe Fan 외 arxiv

Air-ground cross-view referring person detection is a necessary component in the language-to-perception-to-control chain of collective embodied intelligence, grounding a language command into the same physical target bef…

Referring Expression

When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs

2026-08-25 · Zhengxiang Wang, Owen Rambow arxiv

Visual grounding is typically evaluated as a one-shot mapping from an informative referring expression to a visual target. This formulation misses a central property of real-world reference: initial referring expressions…

Referring ExpressionVisual Grounding

GroupForward: Building Referable 3D Scenes via Instance-Grouped Feed-Forward Gaussian Splatting

2026-08-18 · Qijian Tian, Zimeng Wu, Xuhong Wang, Lizhuang Ma 외 arxiv

Simultaneously reconstructing and understanding 3D environments is essential for embodied agents. Toward this goal, feed-forward semantic 3D Gaussian Splatting (3DGS) efficiently constructs semantic scene representations…

Referring Expression

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding

2026-08-13 · Junyi Hu, Tian Bai, Fengyi Wu, Yian Huang 외 arxiv

Referring Expression Comprehension (REC) is commonly studied under dataset-specific fine-tuning, resulting in specialist models with limited cross-dataset generalization. In this work, we revisit REC from the perspective…

Referring ExpressionVisual Grounding

Vision-Language Grounding as Bidirectional Concept Correspondence

2026-08-08 · Jieyu Zhang, Ziqi Gao, Luke Zettlemoyer, Ranjay Krishna hf

Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the cor…

Referring ExpressionImage SegmentationPhrase Grounding

CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA

2026-07-23 · Hanseok Oh, Parishad BehnamGhader, Benno Krojer, Hyunji Lee 외 arxiv

Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions. K…

Visual Question AnsweringReferring ExpressionObject RecognitionVisual Grounding

전체 425편 보기 →