paper-with-me

Papers Referring Expression

“Referring Expression” 태그가 달린 논문 425편 · 필터 해제

A-PAIR: A Benchmark and Identity-Consistent Grounding Framework for Air-Ground Cross-View Referring Person Detection

2026-08-28 · Zhoupeng Guo, Xinjie Yao, Yunqi Zhu, Zhihe Fan 외 arxiv

Air-ground cross-view referring person detection is a necessary component in the language-to-perception-to-control chain of collective embodied intelligence, grounding a language command into the same physical target bef…

Referring Expression

When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs

2026-08-25 · Zhengxiang Wang, Owen Rambow arxiv

Visual grounding is typically evaluated as a one-shot mapping from an informative referring expression to a visual target. This formulation misses a central property of real-world reference: initial referring expressions…

Referring ExpressionVisual Grounding

GroupForward: Building Referable 3D Scenes via Instance-Grouped Feed-Forward Gaussian Splatting

2026-08-18 · Qijian Tian, Zimeng Wu, Xuhong Wang, Lizhuang Ma 외 arxiv

Simultaneously reconstructing and understanding 3D environments is essential for embodied agents. Toward this goal, feed-forward semantic 3D Gaussian Splatting (3DGS) efficiently constructs semantic scene representations…

Referring Expression

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding

2026-08-13 · Junyi Hu, Tian Bai, Fengyi Wu, Yian Huang 외 arxiv

Referring Expression Comprehension (REC) is commonly studied under dataset-specific fine-tuning, resulting in specialist models with limited cross-dataset generalization. In this work, we revisit REC from the perspective…

Referring ExpressionVisual Grounding

Vision-Language Grounding as Bidirectional Concept Correspondence

2026-08-08 · Jieyu Zhang, Ziqi Gao, Luke Zettlemoyer, Ranjay Krishna hf

Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the cor…

Referring ExpressionImage SegmentationPhrase Grounding

CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA

2026-07-23 · Hanseok Oh, Parishad BehnamGhader, Benno Krojer, Hyunji Lee 외 arxiv

Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions. K…

Visual Question AnsweringReferring ExpressionObject RecognitionVisual Grounding

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos

2026-07-22 · Penglei Sun, Yehua Huang, Zhuoli Tao, Xiang Li 외 arxiv

Language-guided aerial perception aims to understand user-specified tiny targets in complex unmanned aerial vehicle (UAV) scenes. In real UAV deployment, the UAV must respond while it flies, so such perception runs in an…

Referring Expression

UniRef-UAV: A Multimodal Benchmark for Universal Referring in UAV Imagery

2026-07-09 · Haibin Tian, Huichao Xie, Xuelin Qian, Ruitao Lu 외 arxiv

Unmanned aerial vehicles (UAVs) increasingly rely on visual grounding capabilities to localize task-relevant targets from diverse instructions in complex aerial scenes. Existing referring expression comprehension (REC) b…

Referring ExpressionVisual Grounding

Learning Structured Visual Compositional Representations for Weakly Supervised Referring Expression Comprehension

2026-07-06 · Lian Xu, Mohammed Bennamoun, Farid Boussaid, Hamid Laga 외 arxiv

Referring expression comprehension (REC) aims to localize the object in an image described by natural language. In Weakly supervised REC (WREC), existing approaches primarily operate on anchor-level visual representation…

Referring Expression

LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension

2026-07-02 · Shunya Kato, Taiki Miyanishi, Shuhei Kurita, Mahiro Ukai 외 arxiv

Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects. In this context, Video Referring Expression Comprehen…

Referring Expression

Disentangling Pictorial Cue Understanding from Language Bias in VLMs via Depth Ordering Task

2026-07-01 · Yiqian Liu, Iuliia Kotseruba, John K. Tsotsos arxiv

In this paper, we study depth perception of vision-language models (VLMs) to isolate the effects of pictorial depth cues and disentangle vision and language influences on model performance. To this end, we combine depth-…

Referring ExpressionImage Generation

Open-Vocabulary and Referring Segmentation for 3D Gaussians Using 2D Detectors

2026-06-29 · Jameel Hassan, Yasiru Ranasinghe, Vishal Patel arxiv

3D Gaussian Splatting (3DGS) has emerged at the forefront of 3D scene reconstruction. Extending 3DGS with language-driven, open-vocabulary understanding has gained significant attention for real-world applications such a…

Referring ExpressionSpatial Reasoning

From Propositional to Perceptual Asymmetry: Extending Frictive Policy Optimization to Asymmetric Partial Information Dialogue

2026-06-29 · Yifan Zhu, Kyeongmin Rim, James Pustejovsky arxiv

Frictive Policy Optimization (FPO; Pustejovsky et al., 2025) treats friction in collaborative dialogue -- misalignment, misunderstanding, repair -- as an epistemic signal essential to common-ground construction, rather t…

Referring Expression

SATURN: Symbolic Spatial Reasoning for Multi-Perspective Grounding

2026-06-21 · Danial Kamali, Tanawan Premsri, Shreya Rajpal, Amir Zadeh 외 arxiv

Vision-Language Models (VLMs) remain unreliable when spatial reasoning requires composing relations whose meanings depend on frames of reference. Existing neuro-symbolic methods make reasoning more explicit, but often de…

Referring ExpressionSpatial Reasoning

Training-Free Open-Vocabulary Visual Grounding for Remote Sensing Images and Videos

2026-06-15 · Ke Li, Di Wang, Yongshan Zhu, Ting Wang 외 arxiv

Remote sensing visual grounding (RSVG) aims to localize a referred target in a remote sensing image or video according to a natural language expression. Existing RSVG methods usually rely on task-specific manual annotati…

Referring ExpressionVisual GroundingVideo Grounding

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs

2026-06-02 · Eshika Khandelwal, Jingjing Pan, Mingfang Zhang, Quan Kong 외 arxiv

Multimodal large language models (MLLMs) are predominantly evaluated on free-form vision-language tasks such as visual question answering, captioning, and summarization. However, their practical use is rapidly expanding …

Visual Question AnsweringReferring ExpressionObject Detection

Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker

2026-05-25 · Zongjian Wu, Lei Zhang arxiv

Referring expression comprehension (REC) aims to localize a target object within an image based on a given expression. Although recent advances in vision-language models have led to substantial improvements in REC tasks,…

Referring Expression

AgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding

2026-05-21 · Haocheng Li, Juepeng Zheng, Zenghao Yang, Kaiqi Du 외 arxiv

Visual grounding, the task of localizing objects described by natural-language expressions, is a foundational capability for agricultural AI systems, enabling applications such as selective weeding, disease monitoring, a…

Referring ExpressionVisual Grounding

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding

2026-05-18 · Boyuan Sun, Bowen Yin, Yuanming Li, Xihan Wei 외 arxiv

We present SWIM (See What I Mean), a novel training strategy that aligns vision and language representations to enable fine-grained object understanding solely from textual prompts. Unlike existing approaches that requir…

Referring Expression

REC-RL: Referring expression counting via Gaussian and range-based reward optimization

2026-05-15 · Hui Liu, Yunlai Teng, Kunlong Bai, Pengfei Qi 외 arxiv

Referring expression counting (REC) is an intention-driven task that requires context-aware visual reasoning. While recent vision-language models incorporate language for visual understanding, most existing REC methods r…

Reinforcement LearningReferring ExpressionVisual Reasoning
1–20 / 425 다음 →