Papers Referring Expression
“Referring Expression” 태그가 달린 논문 425편 · 필터 해제
A-PAIR: A Benchmark and Identity-Consistent Grounding Framework for Air-Ground Cross-View Referring Person Detection
Air-ground cross-view referring person detection is a necessary component in the language-to-perception-to-control chain of collective embodied intelligence, grounding a language command into the same physical target bef…
Referring ExpressionWhen Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs
Visual grounding is typically evaluated as a one-shot mapping from an informative referring expression to a visual target. This formulation misses a central property of real-world reference: initial referring expressions…
Referring ExpressionVisual GroundingGroupForward: Building Referable 3D Scenes via Instance-Grouped Feed-Forward Gaussian Splatting
Simultaneously reconstructing and understanding 3D environments is essential for embodied agents. Toward this goal, feed-forward semantic 3D Gaussian Splatting (3DGS) efficiently constructs semantic scene representations…
Referring ExpressionScaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding
Referring Expression Comprehension (REC) is commonly studied under dataset-specific fine-tuning, resulting in specialist models with limited cross-dataset generalization. In this work, we revisit REC from the perspective…
Referring ExpressionVisual GroundingVision-Language Grounding as Bidirectional Concept Correspondence
Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the cor…
Referring ExpressionImage SegmentationPhrase GroundingCRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA
Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions. K…
Visual Question AnsweringReferring ExpressionObject RecognitionVisual GroundingMemory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos
Language-guided aerial perception aims to understand user-specified tiny targets in complex unmanned aerial vehicle (UAV) scenes. In real UAV deployment, the UAV must respond while it flies, so such perception runs in an…
Referring ExpressionUniRef-UAV: A Multimodal Benchmark for Universal Referring in UAV Imagery
Unmanned aerial vehicles (UAVs) increasingly rely on visual grounding capabilities to localize task-relevant targets from diverse instructions in complex aerial scenes. Existing referring expression comprehension (REC) b…
Referring ExpressionVisual GroundingLearning Structured Visual Compositional Representations for Weakly Supervised Referring Expression Comprehension
Referring expression comprehension (REC) aims to localize the object in an image described by natural language. In Weakly supervised REC (WREC), existing approaches primarily operate on anchor-level visual representation…
Referring ExpressionLongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension
Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects. In this context, Video Referring Expression Comprehen…
Referring ExpressionDisentangling Pictorial Cue Understanding from Language Bias in VLMs via Depth Ordering Task
In this paper, we study depth perception of vision-language models (VLMs) to isolate the effects of pictorial depth cues and disentangle vision and language influences on model performance. To this end, we combine depth-…
Referring ExpressionImage GenerationOpen-Vocabulary and Referring Segmentation for 3D Gaussians Using 2D Detectors
3D Gaussian Splatting (3DGS) has emerged at the forefront of 3D scene reconstruction. Extending 3DGS with language-driven, open-vocabulary understanding has gained significant attention for real-world applications such a…
Referring ExpressionSpatial ReasoningFrom Propositional to Perceptual Asymmetry: Extending Frictive Policy Optimization to Asymmetric Partial Information Dialogue
Frictive Policy Optimization (FPO; Pustejovsky et al., 2025) treats friction in collaborative dialogue -- misalignment, misunderstanding, repair -- as an epistemic signal essential to common-ground construction, rather t…
Referring ExpressionSATURN: Symbolic Spatial Reasoning for Multi-Perspective Grounding
Vision-Language Models (VLMs) remain unreliable when spatial reasoning requires composing relations whose meanings depend on frames of reference. Existing neuro-symbolic methods make reasoning more explicit, but often de…
Referring ExpressionSpatial ReasoningTraining-Free Open-Vocabulary Visual Grounding for Remote Sensing Images and Videos
Remote sensing visual grounding (RSVG) aims to localize a referred target in a remote sensing image or video according to a natural language expression. Existing RSVG methods usually rely on task-specific manual annotati…
Referring ExpressionVisual GroundingVideo GroundingFindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs
Multimodal large language models (MLLMs) are predominantly evaluated on free-form vision-language tasks such as visual question answering, captioning, and summarization. However, their practical use is rapidly expanding …
Visual Question AnsweringReferring ExpressionObject DetectionTowards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker
Referring expression comprehension (REC) aims to localize a target object within an image based on a given expression. Although recent advances in vision-language models have led to substantial improvements in REC tasks,…
Referring ExpressionAgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding
Visual grounding, the task of localizing objects described by natural-language expressions, is a foundational capability for agricultural AI systems, enabling applications such as selective weeding, disease monitoring, a…
Referring ExpressionVisual GroundingSee What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
We present SWIM (See What I Mean), a novel training strategy that aligns vision and language representations to enable fine-grained object understanding solely from textual prompts. Unlike existing approaches that requir…
Referring ExpressionREC-RL: Referring expression counting via Gaussian and range-based reward optimization
Referring expression counting (REC) is an intention-driven task that requires context-aware visual reasoning. While recent vision-language models incorporate language for visual understanding, most existing REC methods r…
Reinforcement LearningReferring ExpressionVisual Reasoning