Papers Visual Grounding
“Visual Grounding” 태그가 달린 논문 1,123편 · 필터 해제
Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision
We present our submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge, which ranked first in the large-model division and second in the <=2B division. The task requires a wearable assistant to decide…
Visual GroundingWhere to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding
Recent zero-shot 3D visual grounding methods leverage vision-language models (VLMs) to localize objects in 3D scenes from natural language queries. However, these methods typically rely on heuristic rules to select which…
Natural Language QueriesVisual GroundingOn the Design Fundamentals of Pixel Text Representation Learning
Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders struggle with fixed resolution pretraining, visual shortcut learning, weak v…
Representation LearningVisual GroundingCost-efficient Active Learning for Referring Image Segmentation and Grounding
Collecting natural-language referring expressions along with region annotations, such as masks or boxes, is a major bottleneck in visual grounding (VG), as annotators must write descriptions that distinguish target regio…
Image SegmentationVisual GroundingActive LearningScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions
Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from thirteen public science competitions in…
Visual GroundingVisER: Visual Evidence and Reliance for Object Hallucination Detection in LVLMs
Object hallucination remains a persistent reliability issue in large vision-language models, where generated object mentions may sound plausible but lack visual grounding. Recent training-free detectors use internal sign…
Visual GroundingSeqAlign3DVG: A Sequence-Aligned Benchmark and Voxel Reasoning Framework for 3D Visual Grounding
Image-based 3D visual grounding is critical for embodied agents, yet existing benchmarks suffer from loose text-observation alignment and neglect temporal ordering. We introduce SeqAlign3DVG, a novel benchmark dedicated …
Visual GroundingPoint CloudsSemantic-Spatial Discriminability Enhancement for Generalized Visual Grounding
Generalized Visual Grounding (GVG) task aims to localize targets in an image based on referring expressions, extends the classical visual grounding paradigm by integrating multi-target and non-target scenarios. Previous …
Visual GroundingReactivating Test-Time Scaling for Plane Geometry Problem Solving
Plane geometry problem (PGP) solving has become a critical benchmark for multimodal reasoning because it requires accurate visual perception and precise multi-step symbolic deduction. Although test-time scaling (TTS) has…
Mathematical ReasoningMultimodal ReasoningVisual GroundingGRAFT: Grounded and Efficient Online Reinforcement Adaptation for Fine-Grained Robot Manipulation
Pretrained vision-language-action (VLA) policies provide strong priors for robot manipulation, yet adapting them online to fine-grained biomedical tasks remains challenging. Task success often hinges on subtle, view-depe…
Robot ManipulationVisual GroundingMapping Written Words to Spoken Words in a Different Language Using Only Visual Grounding
In many low-resource settings, even just eliciting speech for data collection is difficult. One promising approach has been to ask speakers to describe images. But how do we build models from such visually grounded speec…
Visual GroundingImage CaptioningKeyword SpottingWhen Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs
Visual grounding is typically evaluated as a one-shot mapping from an informative referring expression to a visual target. This formulation misses a central property of real-world reference: initial referring expressions…
Referring ExpressionVisual GroundingFOVEA: Focused On-Demand Visual Evidence Adaptation for Cache-Friendly Multimodal Speculative Decoding
Multimodal speculative decoding accelerates vision-language models by allowing a lightweight draft model to propose candidate tokens for parallel verification by a larger target model. Existing methods typically conditio…
multimodal generationVisual GroundingIs Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds
Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an interaction between visual grounding and downstream reasoning. What remains less cle…
Multimodal ReasoningSpatial ReasoningVisual GroundingObject DetectionTLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where product facts are distributed across speech, video frames, product images, overlaid text, and user queries. We pres…
Speech RecognitionVisual GroundingGrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery
Visual grounding in Unmanned Aerial Vehicle (UAV) imagery aims to localize a target object in complex bird's-eye-view scenes according to a natural language description. However, the abundance of small, densely distribut…
Spatial ReasoningVisual GroundingNeurosymbolic Embodied Agents
Language and vision-language models generate plausible embodied plans but do not guarantee executability, as their outputs can violate environment dynamics or act on incorrectly grounded entities. We present a neurosymbo…
Visual GroundingDefake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection
The rapid progress of image generation models calls for AI-generated image (AIGI) detectors that are not only accurate but also explainable and reliable. While MLLM-based detectors can provide natural language explanatio…
Reinforcement LearningVisual GroundingImage GenerationGaussianDWM++: Language-Grounded 3D Gaussian Driving World Model for Unified Scene Understanding, Editing, and Multi-Modal Generation
Driving World Models (DWMs) have recently advanced rapidly with generative models, yet most existing methods mainly focus on conditional scene generation and lack explicit 3D scene understanding, language-grounded reason…
Scene UnderstandingScene GenerationVisual GroundingSCVIB: Editable State-Conditioned Visual Instance Binding forMulti-Turn Personalized Localization
We introduce editable state-conditioned visual instance binding, a multi-turn localization setting in which several support-defined instances are introduced across turns and protocol-defined state events determine the fi…
Visual Grounding