paper-with-me

Papers Visual Grounding

“Visual Grounding” 태그가 달린 논문 1,123편 · 필터 해제

Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision

2026-09-10 · Logesh Kumar Umapathi hf

We present our submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge, which ranked first in the large-model division and second in the <=2B division. The task requires a wearable assistant to decide…

Visual Grounding

Where to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding

2026-09-04 · Tsung-Chih Chiang, Hsuan-Kung Yang, Jou-Min Liu, Ting-Ru Liu 외 arxiv

Recent zero-shot 3D visual grounding methods leverage vision-language models (VLMs) to localize objects in 3D scenes from natural language queries. However, these methods typically rely on heuristic rules to select which…

Natural Language QueriesVisual Grounding

On the Design Fundamentals of Pixel Text Representation Learning

2026-09-01 · Chaohao Yuan, Ruifeng Yuan, Zhuoxu Huang, Yu Rong 외 hf

Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders struggle with fixed resolution pretraining, visual shortcut learning, weak v…

Representation LearningVisual Grounding

Cost-efficient Active Learning for Referring Image Segmentation and Grounding

2026-08-31 · Junbeom Hong, Seonghoon Yu, Hyung Rok Jung, Sundong Kim 외 arxiv

Collecting natural-language referring expressions along with region annotations, such as masks or boxes, is a major bottleneck in visual grounding (VG), as annotators must write descriptions that distinguish target regio…

Image SegmentationVisual GroundingActive Learning

ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions

2026-08-31 · Guangxiang Zhao, Qilong Shi, Xusen Xiao, Wenpu Liu 외 arxiv

Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from thirteen public science competitions in…

Visual Grounding

VisER: Visual Evidence and Reliance for Object Hallucination Detection in LVLMs

2026-08-31 · Afsaneh Hasanebrahimi, Hanxun Huang, Christopher Leckie, Sarah Erfani arxiv

Object hallucination remains a persistent reliability issue in large vision-language models, where generated object mentions may sound plausible but lack visual grounding. Recent training-free detectors use internal sign…

Visual Grounding

SeqAlign3DVG: A Sequence-Aligned Benchmark and Voxel Reasoning Framework for 3D Visual Grounding

2026-08-31 · Yi Zhang, Yi Wang, Yueting Wu, Kaiyue Yang 외 arxiv

Image-based 3D visual grounding is critical for embodied agents, yet existing benchmarks suffer from loose text-observation alignment and neglect temporal ordering. We introduce SeqAlign3DVG, a novel benchmark dedicated …

Visual GroundingPoint Clouds

Semantic-Spatial Discriminability Enhancement for Generalized Visual Grounding

2026-08-31 · Kaiyan Lei, Xu-Yao Zhang arxiv

Generalized Visual Grounding (GVG) task aims to localize targets in an image based on referring expressions, extends the classical visual grounding paradigm by integrating multi-target and non-target scenarios. Previous …

Visual Grounding

Reactivating Test-Time Scaling for Plane Geometry Problem Solving

2026-08-31 · Xiaoqiang Kang, Shengen Wu, Maizhen Ning, Xiaobo Jin 외 arxiv

Plane geometry problem (PGP) solving has become a critical benchmark for multimodal reasoning because it requires accurate visual perception and precise multi-step symbolic deduction. Although test-time scaling (TTS) has…

Mathematical ReasoningMultimodal ReasoningVisual Grounding

GRAFT: Grounded and Efficient Online Reinforcement Adaptation for Fine-Grained Robot Manipulation

2026-08-27 · Yibo Qiu, Haoliang Ye, Shu'ang Sun, Zan Huang 외 arxiv

Pretrained vision-language-action (VLA) policies provide strong priors for robot manipulation, yet adapting them online to fine-grained biomedical tasks remains challenging. Task success often hinges on subtle, view-depe…

Robot ManipulationVisual Grounding

Mapping Written Words to Spoken Words in a Different Language Using Only Visual Grounding

2026-08-27 · Gabriel Pirlogeanu, Dan Oneata, Horia Cucu, Herman Kamper arxiv

In many low-resource settings, even just eliciting speech for data collection is difficult. One promising approach has been to ask speakers to describe images. But how do we build models from such visually grounded speec…

Visual GroundingImage CaptioningKeyword Spotting

When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs

2026-08-25 · Zhengxiang Wang, Owen Rambow arxiv

Visual grounding is typically evaluated as a one-shot mapping from an informative referring expression to a visual target. This formulation misses a central property of real-world reference: initial referring expressions…

Referring ExpressionVisual Grounding

FOVEA: Focused On-Demand Visual Evidence Adaptation for Cache-Friendly Multimodal Speculative Decoding

2026-08-24 · Hengjie Zhu, Dayan Wu, Zihao Zhang, Xinze Liu 외 arxiv

Multimodal speculative decoding accelerates vision-language models by allowing a lightweight draft model to propose candidate tokens for parallel verification by a larger target model. Existing methods typically conditio…

multimodal generationVisual Grounding

Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds

2026-08-21 · Lars Benedikt Kaesberg, Tianyu Yang, Florian Valentin Wunderlich, Terry Ruas 외 arxiv

Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an interaction between visual grounding and downstream reasoning. What remains less cle…

Multimodal ReasoningSpatial ReasoningVisual GroundingObject Detection

TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

2026-08-21 · Yibo Hu, Yu Qian, Mao Gu, Yingfan Tao 외 arxiv

E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where product facts are distributed across speech, video frames, product images, overlaid text, and user queries. We pres…

Speech RecognitionVisual Grounding

GrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery

2026-08-19 · Chaowei Wang, Yan Di, Jingjun Sun, Baozhe Liu 외 arxiv

Visual grounding in Unmanned Aerial Vehicle (UAV) imagery aims to localize a target object in complex bird's-eye-view scenes according to a natural language description. However, the abundance of small, densely distribut…

Spatial ReasoningVisual Grounding

Neurosymbolic Embodied Agents

2026-08-17 · Mohammad Albinhassan, Yuming Feng, Alessandra Russo, Pranava Madhyastha arxiv

Language and vision-language models generate plausible embodied plans but do not guarantee executability, as their outputs can violate environment dynamics or act on incorrectly grounded entities. We present a neurosymbo…

Visual Grounding

Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection

2026-08-17 · Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan 외 arxiv

The rapid progress of image generation models calls for AI-generated image (AIGI) detectors that are not only accurate but also explainable and reliable. While MLLM-based detectors can provide natural language explanatio…

Reinforcement LearningVisual GroundingImage Generation

GaussianDWM++: Language-Grounded 3D Gaussian Driving World Model for Unified Scene Understanding, Editing, and Multi-Modal Generation

2026-08-17 · Tianchen Deng, Xuefeng Chen, Shuang Wu, Qu Chen 외 arxiv

Driving World Models (DWMs) have recently advanced rapidly with generative models, yet most existing methods mainly focus on conditional scene generation and lack explicit 3D scene understanding, language-grounded reason…

Scene UnderstandingScene GenerationVisual Grounding

SCVIB: Editable State-Conditioned Visual Instance Binding forMulti-Turn Personalized Localization

2026-08-14 · Xiongtai Yang, Ziyan He, Tao Wang arxiv

We introduce editable state-conditioned visual instance binding, a multi-turn localization setting in which several support-defined instances are introduced across turns and protocol-defined state events determine the fi…

Visual Grounding
1–20 / 1,123 다음 →