Visual Grounding
4개 벤치마크 · 논문 1,123편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
ViewRefer: Grasp the Multi-view Knowledge for 3D Visual Grounding with GPT and Prototype Guidance
MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding
Towards Visual Grounding: A Survey
Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Papers
Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision
We present our submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge, which ranked first in the large-model division and second in the <=2B division. The task requires a wearable assistant to decide…
Visual GroundingWhere to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding
Recent zero-shot 3D visual grounding methods leverage vision-language models (VLMs) to localize objects in 3D scenes from natural language queries. However, these methods typically rely on heuristic rules to select which…
Natural Language QueriesVisual GroundingOn the Design Fundamentals of Pixel Text Representation Learning
Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders struggle with fixed resolution pretraining, visual shortcut learning, weak v…
Representation LearningVisual GroundingCost-efficient Active Learning for Referring Image Segmentation and Grounding
Collecting natural-language referring expressions along with region annotations, such as masks or boxes, is a major bottleneck in visual grounding (VG), as annotators must write descriptions that distinguish target regio…
Image SegmentationVisual GroundingActive LearningScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions
Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from thirteen public science competitions in…
Visual GroundingVisER: Visual Evidence and Reliance for Object Hallucination Detection in LVLMs
Object hallucination remains a persistent reliability issue in large vision-language models, where generated object mentions may sound plausible but lack visual grounding. Recent training-free detectors use internal sign…
Visual Grounding