paper-with-me

Visual Grounding

4개 벤치마크 · 논문 1,123편 · 이 태스크의 논문 보기 →

Benchmarks

RefCOCO+ testA

결과 7개

RefCOCO+ test B

결과 6개

RefCOCO+ val

결과 6개

RefCOCO testA

결과 1개

Most implemented

Towards Visual Grounding: A Survey

2024-12-28 · 구현 4개

Papers

Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision

2026-09-10 · Logesh Kumar Umapathi hf

We present our submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge, which ranked first in the large-model division and second in the <=2B division. The task requires a wearable assistant to decide…

Visual Grounding

Where to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding

2026-09-04 · Tsung-Chih Chiang, Hsuan-Kung Yang, Jou-Min Liu, Ting-Ru Liu 외 arxiv

Recent zero-shot 3D visual grounding methods leverage vision-language models (VLMs) to localize objects in 3D scenes from natural language queries. However, these methods typically rely on heuristic rules to select which…

Natural Language QueriesVisual Grounding

On the Design Fundamentals of Pixel Text Representation Learning

2026-09-01 · Chaohao Yuan, Ruifeng Yuan, Zhuoxu Huang, Yu Rong 외 hf

Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders struggle with fixed resolution pretraining, visual shortcut learning, weak v…

Representation LearningVisual Grounding

Cost-efficient Active Learning for Referring Image Segmentation and Grounding

2026-08-31 · Junbeom Hong, Seonghoon Yu, Hyung Rok Jung, Sundong Kim 외 arxiv

Collecting natural-language referring expressions along with region annotations, such as masks or boxes, is a major bottleneck in visual grounding (VG), as annotators must write descriptions that distinguish target regio…

Image SegmentationVisual GroundingActive Learning

ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions

2026-08-31 · Guangxiang Zhao, Qilong Shi, Xusen Xiao, Wenpu Liu 외 arxiv

Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from thirteen public science competitions in…

Visual Grounding

VisER: Visual Evidence and Reliance for Object Hallucination Detection in LVLMs

2026-08-31 · Afsaneh Hasanebrahimi, Hanxun Huang, Christopher Leckie, Sarah Erfani arxiv

Object hallucination remains a persistent reliability issue in large vision-language models, where generated object mentions may sound plausible but lack visual grounding. Recent training-free detectors use internal sign…

Visual Grounding

전체 1,123편 보기 →