paper-with-me

Papers

ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding

2025-01-02 · Austin T. Wang, ZeMing Gong, Angel X. Chang

3D visual grounding (3DVG) involves localizing entities in a 3D scene referred to by natural language text. Such models are useful for embodied AI and scene retrieval applications, which involve searching for objects or patterns using natural language descriptions. While recent works have focused on LLM-based scaling of 3DVG datasets, these datasets do not capture the full range of potential prompts which could be specified in the English language. To ensure that we are scaling up and testing against a useful and representative set of prompts, we propose a framework for linguistically analyzing 3DVG prompts and introduce Visual Grounding with Diverse Language in 3D (ViGiL3D), a diagnostic dataset for evaluating visual grounding methods against a diverse set of language patterns. We evaluate existing open-vocabulary 3DVG methods to demonstrate that these methods are not yet proficient in understanding and identifying the targets of more challenging, out-of-distribution prompts, toward real-world applications.

📄 PDF Abstract BibTeX arXiv:2501.01366

Code (0)

등록된 구현이 없습니다.

Tasks

3D visual groundingDiagnosticVisual Grounding

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Scaling Diverse Language Generation for 3D Visual Grounding

2026-06-18 · Austin T. Wang, Dongchen Yang, Angel X. Chang arxiv

Developing robust models for 3D visual grounding (3DVG), the localization of entities in a 3D scene described in natural language, is important for enabling agents to correspond spatial language with objects in the physi…

Visual Grounding

Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs

2026-06-24 · Xi Xiao, Chen Liu, Chih-Ting Liao, Yunbei Zhang 외 arxiv

Multimodal large language models (MLLMs) extend large language models (LLMs) with visual perception, enabling joint reasoning over images and text. Despite inheriting strong reasoning capabilities from LLMs, they remain …

Visual Grounding

Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation

2025-08-06 · Jinxing Zhou, Yanghao Zhou, Mingfei Han, Tong Wang 외 arxiv

Referring Audio-Visual Segmentation (Ref-AVS) aims to segment target objects in audible videos based on given reference expressions. Prior works typically rely on learning latent embeddings via multimodal fusion to promp…

SignAgent: Agentic LLMs for Linguistically-Grounded Sign Language Annotation and Dataset Curation

2026-03-19 · Oliver Cory, Ozge Mercanoglu Sincan, Richard Bowden arxiv

This paper introduces SignAgent, a novel agentic framework that utilises Large Language Models (LLMs) for scalable, linguistically-grounded Sign Language (SL) annotation and dataset curation. Traditional computational me…

SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes

2026-01-09 · Chuhan Wang, Xintong Li, Jennifer Yuntong Zhang, Junda Wu 외 arxiv

Multimodal large language models often struggle with faithful reasoning in complex visual scenes, where intricate entities and relations require precise visual grounding at each step. This reasoning unfaithfulness freque…

Multimodal ReasoningVisual ReasoningVisual Grounding