paper-with-me

홈 › Papers

Toward Explainable and Fine-Grained 3D Grounding through Referring Textual Phrases

2022-07-05 · Zhihao Yuan, Xu Yan, Zhuo Li, Xuhao Li, Yao Guo, Shuguang Cui, Zhen Li

Recent progress in 3D scene understanding has explored visual grounding (3DVG) to localize a target object through a language description. However, existing methods only consider the dependency between the entire sentence and the target object, ignoring fine-grained relationships between contexts and non-target ones. In this paper, we extend 3DVG to a more fine-grained and interpretable task, called 3D Phrase Aware Grounding (3DPAG). The 3DPAG task aims to localize the target objects in a 3D scene by explicitly identifying all phrase-related objects and then conducting the reasoning according to contextual phrases. To tackle this problem, we manually labeled about 227K phrase-level annotations using a self-developed platform, from 88K sentences of widely used 3DVG datasets, i.e., Nr3D, Sr3D and ScanRefer. By tapping on our datasets, we can extend previous 3DVG methods to the fine-grained phrase-aware scenario. It is achieved through the proposed novel phrase-object alignment optimization and phrase-specific pre-training, boosting conventional 3DVG performance as well. Extensive results confirm significant improvements, i.e., previous state-of-the-art method achieves 3.9%, 3.5% and 4.6% overall accuracy gains on Nr3D, Sr3D and ScanRefer respectively.

📄 PDF Abstract BibTeX arXiv:2207.01821

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectRepresentation LearningScene UnderstandingSentenceVisual Grounding

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

DOGE: Towards Versatile Visual Document Grounding and Referring

2024-11-26 · Yinan Zhou, Yuxin Chen, Haokun Lin, Shuyu Yang 외

In recent years, Multimodal Large Language Models (MLLMs) have increasingly emphasized grounding and referring capabilities to achieve detailed understanding and flexible user interaction. However, in the realm of visual…

document understanding

IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring

2026-05-23 · Xinge Peng, Yiting Lu, Xin Li, Zhibo Chen arxiv

We present IQA-Spider, the first image quality assessment (IQA) framework that unifies reasoning, grounding, and referring into a single LMM-based framework for multi-granularity quality understanding. Existing LMM-based…

Image Quality AssessmentQuestion Answering

Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception

2024-03-05 · CVPR 2024 1 · Junwen He, Yifan Wang, Lijun Wang, Huchuan Lu 외

Multimodal Large Language Model (MLLMs) leverages Large Language Models as a cognitive framework for diverse visual-language tasks. Recent efforts have been made to equip MLLMs with visual perceiving and grounding capabi…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+2

Unveiling Parts Beyond Objects:Towards Finer-Granularity Referring Expression Segmentation

2023-12-13 · Wenxuan Wang, Tongtian Yue, Yisi Zhang, Longteng Guo 외

Referring expression segmentation (RES) aims at segmenting the foreground masks of the entities that match the descriptive natural language expression. Previous datasets and methods for classic RES task heavily rely on t…

DescriptiveObjectReferring ExpressionReferring Expression Segmentation+1

Grounding-IQA: Multimodal Language Grounding Model for Image Quality Assessment

2024-11-26 · Zheng Chen, Xun Zhang, Wenbo Li, Renjing Pei 외

The development of multimodal large language models (MLLMs) enables the evaluation of image quality through natural language descriptions. This advancement allows for more detailed assessments. However, these MLLM-based …

Image Quality AssessmentQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)