paper-with-me

홈 › Papers

Unveiling Parts Beyond Objects:Towards Finer-Granularity Referring Expression Segmentation

2023-12-13 · Wenxuan Wang, Tongtian Yue, Yisi Zhang, Longteng Guo, Xingjian He, Xinlong Wang, Jing Liu

Referring expression segmentation (RES) aims at segmenting the foreground masks of the entities that match the descriptive natural language expression. Previous datasets and methods for classic RES task heavily rely on the prior assumption that one expression must refer to object-level targets. In this paper, we take a step further to finer-grained part-level RES task. To promote the object-level RES task towards finer-grained vision-language understanding, we put forward a new multi-granularity referring expression segmentation (MRES) task and construct an evaluation benchmark called RefCOCOm by manual annotations. By employing our automatic model-assisted data engine, we build the largest visual grounding dataset namely MRES-32M, which comprises over 32.2M high-quality masks and captions on the provided 1M images. Besides, a simple yet strong model named UniRES is designed to accomplish the unified object-level and part-level grounding task. Extensive experiments on our RefCOCOm for MRES and three datasets (i.e., RefCOCO(+/g) for classic RES task demonstrate the superiority of our method over previous state-of-the-art methods. To foster future research into fine-grained visual grounding, our benchmark RefCOCOm, the MRES-32M dataset and model UniRES will be publicly available at https://github.com/Rubics-Xuan/MRES

📄 PDF Abstract BibTeX arXiv:2312.08007

Code (1)

rubics-xuan/mres 공식 구현 pytorch

Tasks

DescriptiveObjectReferring ExpressionReferring Expression SegmentationVisual Grounding

Similar Papers 제목 키워드 기반

Unveiling Parts Beyond Objects: Towards Finer-Granularity Referring Expression Segmentation

2024-01-01 · CVPR 2024 1 · Wenxuan Wang, Tongtian Yue, Yisi Zhang, Longteng Guo 외

Referring expression segmentation (RES) aims at segmenting the foreground masks of the entities that match the descriptive natural language expression. Previous datasets and methods for classic RES task heavily rely …

DescriptiveObjectReferring ExpressionReferring Expression Segmentation+1

Search3D: Hierarchical Open-Vocabulary 3D Segmentation

2024-09-27 · Ayca Takmaz, Alexandros Delitzas, Robert W. Sumner, Francis Engelmann 외

Open-vocabulary 3D segmentation enables exploration of 3D spaces using free-form text descriptions. Existing methods for open-vocabulary 3D instance segmentation primarily focus on identifying object-level instances but …

3D Instance Segmentation3D Part SegmentationInstance SegmentationObject+3

A Computer Vision Aided Beamforming Scheme with EM Exposure Control in Outdoor LOS Scenarios

2020-06-14 · Tianqi Xiang, Huiwen Li, Boren Guo, Xin Zhang

Without any radiation control measures, a large-scale mmWave antenna array at close range may lead to a large amount of electromagnetic exposure of human. In this paper, with the aid of pose detection in computer vision,…

Management

FineRMoE: Dimension Expansion for Finer-Grained Expert with Its Upcycling Approach

2026-03-09 · Ning Liao, Xiaoxing Wang, Xiaohan Qin, Junchi Yan arxiv

As revealed by the scaling law of fine-grained MoE, model performance ceases to be improved once the granularity of the intermediate dimension exceeds the optimal threshold, limiting further gains from single-dimension f…

Tweaking UD Annotations to Investigate the Placement of Determiners, Quantifiers and Numerals in the Noun Phrase

2022-07-01 · NAACL (SIGTYP) 2022 7 · Luigi Talamo

We describe a methodology to extract with finer accuracy word order patterns from texts automatically annotated with Universal Dependency (UD) trained parsers. We use the methodology to quantify the word order entropy of…

LEMMA