paper-with-me

홈 › Papers

GENNAV: Polygon Mask Generation for Generalized Referring Navigable Regions

2025-08-28 · Kei Katsumata, Yui Iioka, Naoki Hosomi, Teruhisa Misu, Kentaro Yamada, Komei Sugiura arxiv

We focus on the task of identifying the location of target regions from a natural language instruction and a front camera image captured by a mobility. This task is challenging because it requires both existence prediction and segmentation, particularly for stuff-type target regions with ambiguous boundaries. Existing methods often underperform in handling stuff-type target regions, in addition to absent or multiple targets. To overcome these limitations, we propose GENNAV, which predicts target existence and generates segmentation masks for multiple stuff-type target regions. To evaluate GENNAV, we constructed a novel benchmark called GRiN-Drive, which includes three distinct types of samples: no-target, single-target, and multi-target. GENNAV achieved superior performance over baseline methods on standard evaluation metrics. Furthermore, we conducted real-world experiments with four automobiles operated in five geographically distinct urban areas to validate its zero-shot transfer performance. In these experiments, GENNAV outperformed baseline methods and demonstrated its robustness across diverse real-world environments. The project page is available at https://gennav.vercel.app/.

📄 PDF Abstract BibTeX arXiv:2508.21102

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PolyFormer: Referring Image Segmentation as Sequential Polygon Generation

2023-02-14 · CVPR 2023 1 · Jiang Liu, Hui Ding, Zhaowei Cai, Yuting Zhang 외

In this work, instead of directly predicting the pixel-level segmentation masks, the problem of referring image segmentation is formulated as sequential polygon generation, and the predicted polygons can be later convert…

DecoderImage SegmentationQuantizationReferring Expression Comprehension+5

Moondream Segmentation: From Words to Masks

2026-04-03 · Ethan Reid arxiv

We present Moondream Segmentation, a referring image segmentation extension of Moondream 3, a vision-language model. Given an image and a referring expression, the model autoregressively decodes a vector path and iterati…

Reinforcement LearningReferring ExpressionImage Segmentation

Object Segmentation from Open-Vocabulary Manipulation Instructions Based on Optimal Transport Polygon Matching with Multimodal Foundation Models

2024-07-01 · Takayuki Nishimura, Katsuyuki Kuyo, Motonari Kambara, Komei Sugiura

We consider the task of generating segmentation masks for the target object from an object manipulation instruction, which allows users to give open vocabulary instructions to domestic service robots. Conventional segmen…

SegmentationSemantic Segmentation

Meteorologists and Students: A resource for language grounding of geographical descriptors

2018-09-07 · WS 2018 11 · Alejandro Ramos-Soto, Ehud Reiter, Kees Van Deemter, Jose M. Alonso 외

We present a data resource which can be useful for research purposes on language grounding tasks in the context of geographical referring expression generation. The resource is composed of two data sets that encompass 25…

Referring ExpressionReferring expression generation

Refer to Anything with Vision-Language Prompts

2025-06-05 · Shengcao Cao, Zijun Wei, Jason Kuen, Kangning Liu 외

Recent image segmentation models have advanced to segment images into high-quality masks for visual entities, and yet they cannot provide comprehensive semantic understanding for complex queries based on both language an…

BenchmarkingGeneralized Referring Expression SegmentationImage SegmentationReferring Expression+3