paper-with-me

홈 › Papers

Omni-Referring Image Segmentation

2025-12-07 · Qiancheng Zheng, Yunhang Shen, Gen Luo, Baiyang Song, Xing Sun, Xiaoshuai Sun, Yiyi Zhou, Rongrong Ji arxiv

In this paper, we propose a novel task termed Omni-Referring Image Segmentation (OmniRIS) towards highly generalized image segmentation. Compared with existing unimodally conditioned segmentation tasks, such as RIS and visual RIS, OmniRIS supports the input of text instructions and reference images with masks, boxes or scribbles as omni-prompts. This property makes it can well exploit the intrinsic merits of both text and visual modalities, i.e., granular attribute referring and uncommon object grounding, respectively. Besides, OmniRIS can also handle various segmentation settings, such as one v.s. many and many v.s. many, further facilitating its practical use. To promote the research of OmniRIS, we also rigorously design and construct a large dataset termed OmniRef, which consists of 186,939 omni-prompts for 30,956 images, and establish a comprehensive evaluation system. Moreover, a strong and general baseline termed OmniSegNet is also proposed to tackle the key challenges of OmniRIS, such as omni-prompt encoding. The extensive experiments not only validate the capability of OmniSegNet in following omni-modal instructions, but also show the superiority of OmniRIS for highly generalized image segmentation.

📄 PDF Abstract BibTeX arXiv:2512.06862

Code (0)

등록된 구현이 없습니다.

Tasks

Image Segmentation

Similar Papers 제목 키워드 기반

Towards Omni-supervised Referring Expression Segmentation

2023-11-01 · Minglang Huang, Yiyi Zhou, Gen Luo, Guannan Jiang 외

Referring Expression Segmentation (RES) is an emerging task in computer vision, which segments the target instances in images based on text descriptions. However, its development is plagued by the expensive segmentation …

Referring ExpressionReferring Expression SegmentationSegmentation

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation

2025-07-30 · Kaining Ying, Henghui Ding, Guangquan Jie, Yu-Gang Jiang arxiv

Referring audio-visual segmentation (RAVS) has recently seen significant advancements, yet challenges remain in integrating multimodal information and deeply understanding and reasoning about audiovisual content. To exte…

Multimodal Reasoning

Refer to Anything with Vision-Language Prompts

2025-06-05 · Shengcao Cao, Zijun Wei, Jason Kuen, Kangning Liu 외

Recent image segmentation models have advanced to segment images into high-quality masks for visual entities, and yet they cannot provide comprehensive semantic understanding for complex queries based on both language an…

BenchmarkingGeneralized Referring Expression SegmentationImage SegmentationReferring Expression+3

ORMOT: A Dataset and Framework for Omnidirectional Referring Multi-Object Tracking

2026-03-05 · Sijia Chen, Zihan Zhou, Yanqiu Yu, En Yu 외 arxiv

Multi-Object Tracking (MOT) is a fundamental task in computer vision, aiming to track targets across video frames. Existing MOT methods perform well in general visual scenes, but face significant challenges and limitatio…

Multi-Object Tracking

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

2025-05-26 · Hao Zhong, Muzhi Zhu, Zongze Du, Zheng Huang 외

Long-horizon video-audio reasoning and fine-grained pixel understanding impose conflicting requirements on omnimodal models: dense temporal coverage demands many low-resolution frames, whereas precise grounding calls for…

Domain GeneralizationHallucinationReasoning Video Object SegmentationReferring Audio-Visual Segmentation+4