paper-with-me

Papers

Unified Open-World Segmentation with Multi-Modal Prompts

2025-10-12 · Yang Liu, Yufei Yin, Chenchen Jing, Muzhi Zhu, Hao Chen, Yuling Xi, Bo Feng, Hao Wang, Shiyu Li, Chunhua Shen arxiv

In this work, we present COSINE, a unified open-world segmentation model that consolidates open-vocabulary segmentation and in-context segmentation with multi-modal prompts (e.g., text and image). COSINE exploits foundation models to extract representations for an input image and corresponding multi-modal prompts, and a SegDecoder to align these representations, model their interaction, and obtain masks specified by input prompts across different granularities. In this way, COSINE overcomes architectural discrepancies, divergent learning objectives, and distinct representation learning strategies of previous pipelines for open-vocabulary segmentation and in-context segmentation. Comprehensive experiments demonstrate that COSINE has significant performance improvements in both open-vocabulary and in-context segmentation tasks. Our exploratory analyses highlight that the synergistic collaboration between using visual and textual prompts leads to significantly improved generalization over single-modality approaches.

📄 PDF Abstract BibTeX arXiv:2510.10524

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning

2025-09-23 · Guoxin Wang, Jun Zhao, Xinyi Liu, Yanbo Liu 외 arxiv

Medical imaging provides critical evidence for clinical diagnosis, treatment planning, and surgical decisions, yet most existing imaging models are narrowly focused and require multiple specialized networks, limiting the…

Visual Grounding

OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language Prompts

2025-07-07 · Shiting Xiao, Rishabh Kabra, Yuhang Li, DongHyun Lee 외

The ability to segment objects based on open-ended language prompts remains a critical challenge, requiring models to ground textual semantics into precise spatial masks while handling diverse and unseen categories. We p…

Image SegmentationPanoptic SegmentationSemantic Segmentation

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation

2024-11-20 · Ziyi Wang, Yanbo Wang, Xumin Yu, Jie zhou 외

Existing methodologies in open vocabulary 3D semantic segmentation primarily concentrate on establishing a unified feature space encompassing 3D, 2D, and textual modalities. Nevertheless, traditional techniques such as g…

3D geometry3D Semantic SegmentationDenoisingOpen Vocabulary Semantic Segmentation+3

Open-set Cross Modal Generalization via Multimodal Unified Representation

2025-07-20 · Hai Huang, Yan Xia, Shulei Wang, Hanting Wang 외 arxiv

This paper extends Cross Modal Generalization (CMG) to open-set environments by proposing the more challenging Open-set Cross Modal Generalization (OSCMG) task. This task evaluates multimodal unified representations in o…

Self-Supervised LearningContrastive Learning

Unified Open-Vocabulary Dense Visual Prediction

2023-07-17 · Hengcan Shi, Munawar Hayat, Jianfei Cai

In recent years, open-vocabulary (OV) dense visual prediction (such as OV object detection, semantic, instance and panoptic segmentations) has attracted increasing research attention. However, most of existing approaches…

object-detectionObject DetectionPrediction