paper-with-me

Papers

MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic Segmentation

2023-08-09 · ICCV 2023 1 · Kaixin Cai, Pengzhen Ren, Yi Zhu, Hang Xu, Jianzhuang Liu, Changlin Li, Guangrun Wang, Xiaodan Liang

Recently, semantic segmentation models trained with image-level text supervision have shown promising results in challenging open-world scenarios. However, these models still face difficulties in learning fine-grained semantic alignment at the pixel level and predicting accurate object masks. To address this issue, we propose MixReorg, a novel and straightforward pre-training paradigm for semantic segmentation that enhances a model's ability to reorganize patches mixed across images, exploring both local visual relevance and global semantic coherence. Our approach involves generating fine-grained patch-text pairs data by mixing image patches while preserving the correspondence between patches and text. The model is then trained to minimize the segmentation loss of the mixed images and the two contrastive losses of the original and restored features. With MixReorg as a mask learner, conventional text-supervised semantic segmentation models can achieve highly generalizable pixel-semantic alignment ability, which is crucial for open-world segmentation. After training with large-scale image-text data, MixReorg models can be applied directly to segment visual objects of arbitrary categories, without the need for further fine-tuning. Our proposed framework demonstrates strong performance on popular zero-shot semantic segmentation benchmarks, outperforming GroupViT by significant margins of 5.0%, 6.2%, 2.5%, and 3.4% mIoU on PASCAL VOC2012, PASCAL Context, MS COCO, and ADE20K, respectively.

📄 PDF Abstract BibTeX arXiv:2308.04829

Code (0)

등록된 구현이 없습니다.

Tasks

SegmentationSemantic SegmentationZero-Shot Semantic Segmentation

Similar Papers 제목 키워드 기반

Visible-Infrared Person Re-Identification via Patch-Mixed Cross-Modality Learning

2023-02-16 · Zhihao Qian, Yutian Lin, Bo Du

Visible-infrared person re-identification (VI-ReID) aims to retrieve images of the same pedestrian from different modalities, where the challenges lie in the significant modality discrepancy. To alleviate the modality ga…

Image GenerationPerson Re-IdentificationRepresentation LearningSemantic correspondence

UCM-VeID V2: A Richer Dataset and A Pre-training Method for UAV Cross-Modality Vehicle Re-Identification

2025-01-01 · CVPR 2025 1 · Xingyue Liu, Jiahao Qi, Chen Chen, Kangcheng Bin 외

Cross-Modality Re-Identification (VI-ReID) aims to achieve around-the-clock target matching, benefiting from the strengths of both RGB and infrared (IR) modalities. However, the field is hindered by limited datasets…

Image ReconstructionSelf-Supervised LearningVehicle Re-Identification

Do Large Language Models Have Emotions?

2026-06-03 · Amit Goldenberg, James J. Gross arxiv

Do LLMs have emotions? A recent paper from Anthropic reports finding internal representations of emotion concepts in Claude Sonnet 4.5, concluding that the LLM has 'functional emotions.' We evaluate this claim against wh…

Multimodal Graph Neural Networks for Prognostic Modeling of Brain Network Reorganization

2025-12-06 · Preksha Girish, Rachana Mysore, Kiran K. N., Hiranmayee R. 외 arxiv

Understanding the dynamic reorganization of brain networks is critical for predicting cognitive decline, neurological progression, and individual variability in clinical outcomes. This work proposes a multimodal graph ne…

Graph Neural Network

Learning differentially reorganizes brain activity and connectivity

2020-09-10

Human learning is a complex process in which future behavior is altered via the reorganization of brain activity and connectivity. It remains unknown whether activity and connectivity differentially reorganize during lea…

Functional Connectivity