paper-with-me

Papers

OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation

2025-09-18 · Bo-Wen Yin, Jiao-Long Cao, Xuying Zhang, Yuming Chen, Ming-Ming Cheng, Qibin Hou arxiv

Recent research on representation learning has proved the merits of multi-modal clues for robust semantic segmentation. Nevertheless, a flexible pretrain-and-finetune pipeline for multiple visual modalities remains unexplored. In this paper, we propose a novel multi-modal learning framework, termed OmniSegmentor. It has two key innovations: 1) Based on ImageNet, we assemble a large-scale dataset for multi-modal pretraining, called ImageNeXt, which contains five popular visual modalities. 2) We provide an efficient pretraining manner to endow the model with the capacity to encode different modality information in the ImageNeXt. For the first time, we introduce a universal multi-modal pretraining framework that consistently amplifies the model's perceptual capabilities across various scenarios, regardless of the arbitrary combination of the involved modalities. Remarkably, our OmniSegmentor achieves new state-of-the-art records on a wide range of multi-modal semantic segmentation datasets, including NYU Depthv2, EventScape, MFNet, DeLiVER, SUNRGBD, and KITTI-360.

📄 PDF Abstract BibTeX arXiv:2509.15096

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningSemantic Segmentation

Similar Papers 제목 키워드 기반

Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment

2025-01-01 · CVPR 2025 1 · Mayug Maniparambil, Raiymbek Akshulakov, Yasser Abdelaziz Dahou Djilali, Sanath Narayan 외

Recent contrastive multimodal vision-language models like CLIP have demonstrated robust open-world semantic understanding, becoming the standard image backbones for vision-language applications. However, recent findi…

Semantic SimilaritySemantic Textual Similarity

FlexMUSE: Multimodal Unification and Semantics Enhancement Framework with Flexible interaction for Creative Writing

2025-08-22 · Jiahao Chen, Zhiyong Ma, Wenbiao Du, Qingyuan Chuai arxiv

Multi-modal creative writing (MMCW) aims to produce illustrated articles. Unlike common multi-modal generative (MMG) tasks such as storytelling or caption generation, MMCW is an entirely new and more abstract challenge w…

Relation Learning on Social Networks with Multi-Modal Graph Edge Variational Autoencoders

2019-11-04 · Carl Yang, Jieyu Zhang, Haonan Wang, Sha Li 외

While node semantics have been extensively explored in social networks, little research attention has been paid to profile edge semantics, i.e., social relations. Ideal edge semantics should not only show that two users …

Relation

High-fidelity Generalized Emotional Talking Face Generation with Multi-modal Emotion Space Learning

2023-05-04 · CVPR 2023 1 · Chao Xu, Junwei Zhu, Jiangning Zhang, Yue Han 외

Recently, emotional talking face generation has received considerable attention. However, existing methods only adopt one-hot coding, image, or audio as emotion conditions, thus lacking flexible control in practical appl…

Face GenerationTalking Face Generation

Support-set based Multi-modal Representation Enhancement for Video Captioning

2022-05-19 · Xiaoya Chen, Jingkuan Song, Pengpeng Zeng, Lianli Gao 외

Video captioning is a challenging task that necessitates a thorough comprehension of visual scenes. Existing methods follow a typical one-to-one mapping, which concentrates on a limited sample space while ignoring the in…

Video Captioning