paper-with-me

Papers

Open-set Cross Modal Generalization via Multimodal Unified Representation

2025-07-20 · Hai Huang, Yan Xia, Shulei Wang, Hanting Wang, Minghui Fang, Shengpeng Ji, Sashuai Zhou, Tao Jin, Zhou Zhao arxiv

This paper extends Cross Modal Generalization (CMG) to open-set environments by proposing the more challenging Open-set Cross Modal Generalization (OSCMG) task. This task evaluates multimodal unified representations in open-set conditions, addressing the limitations of prior closed-set cross-modal evaluations. OSCMG requires not only cross-modal knowledge transfer but also robust generalization to unseen classes within new modalities, a scenario frequently encountered in real-world applications. Existing multimodal unified representation work lacks consideration for open-set environments. To tackle this, we propose MICU, comprising two key components: Fine-Coarse Masked multimodal InfoNCE (FCMI) and Cross modal Unified Jigsaw Puzzles (CUJP). FCMI enhances multimodal alignment by applying contrastive learning at both holistic semantic and temporal levels, incorporating masking to enhance generalization. CUJP enhances feature diversity and model uncertainty by integrating modality-agnostic feature selection with self-supervised learning, thereby strengthening the model's ability to handle unknown categories in open-set tasks. Extensive experiments on CMG and the newly proposed OSCMG validate the effectiveness of our approach. The code is available at https://github.com/haihuangcode/CMG.

📄 PDF Abstract BibTeX arXiv:2507.14935

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningContrastive Learning

Similar Papers 제목 키워드 기반

CrossMed: A Multimodal Cross-Task Benchmark for Compositional Generalization in Medical Imaging

2025-11-14 · Pooja Singh, Siddhant Ujjain, Tapan Kumar Gandhi, Sandeep Kumar arxiv

Recent advances in multimodal large language models have enabled unified processing of visual and textual inputs, offering promising applications in general-purpose medical AI. However, their ability to generalize compos…

Visual Question Answering

Achieving Cross Modal Generalization with Multimodal Unified Representation

2023-09-21 · NeurIPS 2023 11

This paper introduces a novel task called Cross Modal Generalization (CMG), which addresses the challenge of learning a unified discrete representation from paired multimodal data during pre-training. Then in downstream …

MMaDA: Multimodal Large Diffusion Language Models

2025-05-21 · Ling Yang, Ye Tian, Bowen Li, Xinchen Zhang 외

We introduce MMaDA, a novel class of multimodal diffusion foundation models designed to achieve superior performance across diverse domains such as textual reasoning, multimodal understanding, and text-to-image generatio…

Image GenerationReinforcement Learning (RL)Text to Image GenerationText-to-Image Generation

Towards Multimodal Open-Set Domain Generalization and Adaptation through Self-supervision

2024-07-01 · Hao Dong, Eleni Chatzi, Olga Fink

The task of open-set domain generalization (OSDG) involves recognizing novel classes within unseen domains, which becomes more challenging with multiple modalities as input. Existing works have only addressed unimodal OS…

Domain AdaptationDomain GeneralizationMeta-Learning

Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning

2025-09-23 · Guoxin Wang, Jun Zhao, Xinyi Liu, Yanbo Liu 외 arxiv

Medical imaging provides critical evidence for clinical diagnosis, treatment planning, and surgical decisions, yet most existing imaging models are narrowly focused and require multiple specialized networks, limiting the…

Visual Grounding