paper-with-me

Papers

CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation

2026-04-16 · Inseok Jeon, Suhwan Cho, Minhyeok Lee, Seunghoon Lee, Minseok Kang, Jungho Lee, Chaewon Park, Donghyeong Kim, Sangyoun Lee arxiv

Recent advances in unsupervised video object segmentation have highlighted the potential of two-stream architectures that integrate appearance and motion cues. However, fully leveraging these complementary sources of information requires effectively modeling their interdependencies. In this paper, we introduce cross-modality token modulation, a novel approach designed to strengthen the interaction between appearance and motion cues. Our method establishes dense connections between tokens from each modality, enabling efficient intra-modal and inter-modal information propagation through relation transformer blocks. To improve learning efficiency, we incorporate a token masking strategy that addresses the limitations of relying solely on increased model complexity. Our approach achieves state-of-the-art performance across all public benchmarks, outperforming existing methods.

📄 PDF Abstract BibTeX arXiv:2604.14630

Code (0)

등록된 구현이 없습니다.

Tasks

Unsupervised Video Object Segmentation

Similar Papers 제목 키워드 기반

Code and Pixels: Multi-Modal Contrastive Pre-training for Enhanced Tabular Data Analysis

2025-01-13 · Kankana Roy, Lars Krämer, Sebastian Domaschke, Malik Haris 외

Learning from tabular data is of paramount importance, as it complements the conventional analysis of image and video data by providing a rich source of structured information that is often critical for comprehensive und…

Contrastive Learning

Infusing Sequential Information into Conditional Masked Translation Model with Self-Review Mechanism

2020-10-19 · COLING 2020 8 · Pan Xie, Zhi Cui, Xiuyin Chen, Xiaohui Hu 외

Non-autoregressive models generate target words in a parallel way, which achieve a faster decoding speed but at the sacrifice of translation accuracy. To remedy a flawed translation by non-autoregressive models, a promis…

DecoderKnowledge DistillationTranslation

STMI: Segmentation-Guided Token Modulation with Cross-Modal Hypergraph Interaction for Multi-Modal Object Re-Identification

2026-02-28 · Xingguo Xu, Zhanyu Liu, Weixiang Zhou, Yuansheng Gao 외 arxiv

Multi-modal object Re-Identification (ReID) aims to exploit complementary information from different modalities to retrieve specific objects. However, existing methods often rely on hard token filtering or simple fusion …

Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising Utility

2026-02-03 · Mengxuan Wang, Yuxin Chen, Gang Xu, Tao He 외 arxiv

Vision language models (VLMs) extend the reasoning capabilities of large language models (LLMs) to cross-modal settings, yet remain highly vulnerable to multimodal jailbreak attacks. Existing defenses predominantly rely …

A Dual-Modulation Framework for RGB-T Crowd Counting via Spatially Modulated Attention and Adaptive Fusion

2025-09-21 · Yuhong Feng, Hongtao Chen, Qi Zhang, Jie Chen 외 arxiv

Accurate RGB-Thermal (RGB-T) crowd counting is crucial for public safety in challenging conditions. While recent Transformer-based methods excel at capturing global context, their inherent lack of spatial inductive bias …

Crowd Counting