paper-with-me

Papers

DIMOS: Disentangling Instance-level Moving Object Segmentation

2026-06-11 · Hongxiang Huang, Hongwei Ren, Xiaopeng Lin, Yulong Huang, Zeke Xie, Bojun Cheng arxiv

Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, and animal tracking. Event cameras record asynchronous brightness changes, providing high temporal resolution and dynamic range, which makes them highly sensitive to motion information. By fusing event and image features, motion cues from events can complement spatial details from images, enhancing the performance of MIS. However, current multimodal MIS methods still struggle to segment small moving instances, as event cameras often yield sparse features under limited resolution. Moreover, event features entangle appearance attributes with motion cues, which further restricts effective cross-modal fusion. To address these challenges, we first propose a dual-disentangling feature extraction framework that separates and extracts appearance and motion information within both image and event modalities, thereby improving feature density. Subsequently, a multi-granularity cross-modal alignment is introduced to align distributionally and semantically consistent features across modalities, enabling more effective fusion with rich spatial and temporal details. The experiment results demonstrate that our method achieves state-of-the-art performance in multimodal MIS, especially for small instances under challenging conditions such as fast motion and low-light settings.

📄 PDF Abstract BibTeX arXiv:2606.12826

Code (0)

등록된 구현이 없습니다.

Tasks

Instance SegmentationObject SegmentationAutonomous Driving

Similar Papers 제목 키워드 기반

DiMoSR: Feature Modulation via Multi-Branch Dilated Convolutions for Efficient Image Super-Resolution

2025-05-27 · M. Akin Yilmaz, Ahmet Bilican, A. Murat Tekalp

Balancing reconstruction quality versus model efficiency remains a critical challenge in lightweight single image super-resolution (SISR). Despite the prevalence of attention mechanisms in recent state-of-the-art SISR ap…

Computational EfficiencyImage Super-ResolutionSSIMSuper-Resolution

Fast, Modular, and Differentiable Framework for Machine Learning-Enhanced Molecular Simulations

2025-03-26 · Henrik Christiansen, Takashi Maruyama, Federico Errica, Viktor Zaverkin 외

We present an end-to-end differentiable molecular simulation framework (DIMOS) for molecular dynamics and Monte Carlo simulations. DIMOS easily integrates machine-learning-based interatomic potentials and implements clas…

Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion

2023-12-29 · Yun Chen, Lingxiao Yang, Qi Chen, Jian-Huang Lai 외

Emotional Voice Conversion aims to manipulate a speech according to a given emotion while preserving non-emotion components. Existing approaches cannot well express fine-grained emotional attributes. In this paper, we pr…

Contrastive LearningDisentanglementVoice Conversion

Towards Instance-level Image-to-Image Translation

2019-05-05 · CVPR 2019 6 · Zhiqiang Shen, Mingyang Huang, Jianping Shi, xiangyang xue 외

Unpaired Image-to-image Translation is a new rising and challenging vision problem that aims to learn a mapping between unaligned image pairs in diverse domains. Recent advances in this field like MUNIT and DRIT mainly f…

AttributeImage-to-Image Translationobject-detectionObject Detection+1

InstMove: Instance Motion for Object-centric Video Segmentation

2023-03-14 · CVPR 2023 1 · Qihao Liu, Junfeng Wu, Yi Jiang, Xiang Bai 외

Despite significant efforts, cutting-edge video segmentation methods still remain sensitive to occlusion and rapid movement, due to their reliance on the appearance of objects in the form of object embeddings, which are …

ObjectOptical Flow EstimationSegmentationVideo Segmentation+1