paper-with-me

Papers

Focus Through Motion: RGB-Event Collaborative Token Sparsification for Efficient Object Detection

2025-09-04 · Nan Yang, Yang Wang, Zhanwen Liu, Yuchao Dai, Yang Liu, Xiangmo Zhao arxiv

Existing RGB-Event detection methods process the low-information regions of both modalities (background in images and non-event regions in event data) uniformly during feature extraction and fusion, resulting in high computational costs and suboptimal performance. To mitigate the computational redundancy during feature extraction, researchers have respectively proposed token sparsification methods for the image and event modalities. However, these methods employ a fixed number or threshold for token selection, hindering the retention of informative tokens for samples with varying complexity. To achieve a better balance between accuracy and efficiency, we propose FocusMamba, which performs adaptive collaborative sparsification of multimodal features and efficiently integrates complementary information. Specifically, an Event-Guided Multimodal Sparsification (EGMS) strategy is designed to identify and adaptively discard low-information regions within each modality by leveraging scene content changes perceived by the event camera. Based on the sparsification results, a Cross-Modality Focus Fusion (CMFF) module is proposed to effectively capture and integrate complementary features from both modalities. Experiments on the DSEC-Det and PKU-DAVIS-SOD datasets demonstrate that the proposed method achieves superior performance in both accuracy and efficiency compared to existing methods. The code will be available at https://github.com/Zizzzzzzz/FocusMamba.

📄 PDF Abstract BibTeX arXiv:2509.03872

Code (0)

등록된 구현이 없습니다.

Tasks

Object Detection

Similar Papers 제목 키워드 기반

EventPrune: Cascaded Event-Assisted Token Pruning for Efficient First-Person Dynamic Spatial Reasoning

2026-05-19 · Pengtao Ma, Ziliang Zhou, Ciyu Ruan, Haoyang Wang 외 arxiv

First-person dynamic spatial reasoning requires models to track continuous motion and precise geometric structure, but the quadratic attention cost of Transformer-based Video-LLMs makes dense visual tokens computationall…

Spatial Reasoning

Can Emotion Carriers Explain Automatic Sentiment Prediction? A Study on Personal Narratives

2022-05-01 · WASSA (ACL) 2022 5 · Seyed Mahed Mousavi, Gabriel Roccabruna, Aniruddha Tammewar, Steve Azzolin 외

Deep Neural Networks (DNN) models have achieved acceptable performance in sentiment prediction of written text. However, the output of these machine learning (ML) models cannot be natively interpreted. In this paper, we …

Motion-Based Tokenization for Cross-Dataset Egocentric Gaze Modeling

2026-08-24 · Virmarie Maquiling, Zhuojiang Cai, Enkelejda Kasneci arxiv

Gaze is increasingly used as an input signal for vision and multimodal models, yet no consensus exists on how to represent it across datasets. Raw traces preserve detail but are noisy and device-dependent, while coarse e…

InterMask: 3D Human Interaction Generation via Collaborative Masked Modelling

2024-10-13 · Muhammad Gohar Javed, Chuan Guo, Li Cheng, Xingyu Li

Generating realistic 3D human-human interactions from textual descriptions remains a challenging task. Existing approaches, typically based on diffusion models, often generate unnatural and unrealistic results. In this w…

Motion Synthesis

Motion-adaptive Separable Collaborative Filters for Blind Motion Deblurring

2024-04-19 · CVPR 2024 1 · Chengxu Liu, Xuan Wang, Xiangyu Xu, Ruhao Tian 외

Eliminating image blur produced by various kinds of motion has been a challenging problem. Dominant approaches rely heavily on model capacity to remove blurring by reconstructing residual from blurry observation in featu…

DeblurringMotion Estimation