paper-with-me

홈 › Papers

Dual-Stream Attention with Multi-Modal Queries for Object Detection in Transportation Applications

2025-08-06 · Noreen Anwar, Guillaume-Alexandre Bilodeau, Wassim Bouachir arxiv

Transformer-based object detectors often struggle with occlusions, fine-grained localization, and computational inefficiency caused by fixed queries and dense attention. We propose DAMM, Dual-stream Attention with Multi-Modal queries, a novel framework introducing both query adaptation and structured cross-attention for improved accuracy and efficiency. DAMM capitalizes on three types of queries: appearance-based queries from vision-language models, positional queries using polygonal embeddings, and random learned queries for general scene coverage. Furthermore, a dual-stream cross-attention module separately refines semantic and spatial features, boosting localization precision in cluttered scenes. We evaluated DAMM on four challenging benchmarks, and it achieved state-of-the-art performance in average precision (AP) and recall, demonstrating the effectiveness of multi-modal query adaptation and dual-stream attention. Source code is at: \href{https://github.com/DET-LIP/DAMM}{GitHub}.

📄 PDF Abstract BibTeX arXiv:2508.04868

Code (0)

등록된 구현이 없습니다.

Tasks

Object Detection

Similar Papers 제목 키워드 기반

Agentic Learner with Grow-and-Refine Multimodal Semantic Memory

2025-11-26 · Weihao Bo, Shan Zhang, Yanpeng Sun, Jingjing Wu 외 arxiv

MLLMs exhibit strong reasoning on isolated queries, yet they operate de novo -- solving each problem independently and often repeating the same mistakes. Existing memory-augmented agents mainly store past trajectories fo…

Logical Reasoning

Multi-Attention Network for Compressed Video Referring Object Segmentation

2022-07-26 · Weidong Chen, Dexiang Hong, Yuankai Qi, Zhenjun Han 외

Referring video object segmentation aims to segment the object referred by a given language expression. Existing works typically require compressed video bitstream to be decoded to RGB frames before being segmented, whic…

ObjectReferring Expression SegmentationReferring Video Object SegmentationSegmentation+3

Align Your Query: Representation Alignment for Multimodality Medical Object Detection

2025-10-03 · Ara Seo, Bryan Sangwoo Kim, Hyungjin Chung, Jong Chul Ye arxiv

Medical object detection suffers when a single detector is trained on mixed medical modalities (e.g., CXR, CT, MRI) due to heterogeneous statistics and disjoint representation spaces. To address this challenge, we turn t…

Medical Object Detection

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

2026-06-29 · Kien T. Pham, I Chieh Chen, Qifeng Chen, Long Chen hf

Audio-video generation has recently gained unprecedented research attention, aiming to synthesize high-quality sounding video content with fine-grained synchronization and semantic alignment between the auditory and visu…

Video ReconstructionVideo Generation

DREAM: Extending Vision-Language Models with Dual-Objective Encoding for Cross-Modal Retrieval

2026-06-17 · Kaleem Ullah, Altaf Hussain, Muhammad Munsif, Sung Wook Baik arxiv

In today's media-driven world, the exponential growth of video content across domains such as surveillance, education, and entertainment has made retrieving semantically relevant videos via natural language queries incre…

Natural Language QueriesRepresentation LearningCross-Modal RetrievalVideo Retrieval