paper-with-me

홈 › Papers

AMANet: Advancing SAR Ship Detection with Adaptive Multi-Hierarchical Attention Network

2024-01-24 · Xiaolin Ma, Junkai Cheng, Aihua Li, Yuhua Zhang, Zhilong Lin

Recently, methods based on deep learning have been successfully applied to ship detection for synthetic aperture radar (SAR) images. Despite the development of numerous ship detection methodologies, detecting small and coastal ships remains a significant challenge due to the limited features and clutter in coastal environments. For that, a novel adaptive multi-hierarchical attention module (AMAM) is proposed to learn multi-scale features and adaptively aggregate salient features from various feature layers, even in complex environments. Specifically, we first fuse information from adjacent feature layers to enhance the detection of smaller targets, thereby achieving multi-scale feature enhancement. Then, to filter out the adverse effects of complex backgrounds, we dissect the previously fused multi-level features on the channel, individually excavate the salient regions, and adaptively amalgamate features originating from different channels. Thirdly, we present a novel adaptive multi-hierarchical attention network (AMANet) by embedding the AMAM between the backbone network and the feature pyramid network (FPN). Besides, the AMAM can be readily inserted between different frameworks to improve object detection. Lastly, extensive experiments on two large-scale SAR ship detection datasets demonstrate that our AMANet method is superior to state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2401.13214

Code (0)

등록된 구현이 없습니다.

Tasks

object-detectionObject DetectionSAR Ship Detection

Similar Papers 제목 키워드 기반

CAMANet: Class Activation Map Guided Attention Network for Radiology Report Generation

2022-11-02 · Jun Wang, Abhir Bhalerao, Terry Yin, Simon See 외

Radiology report generation (RRG) has gained increasing research attention because of its huge potential to mitigate medical resource shortages and aid the process of disease decision making by radiologists. Recent advan…

cross-modal alignmentDecision Making

ObamaNet: Photo-realistic lip-sync from text

2017-12-06 · Rithesh Kumar, Jose Sotelo, Kundan Kumar, Alexandre de Brebisson 외

We present ObamaNet, the first architecture that generates both audio and synchronized photo-realistic lip-sync videos from any new text. Contrary to other published lip-sync approaches, ours is only composed of fully tr…

Constrained Lip-synchronizationtext-to-speechText to Speech

Semantic-Aware Ship Detection with Vision-Language Integration

2025-08-21 · Jiahao Li, Jiancheng Pan, Yuze Sun, Xiaomeng Huang arxiv

Ship detection in remote sensing imagery is a critical task with wide-ranging applications, such as maritime activity monitoring, shipping logistics, and environmental studies. However, existing methods often struggle to…

AUD-TGN: Advancing Action Unit Detection with Temporal Convolution and GPT-2 in Wild Audiovisual Contexts

2024-03-20 · Jun Yu, Zerui Zhang, Zhihong Wei, Gongpeng Zhao 외

Leveraging the synergy of both audio data and visual data is essential for understanding human emotions and behaviors, especially in in-the-wild setting. Traditional methods for integrating such multimodal information of…

Action Unit DetectionFacial Action Unit Detection

ARM: Refining Multivariate Forecasting with Adaptive Temporal-Contextual Learning

2023-10-14 · Jiecheng Lu, Xu Han, Shihao Yang

Long-term time series forecasting (LTSF) is important for various domains but is confronted by challenges in handling the complex temporal-contextual relationships. As multivariate input models underperforming some recen…

Time SeriesTime Series Forecasting