paper-with-me

Papers

High-Frequency Semantics and Geometric Priors for End-to-End Detection Transformers in Challenging UAV Imagery

2025-07-01 · Hongxing Peng, Lide Chen, Hui Zhu, Yan Chen arxiv

Object detection in Unmanned Aerial Vehicle (UAV) imagery is fundamentally challenged by a prevalence of small, densely packed, and occluded objects within cluttered backgrounds. Conventional detectors struggle with this domain, as they rely on hand-crafted components like pre-defined anchors and heuristic-based Non-Maximum Suppression (NMS), creating a well-known performance bottleneck in dense scenes. Even recent end-to-end frameworks have not been purpose-built to overcome these specific aerial challenges, resulting in a persistent performance gap. To bridge this gap, we introduce HEDS-DETR, a holistically enhanced real-time Detection Transformer tailored for aerial scenes. Our framework features three key innovations. First, we propose a novel High-Frequency Enhanced Semantics Network (HFESNet) backbone, which yields highly discriminative features by preserving critical high-frequency details alongside robust semantic context. Second, our Efficient Small Object Pyramid (ESOP) counteracts information loss by efficiently fusing high-resolution features, significantly boosting small object detection. Finally, we enhance decoder stability and localization precision with two synergistic components: Selective Query Recollection (SQR) and Geometry-Aware Positional Encoding (GAPE), which stabilize optimization and provide explicit spatial priors for dense object arrangements. On the VisDrone dataset, HEDS-DETR achieves a +3.8% AP and +5.1% AP50 gain over its baseline while reducing parameters by 4M and maintaining real-time speeds. This demonstrates a highly competitive accuracy-efficiency balance, especially for detecting dense and small objects in aerial scenes.

📄 PDF Abstract BibTeX arXiv:2507.00825

Code (0)

등록된 구현이 없습니다.

Tasks

Small Object Detection

Similar Papers 제목 키워드 기반

CFSR: Geometry-Conditioned Shadow Removal via Physical Disentanglement

2026-04-20 · Pan Wang, Yihao Hu, Xiujin Liu, Hang Wang arxiv

Traditional shadow removal networks often treat image restoration as an unconstrained mapping, lacking the physical interpretability required to balance localized texture recovery with global illumination consistency. To…

Image RestorationShadow Removal

Enhancing Event-based Object Detection with Monocular Normal Maps

2025-08-04 · Mingjie Liu, Hanqing Liu, Luoping Cui, Chuang Zhu arxiv

Object detection in autonomous driving is frequently compromised by complex illumination. While event cameras offer a robust solution, they are susceptible to sudden contrast changes such as reflections which often trigg…

Autonomous DrivingObject Detection

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation

2026-08-01 · Chenggong Hu, Shaoyin Ma, Yi Wang, Li Sun 외 arxiv

Audio-driven emotional talking face generation aims to synthesize realistic videos with expressive facial dynamics. However, existing methods struggle to balance controllability and visual fidelity. Although implicit rep…

Talking Face GenerationContinuous Control

Text-Guided Multimodal Unified Industrial Anomaly Detection

2026-04-24 · Zewen Li, Shuo Ye, Zitong Yu, Weicheng Xie 외 arxiv

Industrial anomaly detection based on RGB-3D multimodal data has emerged as a mainstream paradigm for intelligent quality inspection. However, existing unsupervised methods suffer from two critical limitations: ambiguous…

Anomaly Detection

Disentangled Textual Priors for Diffusion-based Image Super-Resolution

2026-03-08 · Lei Jiang, Xin Liu, Xinze Tong, Zhiliang Li 외 arxiv

Image Super-Resolution (SR) aims to reconstruct high-resolution images from degraded low-resolution inputs. While diffusion-based SR methods offer powerful generative capabilities, their performance heavily depends on ho…

Image Super-Resolution