paper-with-me

Papers

FinSight-Net:A Physics-Aware Decoupled Network with Frequency-Domain Compensation for Underwater Fish Detection in Smart Aquaculture

2026-02-23 · Jinsong Yang, Zeyuan Hu, Yichen Li, Hong Yu arxiv

Underwater fish detection (UFD) is a core capability for smart aquaculture and marine ecological monitoring. While recent detectors improve accuracy by stacking feature extractors or introducing heavy attention modules, they often incur substantial computational overhead and, more importantly, neglect the physics that fundamentally limits UFD: wavelength-dependent absorption and turbidity-induced scattering significantly degrade contrast, blur fine structures, and introduce backscattering noise, leading to unreliable localization and recognition. To address these challenges, we propose FinSight-Net, an efficient and physics-aware detection framework tailored for complex aquaculture environments. FinSight-Net introduces a Multi-Scale Decoupled Dual-Stream Processing (MS-DDSP) bottleneck that explicitly targets frequency-specific information loss via heterogeneous convolutional branches, suppressing backscattering artifacts while compensating distorted biological cues through scale-aware and channel-weighted pathways. We further design an Efficient Path Aggregation FPN (EPA-FPN) as a detail-filling mechanism: it restores high-frequency spatial information typically attenuated in deep layers by establishing long-range skip connections and pruning redundant fusion routes, enabling robust detection of non-rigid fish targets under severe blur and turbidity. Extensive experiments on DeepFish, AquaFishSet, and our challenging UW-BlurredFish benchmark demonstrate that FinSight-Net achieves state-of-the-art performance. In particular, on UW-BlurredFish, FinSight-Net reaches 92.8% mAP, outperforming YOLOv11s by 4.8% while reducing parameters by 29.0%, providing a strong and lightweight solution for real-time automated monitoring in smart aquaculture.

📄 PDF Abstract BibTeX arXiv:2602.19437

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FinSight: Towards Real-World Financial Deep Research

2025-10-19 · Jiajie Jin, Yuyao Zhang, Yimeng Xu, Hongjin Qian 외 arxiv

Generating professional financial reports is a labor-intensive and intellectually demanding process that current AI systems struggle to fully automate. To address this challenge, we introduce FinSight (Financial InSight)…

Dual-Domain Manifold Modeling for Hyperspectral Image Fusion

2026-07-28 · Chengxin Xie, Qiya Song, Yangbangyan Jiang, Renwei Dian 외 arxiv

Achieving a coherent integration of spectral richness and spatial fidelity remains a central objective in hyperspectral image fusion. However, existing hyperspectral image fusion methods struggle to effectively model geo…

Representation LearningSpectral Reconstruction

FUSE: Label-Free Image-Event Joint Monocular Depth Estimation via Frequency-Decoupled Alignment and Degradation-Robust Fusion

2025-03-25 · Pihai Sun, Junjun Jiang, Yuanqi Yao, Youyu Chen 외

Image-event joint depth estimation methods leverage complementary modalities for robust perception, yet face challenges in generalizability stemming from two factors: 1) limited annotated image-event-depth datasets causi…

Depth EstimationMonocular Depth EstimationTransfer Learning

FCDM: A Physics-Guided Bidirectional Frequency Aware Convolution and Diffusion-Based Model for Sinogram Inpainting

2024-08-26 · Jiaze E, Srutarshi Banerjee, Tekin Bicer, Guannan Wang 외

Computed tomography (CT) is widely used in industrial and medical imaging, but sparse-view scanning reduces radiation exposure at the cost of incomplete sinograms and challenging reconstruction. Existing RGB-based inpain…

Computed Tomography (CT)CT ReconstructionImage ReconstructionScheduling+1

Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility

2025-09-29 · Yutong Hao, Chen Chen, Ajmal Saeed Mian, Chang Xu 외 arxiv

Diffusion models can generate realistic videos, but existing methods rely on implicitly learning physical reasoning from large-scale text-video datasets, which is costly, difficult to scale, and still prone to producing …

Video Generation