paper-with-me

Papers

Multi-modal Iterative and Deep Fusion Frameworks for Enhanced Passive DOA Sensing via a Green Massive H2AD MIMO Receiver

2024-11-11 · Jiatong Bai, Minghao Chen, Wankai Tang, YiFan Li, Cunhua Pan, Yongpeng Wu, Feng Shu

Most existing DOA estimation methods assume ideal source incident angles with minimal noise. Moreover, directly using pre-estimated angles to calculate weighted coefficients can lead to performance loss. Thus, a green multi-modal (MM) fusion DOA framework is proposed to realize a more practical, low-cost and high time-efficiency DOA estimation for a H$^2$AD array. Firstly, two more efficient clustering methods, global maximum cos\_similarity clustering (GMaxCS) and global minimum distance clustering (GMinD), are presented to infer more precise true solutions from the candidate solution sets. Based on this, an iteration weighted fusion (IWF)-based method is introduced to iteratively update weighted fusion coefficients and the clustering center of the true solution classes by using the estimated values. Particularly, the coarse DOA calculated by fully digital (FD) subarray, serves as the initial cluster center. The above process yields two methods called MM-IWF-GMaxCS and MM-IWF-GMinD. To further provide a higher-accuracy DOA estimation, a fusion network (fusionNet) is proposed to aggregate the inferred two-part true angles and thus generates two effective approaches called MM-fusionNet-GMaxCS and MM-fusionNet-GMinD. The simulation outcomes show the proposed four approaches can achieve the ideal DOA performance and the CRLB. Meanwhile, proposed MM-fusionNet-GMaxCS and MM-fusionNet-GMinD exhibit superior DOA performance compared to MM-IWF-GMaxCS and MM-IWF-GMinD, especially in extremely-low SNR range.

📄 PDF Abstract BibTeX arXiv:2411.06927

Code (0)

등록된 구현이 없습니다.

Tasks

Clustering

Similar Papers 제목 키워드 기반

Joint Segmentation and Grading with Iterative Optimization for Multimodal Glaucoma Diagnosis

2026-03-15 · Zhiwei Wang, Yuxing Li, Meilu Zhu, Defeng He 외 arxiv

Accurate diagnosis of glaucoma is challenging, as early-stage changes are subtle and often lack clear structural or appearance cues. Most existing approaches rely on a single modality, such as fundus or optical coherence…

CMFN: Cross-Modal Fusion Network for Irregular Scene Text Recognition

2024-01-18 · Jinzhi Zheng, Ruyi Ji, Libo Zhang, Yanjun Wu 외

Scene text recognition, as a cross-modal task involving vision and text, is an important research topic in computer vision. Most existing methods use language models to extract semantic information for optimizing visual …

PositionScene Text Recognition

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs

2025-07-22 · Yangshu Yuan, Heng Chen, Xinyi Jiang, Christian Ng 외 arxiv

The rapid advancement of Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) has enhanced our ability to process and generate human language and visual information. However, these models often struggle …

Instruction FollowingMultimodal ReasoningResponse GenerationLogical Reasoning

Vision-Enhanced Time Series Forecasting via Latent Diffusion Models

2025-02-16 · Weilin Ruan, Siru Zhong, Haomin Wen, Yuxuan Liang

Diffusion models have recently emerged as powerful frameworks for generating high-quality images. While recent studies have explored their application to time series forecasting, these approaches face significant challen…

Image ReconstructionTime SeriesTime Series Forecasting

Human Action Recognition Using Deep Multilevel Multimodal (M2) Fusion of Depth and Inertial Sensors

2019-10-25 · Zeeshan Ahmad, Naimul Khan

Multimodal fusion frameworks for Human Action Recognition (HAR) using depth and inertial sensor data have been proposed over the years. In most of the existing works, fusion is performed at a single level (feature level …

Action RecognitionTemporal Action Localization