paper-with-me

Papers

DOA Estimation with Lightweight Network on LLM-Aided Simulated Acoustic Scenes

2025-11-11 · Haowen Li, Zhengding Luo, Dongyuan Shi, Boxiang Wang, Junwei Ji, Ziyi Yang, Woon-Seng Gan arxiv

Direction-of-Arrival (DOA) estimation is critical in spatial audio and acoustic signal processing, with wide-ranging applications in real-world. Most existing DOA models are trained on synthetic data by convolving clean speech with room impulse responses (RIRs), which limits their generalizability due to constrained acoustic diversity. In this paper, we revisit DOA estimation using a recently introduced dataset constructed with the assistance of large language models (LLMs), which provides more realistic and diverse spatial audio scenes. We benchmark several representative neural-based DOA methods on this dataset and propose LightDOA, a lightweight DOA estimation model based on depthwise separable convolutions, specifically designed for mutil-channel input in varying environments. Experimental results show that LightDOA achieves satisfactory accuracy and robustness across various acoustic scenes while maintaining low computational complexity. This study not only highlights the potential of spatial audio synthesized with the assistance of LLMs in advancing robust and efficient DOA estimation research, but also highlights LightDOA as efficient solution for resource-constrained applications.

📄 PDF Abstract BibTeX arXiv:2511.08012

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sound Source Localization for Spatial Mapping of Surgical Actions in Dynamic Scenes

2025-10-28 · Jonas Hein, Lazaros Vlachopoulos, Maurits Geert Laurent Olthof, Bastian Sigrist 외 arxiv

Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual…

Sound Source LocalizationScene UnderstandingPoint Clouds

Relative Acoustic Features for Distance Estimation in Smart-Homes

2022-12-02 · Francesco Nespoli, Daniel Barreda, Patrick A. Naylor

Any audio recording encapsulates the unique fingerprint of the associated acoustic environment, namely the background noise and reverberation. Considering the scenario of a room equipped with a fixed smart speaker device…

Room Impulse Response (RIR)

Single Microphone Own Voice Detection based on Simulated Transfer Functions for Hearing Aids

2026-03-03 · Mathuranathan Mayuravaani, W. Bastiaan Kleijn, Andrew Lensen, Charlotte Sørensen arxiv

This paper presents a simulation-based approach to own voice detection (OVD) in hearing aids using a single microphone. While OVD can significantly improve user comfort and speech intelligibility, existing solutions ofte…

Data Augmentation

Enhanced Scale-aware Depth Estimation for Monocular Endoscopic Scenes with Geometric Modeling

2024-08-14 · Ruofeng Wei, Bin Li, Kai Chen, Yiyao Ma 외

Scale-aware monocular depth estimation poses a significant challenge in computer-aided endoscopic navigation. However, existing depth estimation methods that do not consider the geometric priors struggle to learn the abs…

Depth EstimationMonocular Depth Estimation

A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation

2024-09-19 · Jingyuan Wang, Jie Zhang, Shihao Chen, Miao Sun

Binaural speech enhancement (BSE) aims to jointly improve the speech quality and intelligibility of noisy signals received by hearing devices and preserve the spatial cues of the target for natural listening. Existing me…

Speech Enhancement