paper-with-me

홈 › Papers

Exploring the Potential of Data-Driven Spatial Audio Enhancement Using a Single-Channel Model

2024-04-22 · Arthur N. dos Santos, Bruno S. Masiero, Túlio C. L. Mateus

One key aspect differentiating data-driven single- and multi-channel speech enhancement and dereverberation methods is that both the problem formulation and complexity of the solutions are considerably more challenging in the latter case. Additionally, with limited computational resources, it is cumbersome to train models that require the management of larger datasets or those with more complex designs. In this scenario, an unverified hypothesis that single-channel methods can be adapted to multi-channel scenarios simply by processing each channel independently holds significant implications, boosting compatibility between sound scene capture and system input-output formats, while also allowing modern research to focus on other challenging aspects, such as full-bandwidth audio enhancement, competitive noise suppression, and unsupervised learning. This study verifies this hypothesis by comparing the enhancement promoted by a basic single-channel speech enhancement and dereverberation model with two other multi-channel models tailored to separate clean speech from noisy 3D mixes. A direction of arrival estimation model was used to objectively evaluate its capacity to preserve spatial information by comparing the output signals with ground-truth coordinate values. Consequently, a trade-off arises between preserving spatial information with a more straightforward single-channel solution at the cost of obtaining lower gains in intelligibility scores.

📄 PDF Abstract BibTeX arXiv:2404.14564

Code (0)

등록된 구현이 없습니다.

Tasks

Direction of Arrival EstimationSpeech Enhancement

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation

2025-08-01 · Kien T. Pham, Yingqing He, Yazhou Xing, Qifeng Chen 외 arxiv

Audio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focu…

Video Generation

EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian Deformation

2025-10-03 · Tianheng Zhu, Yinfeng Yu, Liejun Wang, Fuchun Sun 외 arxiv

This paper presents EGSTalker, a real-time audio-driven talking head generation framework based on 3D Gaussian Splatting (3DGS). Designed to enhance both speed and visual fidelity, EGSTalker requires only 3-5 minutes of …

Talking Head Generation

MOSPA: Human Motion Generation Driven by Spatial Audio

2025-07-16 · Shuyang Xu, Zhiyang Dou, Mingyi Shi, Liang Pan 외 arxiv

Enabling virtual humans to dynamically and realistically respond to diverse auditory stimuli remains a key challenge in character animation, demanding the integration of perceptual modeling and motion synthesis. Despite …

Motion Synthesis

DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis

2024-09-16 · Fa-Ting Hong, Yunfei Liu, Yu Li, Changyin Zhou 외

Audio-driven talking head synthesis strives to generate lifelike video portraits from provided audio. The diffusion model, recognized for its superior quality and robust generalization, has been explored for this task. H…

Talking Head Generation

PGSTalker: Real-Time Audio-Driven Talking Head Generation via 3D Gaussian Splatting with Pixel-Aware Density Control

2025-09-21 · Tianheng Zhu, Yinfeng Yu, Liejun Wang, Fuchun Sun 외 arxiv

Audio-driven talking head generation is crucial for applications in virtual reality, digital avatars, and film production. While NeRF-based methods enable high-fidelity reconstruction, they suffer from low rendering effi…

Talking Head Generation