paper-with-me

Papers

Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition

2025-06-17 · Jiamin Xie, Ju Lin, Yiteng Huang, Tyler Vuong, Zhaojiang Lin, Zhaojun Yang, Peng Su, Prashant Rawat, Sangeeta Srivastava, Ming Sun, Florian Metze

Recent studies have demonstrated that prompting large language models (LLM) with audio encodings enables effective speech recognition capabilities. However, the ability of Speech LLMs to comprehend and process multi-channel audio with spatial cues remains a relatively uninvestigated area of research. In this work, we present directional-SpeechLlama, a novel approach that leverages the microphone array of smart glasses to achieve directional speech recognition, source localization, and bystander cross-talk suppression. To enhance the model's ability to understand directivity, we propose two key techniques: serialized directional output training (S-DOT) and contrastive direction data augmentation (CDDA). Experimental results show that our proposed directional-SpeechLlama effectively captures the relationship between textual cues and spatial audio, yielding strong performance in both speech recognition and source localization tasks.

📄 PDF Abstract BibTeX arXiv:2506.14973

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationLanguage ModelingLanguage ModellingLarge Language Modelspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Equipping LLM with Directional Multi-Talker Speech Understanding Capabilities

2026-02-06 · Ju Lin, Jing Pan, Ruizhi Li, Ming Sun 외 arxiv

Recent studies have demonstrated that prompting large language models (LLM) with audio encodings enables effective speech understanding capabilities. However, most speech LLMs are trained on single-channel, single-talker…

Speech Recognition

Reconstructing the Dynamic Directivity of Unconstrained Speech

2022-09-09 · Camille Noufi, Dejan Markovic, Peter Dodds

This article presents a method for estimating and reconstructing the spatial energy distribution pattern of natural speech, which is crucial for achieving realistic vocal presence in virtual communication settings. The m…

Neural Directional Filtering: Far-Field Directivity Control With a Small Microphone Array

2024-09-20 · Julian Wechsler, Srikanth Raj Chetupalli, Mhd Modar Halimeh, Oliver Thiergart 외

Capturing audio signals with specific directivity patterns is essential in speech communication. This study presents a deep neural network (DNN)-based approach to directional filtering, alleviating the need for explicit …

MB-RIRs: a Synthetic Room Impulse Response Dataset with Frequency-Dependent Absorption Coefficients

2025-07-13 · Enric Gusó, Joanna Luberadzka, Umut Sayin, Xavier Serra arxiv

We investigate the effects of four strategies for improving the ecological validity of synthetic room impulse response (RIR) datasets for monoaural Speech Enhancement (SE). We implement three features on top of the tradi…

Speech Enhancement

GA-Aided Directivity in Volumetric and Planar Massive-Antenna Array Design

2023-01-07 · Bruno Felipe Costa, Taufik Abrão

The problem of directivity enhancement, leading to the increase in the directivity gain over a certain desired angle of arrival/departure (AoA/AoD), is considered in this work. A new formulation of the volumetric array d…