paper-with-me

홈 › Papers

Treble10: A high-quality dataset for far-field speech recognition, dereverberation, and enhancement

2025-10-27 · Sarabeth S. Mullins, Georg Götz, Eric Bezzam, Steven Zheng, Daniel Gert Nielsen arxiv

Accurate far-field speech datasets are critical for tasks such as automatic speech recognition (ASR), dereverberation, speech enhancement, and source separation. However, current datasets are limited by the trade-off between acoustic realism and scalability. Measured corpora provide faithful physics but are expensive, low-coverage, and rarely include paired clean and reverberant data. In contrast, most simulation-based datasets rely on simplified geometrical acoustics, thus failing to reproduce key physical phenomena like diffraction, scattering, and interference that govern sound propagation in complex environments. We introduce Treble10, a large-scale, physically accurate room-acoustic dataset. Treble10 contains over 3000 broadband room impulse responses (RIRs) simulated in 10 fully furnished real-world rooms, using a hybrid simulation paradigm implemented in the Treble SDK that combines a wave-based and geometrical acoustics solver. The dataset provides six complementary subsets, spanning mono, 8th-order Ambisonics, and 6-channel device RIRs, as well as pre-convolved reverberant speech scenes paired with LibriSpeech utterances. All signals are simulated at 32 kHz, accurately modelling low-frequency wave effects and high-frequency reflections. Treble10 bridges the realism gap between measurement and simulation, enabling reproducible, physically grounded evaluation and large-scale data augmentation for far-field speech tasks. The dataset is openly available via the Hugging Face Hub, and is intended as both a benchmark and a template for next-generation simulation-driven audio research.

📄 PDF Abstract BibTeX arXiv:2510.23141

Code (0)

등록된 구현이 없습니다.

Tasks

Speech RecognitionSpeech EnhancementData Augmentation

Similar Papers 제목 키워드 기반

Towards Improving Speaker Distance Estimation through Generative Impulse Response Augmentation

2026-05-01 · Anton Ratnarajah, Mehmet Ergezer, Arun Nair, Mrudula Athi arxiv

The Room Acoustics and Speaker Distance Estimation (SDE) Challenge at ICASSP 2025 explores the effectiveness of augmented room impulse response (RIR) data for improving SDE model performance. This challenge at GenDARA in…

Hyperparameter Optimization

Treble Counterfactual VLMs: A Causal Approach to Hallucination

2025-03-08 · Li Li, Jiashu Qu, Yuxiao Zhou, Yuehan Qin 외

Vision-Language Models (VLMs) have advanced multi-modal tasks like image captioning, visual question answering, and reasoning. However, they often generate hallucinated outputs inconsistent with the visual context or pro…

Autonomous DrivingcounterfactualHallucinationImage Captioning+2

ctPuLSE: Close-Talk, and Pseudo-Label Based Far-Field, Speech Enhancement

2024-07-28 · Zhong-Qiu Wang

The current dominant approach for neural speech enhancement is via purely-supervised deep learning on simulated pairs of far-field noisy-reverberant speech (i.e., mixtures) and clean speech. The trained models, however, …

Pseudo LabelSpeech Enhancement

SpeechAlign: a Framework for Speech Translation Alignment Evaluation

2023-09-20 · Belen Alastruey, Aleix Sant, Gerard I. Gállego, David Dale 외

Speech-to-Speech and Speech-to-Text translation are currently dynamic areas of research. In our commitment to advance these fields, we present SpeechAlign, a framework designed to evaluate the underexplored field of sour…

Speech-to-TextSpeech-to-Text TranslationTranslation

MANNER: Multi-view Attention Network for Noise Erasure

2022-03-04 · Hyun Joon Park, Byung Ha Kang, WooSeok Shin, Jin Sob Kim 외

In the field of speech enhancement, time domain methods have difficulties in achieving both high performance and efficiency. Recently, dual-path models have been adopted to represent long sequential features, but they st…

DecoderSpeech Enhancement