paper-with-me

Papers

HebDB: a Weakly Supervised Dataset for Hebrew Speech Processing

2024-07-10 · Arnon Turetzky, Or Tal, Yael Segal-Feldman, Yehoshua Dissen, Ella Zeldes, Amit Roth, Eyal Cohen, Yosi Shrem, Bronya R. Chernyak, Olga Seleznova, Joseph Keshet, Yossi Adi

We present HebDB, a weakly supervised dataset for spoken language processing in the Hebrew language. HebDB offers roughly 2500 hours of natural and spontaneous speech recordings in the Hebrew language, consisting of a large variety of speakers and topics. We provide raw recordings together with a pre-processed, weakly supervised, and filtered version. The goal of HebDB is to further enhance research and development of spoken language processing tools for the Hebrew language. Hence, we additionally provide two baseline systems for Automatic Speech Recognition (ASR): (i) a self-supervised model; and (ii) a fully supervised model. We present the performance of these two methods optimized on HebDB and compare them to current multi-lingual ASR alternatives. Results suggest the proposed method reaches better results than the evaluated baselines considering similar model sizes. Dataset, code, and models are publicly available under https://pages.cs.huji.ac.il/adiyoss-lab/HebDB/.

📄 PDF Abstract BibTeX arXiv:2407.07566

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

A Language Modeling Approach to Diacritic-Free Hebrew TTS

2024-07-16 · Amit Roth, Arnon Turetzky, Yossi Adi

We tackle the task of text-to-speech (TTS) in Hebrew. Traditional Hebrew contains Diacritics, which dictate the way individuals should pronounce given words, however, modern Hebrew rarely uses them. The lack of diacritic…

Language ModelingLanguage Modellingtext-to-speechText to Speech

ivrit.ai: A Comprehensive Dataset of Hebrew Speech for AI Research and Development

2023-07-17 · Yanir Marmor, Kinneret Misgav, Yair Lifshitz

We introduce "ivrit.ai", a comprehensive Hebrew speech dataset, addressing the distinct lack of extensive, high-quality resources for advancing Automated Speech Recognition (ASR) technology in Hebrew. With over 3,300 spe…

Action DetectionActivity Detectionspeech-recognitionSpeech Recognition

Phonikud: Hebrew Grapheme-to-Phoneme Conversion for Real-Time Text-to-Speech

2025-06-14 · Yakov Kolani, Maxim Melichov, Cobi Calev, Morris Alper

Real-time text-to-speech (TTS) for Modern Hebrew is challenging due to the language's orthographic complexity. Existing solutions ignore crucial phonetic features such as stress that remain underspecified even when vowel…

Grapheme-to-Phoneme Conversiontext-to-speechText to Speech

ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion

2026-06-18 · Maxim Melichov, Yakov Kolani, Morris Alper arxiv

Grapheme-to-phoneme (G2P) conversion for Modern Hebrew is needed for applications like text-to-speech (TTS), but is challenging due to the language's abjad writing system, which leaves vowels largely unwritten, creating …

Speech Recognition

VoxKnesset: A Large-Scale Longitudinal Hebrew Speech Dataset for Aging Speaker Modeling

2026-03-01 · Yanir Marmor, Arad Zulti, David Krongauz, Adam Gabet 외 arxiv

Speech processing systems face a fundamental challenge: the human voice changes with age, yet few datasets support rigorous longitudinal evaluation. We introduce VoxKnesset, an open-access dataset of ~2,300 hours of Hebr…

Speaker Verification