paper-with-me

Papers

VoxKnesset: A Large-Scale Longitudinal Hebrew Speech Dataset for Aging Speaker Modeling

2026-03-01 · Yanir Marmor, Arad Zulti, David Krongauz, Adam Gabet, Yoad Snapir, Yair Lifshitz, Eran Segal arxiv

Speech processing systems face a fundamental challenge: the human voice changes with age, yet few datasets support rigorous longitudinal evaluation. We introduce VoxKnesset, an open-access dataset of ~2,300 hours of Hebrew parliamentary speech spanning 2009-2025, comprising 393 speakers with recording spans of up to 15 years. Each segment includes aligned transcripts and verified demographic metadata from official parliamentary records. We benchmark modern speech embeddings (WavLM-Large, ECAPA-TDNN, Wav2Vec2-XLSR-1B) on age prediction and speaker verification under longitudinal conditions. Speaker verification EER rises from 2.15\% to 4.58\% over 15 years for the strongest model, and cross-sectionally trained age regressors fail to capture within-speaker aging, while longitudinally trained models recover a meaningful temporal signal. We publicly release the dataset and pipeline to support aging-robust speech systems and Hebrew speech processing.

📄 PDF Abstract BibTeX arXiv:2603.01270

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Verification

Similar Papers 제목 키워드 기반

ivrit.ai: A Comprehensive Dataset of Hebrew Speech for AI Research and Development

2023-07-17 · Yanir Marmor, Kinneret Misgav, Yair Lifshitz

We introduce "ivrit.ai", a comprehensive Hebrew speech dataset, addressing the distinct lack of extensive, high-quality resources for advancing Automated Speech Recognition (ASR) technology in Hebrew. With over 3,300 spe…

Action DetectionActivity Detectionspeech-recognitionSpeech Recognition

Phonikud: Hebrew Grapheme-to-Phoneme Conversion for Real-Time Text-to-Speech

2025-06-14 · Yakov Kolani, Maxim Melichov, Cobi Calev, Morris Alper

Real-time text-to-speech (TTS) for Modern Hebrew is challenging due to the language's orthographic complexity. Existing solutions ignore crucial phonetic features such as stress that remain underspecified even when vowel…

Grapheme-to-Phoneme Conversiontext-to-speechText to Speech

A Language Modeling Approach to Diacritic-Free Hebrew TTS

2024-07-16 · Amit Roth, Arnon Turetzky, Yossi Adi

We tackle the task of text-to-speech (TTS) in Hebrew. Traditional Hebrew contains Diacritics, which dictate the way individuals should pronounce given words, however, modern Hebrew rarely uses them. The lack of diacritic…

Language ModelingLanguage Modellingtext-to-speechText to Speech

AlephBERT:A Hebrew Large Pre-Trained Language Model to Start-off your Hebrew NLP Application With

2021-04-08 · Amit Seker, Elron Bandel, Dan Bareket, Idan Brusilovsky 외

Large Pre-trained Language Models (PLMs) have become ubiquitous in the development of language understanding technology and lie at the heart of many artificial intelligence advances. While advances reported for English u…

Language ModelingLanguage ModellingMorphological Taggingnamed-entity-recognition+4

The Enemy from Within: A Study of Political Delegitimization Discourse in Israeli Political Speech

2025-08-21 · Naama Rivlin-Angert, Guy Mor-Lan arxiv

We present the first large-scale computational study of political delegitimization discourse (PDD), defined as symbolic attacks on the normative validity of political entities. We curate and manually annotate a novel Heb…