paper-with-me

Papers

Generating a Yiddish Speech Corpus, Forced Aligner and Basic ASR System for the AHEYM Project

2016-05-01 · LREC 2016 5 · Malgorzata {\'C}avar, Damir {\'C}avar, Dov-Ber Kerler, Anya Quilitzsch

To create automatic transcription and annotation tools for the AHEYM corpus of recorded interviews with Yiddish speakers in Eastern Europe we develop initial Yiddish language resources that are used for adaptations of speech and language technologies. Our project aims at the development of resources and technologies that can make the entire AHEYM corpus and other Yiddish resources more accessible to not only the community of Yiddish speakers or linguists with language expertise, but also historians and experts from other disciplines or the general public. In this paper we describe the rationale behind our approach, the procedures and methods, and challenges that are not specific to the AHEYM corpus, but apply to all documentary language data that is collected in the field. To the best of our knowledge, this is the first attempt to create a speech corpus and speech technologies for Yiddish. This is also the first attempt to work out speech and language technologies to transcribe and translate a large collection of Yiddish spoken language resources.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Corpus Phonetics Tutorial

2018-11-13 · Eleanor Chodroff

Corpus phonetics has become an increasingly popular method of research in linguistic analysis. With advances in speech technology and computational power, large scale processing of speech data has become a viable techniq…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

A Part-of-Speech Tagger for Yiddish

2022-04-03 · Seth Kulick, Neville Ryant, Beatrice Santorini, Joel Wallenberg 외

We describe the construction and evaluation of a part-of-speech tagger for Yiddish. This is the first step in a larger project of automatically assigning part-of-speech tags and syntactic structure to Yiddish text for pu…

Word Embeddings

LDC Forced Aligner

2012-05-01 · LREC 2012 5 · Xiaoyi Ma

This paper describes the LDC forced aligner which was designed to align audio and transcripts. Unlike existing forced aligners, LDC forced aligner can align partially transcribed audio files, and also audio files with la…

SentenceSpeech RecognitionSpeech SynthesisText-To-Speech Synthesis

FASA: a Flexible and Automatic Speech Aligner for Extracting High-quality Aligned Children Speech Data

2024-06-25 · Dancheng Liu, JinJun Xiong

Automatic Speech Recognition (ASR) for adults' speeches has made significant progress by employing deep neural network (DNN) models recently, but improvement in children's speech is still unsatisfactory due to children's…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Jochre 3 and the Yiddish OCR corpus

2025-01-14 · Assaf Urieli, Amber Clooney, Michelle Sigiel, Grisha Leyfer

We describe the construction of a publicly available Yiddish OCR Corpus, and describe and evaluate the open source OCR tool suite Jochre 3, including an Alto editor for corpus annotation, OCR software for Alto OCR layer …

Optical Character Recognition (OCR)