paper-with-me

홈 › Papers

FROST-EMA: Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography Measurements with L1, L2 and Imitated L2 Accents

2025-06-10 · Satu Hopponen, Tomi Kinnunen, Alexandre Nikolaev, Rosa González Hautamäki, Lauri Tavi, Einar Meister

We introduce a new FROST-EMA (Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography) corpus. It consists of 18 bilingual speakers, who produced speech in their native language (L1), second language (L2), and imitated L2 (fake foreign accent). The new corpus enables research into language variability from phonetic and technological points of view. Accordingly, we include two preliminary case studies to demonstrate both perspectives. The first case study explores the impact of L2 and imitated L2 on the performance of an automatic speaker verification system, while the second illustrates the articulatory patterns of one speaker in L1, L2, and a fake accent.

📄 PDF Abstract BibTeX arXiv:2506.08981

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Verification

Similar Papers 제목 키워드 기반

Perceptual and acoustic analysis of voice similarities between parents and young children

2019-09-01 · WS (NoDaLiDa) 2019 9 · Evgeniia Rykova, Stefan Werner

Human voice provides the means for verbal communication and forms a part of personal identity. Due to genetic and environmental factors, a voice of a child should resemble the voice of her parent(s), but voice similariti…

MaSS: A Large and Clean Multilingual Corpus of Sentence-aligned Spoken Utterances Extracted from the Bible

2019-07-30 · LREC 2020 5 · Marcely Zanon Boito, William N. Havard, Mahault Garnerin, Éric Le Ferrand 외

The CMU Wilderness Multilingual Speech Dataset (Black, 2019) is a newly published multilingual speech dataset based on recorded readings of the New Testament. It provides data to build Automatic Speech Recognition (ASR) …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)RetrievalSentence+5

Unsupervised Word Segmentation from Discrete Speech Units in Low-Resource Settings

2021-06-08 · SIGUL (LREC) 2022 6 · Marcely Zanon Boito, Bolaji Yusuf, Lucas Ondel, Aline Villavicencio 외

Documenting languages helps to prevent the extinction of endangered dialects, many of which are otherwise expected to disappear by the end of the century. When documenting oral languages, unsupervised word segmentation (…

Linked Multi-Model Data on Russian Domestic and Foreign Policy Speeches

2026-05-15 · Daria Blinova, Gayathri Emuru, Rakesh Emuru, Kushagradheer Shridheer Srivastava 외 arxiv

This paper introduces a dataset of interlinked multimodal political communications from the Russian government, addressing persistent deficiencies in the availability of social text- and image-based data for authoritaria…

Hybrid Physics-ML Framework for Pan-Arctic Permafrost Infrastructure Risk at Record 2.9-Million Observation Scale

2025-10-02 · Boris Kriuk arxiv

Arctic warming threatens over 100 billion in permafrost-dependent infrastructure across Northern territories, yet existing risk assessment frameworks lack spatiotemporal validation, uncertainty quantification, and operat…