paper-with-me

홈 › Papers

Recovering dialect geography from an unaligned comparable corpus

2012-04-01 · WS 2012 4 · Yves Scherrer
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

RegSpeech12: A Regional Corpus of Bengali Spontaneous Speech Across Dialects

2025-10-28 · Md. Rezuwan Hassan, Azmol Hossain, Kanij Fatema, Rubayet Sabbir Faruque 외 arxiv

The Bengali language, spoken extensively across South Asia and among diasporic communities, exhibits considerable dialectal diversity shaped by geography, culture, and history. Phonological and pronunciation-based classi…

Speech Recognition

ACTIV-ES: a comparable, cross-dialect corpus of `everyday' Spanish from Argentina, Mexico, and Spain

2014-05-01 · LREC 2014 5 · Jerid Francom, Mans Hulden, Adam Ussishkin

Corpus resources for Spanish have proved invaluable for a number of applications in a wide variety of fields. However, a majority of resources are based on formal, written language and/or are not built to model language …

Part-Of-Speech Tagging

WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing

2025-09-22 · Yuhang Dai, Ziyu Zhang, Shuai Wang, Longhao Li 외 arxiv

The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects of Chinese. To address this critical gap…

Bhasacitra: Visualising the dialect geography of South Asia

2021-05-28 · Aryaman Arora, Adam Farris, Gopalakrishnan R, Samopriya Basu

We present Bhasacitra, a dialect mapping system for South Asia built on a database of linguistic studies of languages of the region annotated for topic and location data. We analyse language coverage and look towards app…

Bhāṣācitra: Visualising the dialect geography of South Asia

2021-08-01 · ACL (LChange) 2021 8 · Aryaman Arora, Adam Farris, Gopalakrishnan R, Samopriya Basu

We present Bhāṣācitra, a dialect mapping system for South Asia built on a database of linguistic studies of languages of the region annotated for topic and location data. We analyse language coverage and look towards app…