paper-with-me

홈 › Papers

ACTIV-ES: a comparable, cross-dialect corpus of `everyday' Spanish from Argentina, Mexico, and Spain

2014-05-01 · LREC 2014 5 · Jerid Francom, Mans Hulden, Adam Ussishkin

Corpus resources for Spanish have proved invaluable for a number of applications in a wide variety of fields. However, a majority of resources are based on formal, written language and/or are not built to model language variation between varieties of the Spanish language, despite the fact that most language in ‘everyday’ use is informal/ dialogue-based and shows rich regional variation. This paper outlines the development and evaluation of the ACTIV-ES corpus, a first-step to produce a comparable, cross-dialect corpus representative of the ‘everyday’ language of various regions of the Spanish-speaking world.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Part-Of-Speech Tagging

Similar Papers 제목 키워드 기반

Recovering dialect geography from an unaligned comparable corpus

2012-04-01 · WS 2012 4 · Yves Scherrer
Machine Translation

WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing

2025-09-22 · Yuhang Dai, Ziyu Zhang, Shuai Wang, Longhao Li 외 arxiv

The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects of Chinese. To address this critical gap…

ArchiMob - A Corpus of Spoken Swiss German

2016-05-01 · LREC 2016 5 · Tanja Samard{\v{z}}i{\'c}, Yves Scherrer, Elvira Glaser

Swiss dialects of German are, unlike most dialects of well standardised languages, widely used in everyday communication. Despite this fact, automatic processing of Swiss German is still a considerable challenge due to t…

Machine TranslationPart-Of-Speech TaggingTranslation

Survey of Conversational Behavior: Towards the Design of a Balanced Corpus of Everyday Japanese Conversation

2016-05-01 · LREC 2016 5 · Hanae Koiso, Tomoyuki Tsuchiya, Ryoko Watanabe, Daisuke Yokomori 외

In 2016, we set about building a large-scale corpus of everyday Japanese conversation―a collection of conversations embedded in naturally occurring activities in daily life. We will collect more than 200 hours of recor…

Survey

BIDWESH: A Bangla Regional Based Hate Speech Detection Dataset

2025-07-22 · Azizul Hakim Fayaz, MD. Shorif Uddin, Rayhan Uddin Bhuiyan, Zakia Sultana 외 arxiv

Hate speech on digital platforms has become a growing concern globally, especially in linguistically diverse countries like Bangladesh, where regional dialects play a major role in everyday communication. Despite progres…

Hate Speech Detection