paper-with-me

홈 › Papers

Sentence-level dialects identification in the greater China region

2017-01-08 · Fan Xu, Mingwen Wang, Maoxi Li

Identifying the different varieties of the same language is more challenging than unrelated languages identification. In this paper, we propose an approach to discriminate language varieties or dialects of Mandarin Chinese for the Mainland China, Hong Kong, Taiwan, Macao, Malaysia and Singapore, a.k.a., the Greater China Region (GCR). When applied to the dialects identification of the GCR, we find that the commonly used character-level or word-level uni-gram feature is not very efficient since there exist several specific problems such as the ambiguity and context-dependent characteristic of words in the dialects of the GCR. To overcome these challenges, we use not only the general features like character-level n-gram, but also many new word-level features, including PMI-based and word alignment-based features. A series of evaluation results on both the news and open-domain dataset from Wikipedia show the effectiveness of the proposed approach.

📄 PDF Abstract BibTeX arXiv:1701.01908

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceWord Alignment

Similar Papers 제목 키워드 기반

ArbDialectID at MADAR Shared Task 1: Language Modelling and Ensemble Learning for Fine Grained Arabic Dialect Identification

2019-08-01 · WS 2019 8 · Kathrein Abu Kwaik, Motaz Saad

In this paper, we present a Dialect Identification system (ArbDialectID) that competed at Task 1 of the MADAR shared task, MADARTravel Domain Dialect Identification. We build a course and a fine-grained identification mo…

Dialect IdentificationEnsemble LearningFeature EngineeringLanguage Modelling+1

Dialects Identification of Armenian Language

2022-06-01 · DigitAm (LREC) 2022 6 · Karen Avetisyan

The Armenian language has many dialects that differ from each other syntactically, morphologically, and phonetically. In this work, we implement and evaluate models that determine the dialect of a given passage of text. …

Dialect IdentificationLanguage IdentificationSentenceWord Embeddings

Simple But Not Na\"\ive: Fine-Grained Arabic Dialect Identification Using Only N-Grams

2019-08-01 · WS 2019 8 · Sohaila Eltanbouly, May Bashendy, Tamer Elsayed

This paper presents the participation of Qatar University team in MADAR shared task, which addresses the problem of sentence-level fine-grained Arabic Dialect Identification over 25 different Arabic dialects in addition …

Dialect IdentificationSentence

Phonemic evidence reveals interwoven evolution of Chinese dialects

2018-02-16

Han Chinese experienced substantial population migrations and admixture in history, yet little is known about the evolutionary process of Chinese dialects. Here, we used phylogenetic approaches and admixture inference to…

Diversity

The Unreasonable Effectiveness of Machine Learning in Moldavian versus Romanian Dialect Identification

2020-07-30 · Mihaela Găman, Radu Tudor Ionescu

Motivated by the seemingly high accuracy levels of machine learning models in Moldavian versus Romanian dialect identification and the increasing research interest on this topic, we provide a follow-up on the Moldavian v…

ArticlesBIG-bench Machine LearningDialect IdentificationEnsemble Learning+1