paper-with-me

홈 › Papers

Ensemble Methods to Distinguish Mainland and Taiwan Chinese

2019-06-01 · WS 2019 6 · Hai Hu, Wen Li, He Zhou, Zuoyu Tian, Yiwen Zhang, Liang Zou

This paper describes the IUCL system at VarDial 2019 evaluation campaign for the task of discriminating between Mainland and Taiwan variation of mandarin Chinese. We first build several base classifiers, including a Naive Bayes classifier with word n-gram as features, SVMs with both character and syntactic features, and neural networks with pre-trained character/word embeddings. Then we adopt ensemble methods to combine output from base classifiers to make final predictions. Our ensemble models achieve the highest F1 score (0.893) in simplified Chinese track and the second highest (0.901) in traditional Chinese track. Our results demonstrate the effectiveness and robustness of the ensemble methods.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Word Embeddings

Similar Papers 제목 키워드 기반

A Topic-aware Comparable Corpus of Chinese Variations

2024-11-17 · Da-Chen Lian, Shu-Kai Hsieh

This study aims to fill the gap by constructing a topic-aware comparable corpus of Mainland Chinese Mandarin and Taiwanese Mandarin from the social media in Mainland China and Taiwan, respectively. Using Dcard for Taiwan…

Naive Bayes and BiLSTM Ensemble for Discriminating between Mainland and Taiwan Variation of Mandarin Chinese

2019-06-01 · WS 2019 6 · Li Yang, Yang Xiang

Automatic dialect identification is a more challengingctask than language identification, as it requires the ability to discriminate between varieties of one language. In this paper, we propose an ensemble based system, …

Dialect IdentificationLanguage IdentificationWord Embeddings

Cross-strait Variations on Two Near-synonymous Loanwords xie2shang1 and tan2pan4: A Corpus-based Comparative Study

2022-10-09 · Yueyue Huang, Chu-Ren Huang

This study attempts to investigate cross-strait variations on two typical synonymous loanwords in Chinese, i.e. xie2shang1 and tan2pan4, drawn on MARVS theory. Through a comparative analysis, the study found some distrib…

De-verbalization and Nominal Categories in Mandarin Chinese: A corpus-driven study in both Mainland Mandarin and Taiwan Mandarin

2015-10-01 · PACLIC 2015 10 · Jiajuan Xiong, Chu-Ren Huang

Using Contextually Aligned Online Reviews to Measure LLMs' Performance Disparities Across Language Varieties

2025-02-10 · Zixin Tang, Chieh-Yang Huang, Tsung-Che Li, Ho Yin Sam Ng 외

A language can have different varieties. These varieties can affect the performance of natural language processing (NLP) models, including large language models (LLMs), which are often trained on data from widely spoken …

Sentiment Analysis