paper-with-me

홈 › Papers

CINO: A Chinese Minority Pre-trained Language Model

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Multilingual pre-trained language models have shown impressive performance on cross-lingual tasks. It greatly facilitates the applications of natural language processing on low-resource languages. However, there are still some languages that the existing multilingual model does not perform well on. In this paper, we propose CINO (Chinese Minority Pre-trained Language Model), a multilingual pre-trained language model for Chinese minority languages. It covers Standard Chinese, Cantonese, and six other Chinese minority languages. To evaluate the cross-lingual ability of the multilingual models on the minority languages, we collect documents from Wikipedia and build a text classification dataset WCM (Wiki-Chinese-Minority). We test CINO on WCM and two other text classification tasks. Experiments show that CINO outperforms the baselines notably. The CINO model and the WCM dataset will be made publicly available.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modellingmodeltext-classificationText Classification

Similar Papers 제목 키워드 기반

CINO: A Chinese Minority Pre-trained Language Model

2022-02-28 · COLING 2022 10 · Ziqing Yang, Zihang Xu, Yiming Cui, Baoxin Wang 외

Multilingual pre-trained language models have shown impressive performance on cross-lingual tasks. It greatly facilitates the applications of natural language processing on low-resource languages. However, there are stil…

Language ModelingLanguage Modellingmodeltext-classification+1

Pay Attention to the Robustness of Chinese Minority Language Models! Syllable-level Textual Adversarial Attack on Tibetan Script

2024-12-03 · Xi Cao, Dolma Dawa, Nuo Qun, Trashi Nyima

The textual adversarial attack refers to an attack method in which the attacker adds imperceptible perturbations to the original texts by elaborate design so that the NLP (natural language processing) model produces fals…

Adversarial Attack

Do Chinese models speak Chinese languages?

2025-03-31 · Andrea W Wen-Yi, Unso Eun Seo Jo, David Mimno

The release of top-performing open-weight LLMs has cemented China's role as a leading force in AI development. Do these models support languages spoken in China? Or do they speak the same languages as Western models? Com…

Reading Comprehension

CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China

2025-09-12 · Guixian Xu, Zeli Su, Ziyin Zhang, Jianing Liu 외 arxiv

Minority languages in China, such as Tibetan, Uyghur, and Traditional Mongolian, face significant challenges due to their unique writing systems, which differ from international standards. This discrepancy has led to a s…

Headline Generation

Multi-Granularity Tibetan Textual Adversarial Attack Method Based on Masked Language Model

2024-12-03 · Xi Cao, Nuo Qun, Quzong Gesang, Yulei Zhu 외

In social media, neural network models have been applied to hate speech detection, sentiment analysis, etc., but neural network models are susceptible to adversarial attacks. For instance, in a text classification task, …

Adversarial AttackHate Speech DetectionLanguage ModelingLanguage Modelling+3