paper-with-me

홈 › Papers

How Chinese are Chinese Language Models? The Puzzling Lack of Language Policy in China's LLMs

2024-07-12 · Andrea W Wen-Yi, Unso Eun Seo Jo, Lu Jia Lin, David Mimno

Contemporary language models are increasingly multilingual, but Chinese LLM developers must navigate complex political and business considerations of language diversity. Language policy in China aims at influencing the public discourse and governing a multi-ethnic society, and has gradually transitioned from a pluralist to a more assimilationist approach since 1949. We explore the impact of these influences on current language technology. We evaluate six open-source multilingual LLMs pre-trained by Chinese companies on 18 languages, spanning a wide range of Chinese, Asian, and Anglo-European languages. Our experiments show Chinese LLMs performance on diverse languages is indistinguishable from international LLMs. Similarly, the models' technical reports also show lack of consideration for pretraining data language coverage except for English and Mandarin Chinese. Examining Chinese AI policy, model experiments, and technical reports, we find no sign of any consistent policy, either for or against, language diversity in China's LLM development. This leaves a puzzling fact that while China regulates both the languages people use daily as well as language model development, they do not seem to have any policy on the languages in language models.

📄 PDF Abstract BibTeX arXiv:2407.09652

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityLanguage ModellingNavigate

Similar Papers 제목 키워드 기반

CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation

2024-07-01 · Yuxuan Wang, Yijun Liu, Fei Yu, Chen Huang 외

Despite the rapid development of Chinese vision-language models (VLMs), most existing Chinese vision-language (VL) datasets are constructed on Western-centric images from existing English VL datasets. The cultural bias i…

Image-text RetrievalQuestion AnsweringText RetrievalVisual Grounding+1

The CUHK Discourse TreeBank for Chinese: Annotating Explicit Discourse Connectives for the Chinese TreeBank

2014-05-01 · LREC 2014 5 · Lanjun Zhou, Binyang Li, Zhongyu Wei, Kam-Fai Wong

The lack of open discourse corpus for Chinese brings limitations for many natural language processing tasks. In this work, we present the first open discourse treebank for Chinese, namely, the Discourse Treebank for Chin…

Part-Of-Speech TaggingQuestion AnsweringSentenceSentence Compression+2

SCTB: A Chinese Treebank in Scientific Domain

2016-12-01 · WS 2016 12 · Chenhui Chu, Toshiaki Nakazawa, Daisuke Kawahara, Sadao Kurohashi

Treebanks are curial for natural language processing (NLP). In this paper, we present our work for annotating a Chinese treebank in scientific domain (SCTB), to address the problem of the lack of Chinese treebanks in thi…

Chinese Word SegmentationMachine TranslationTranslation

Modern Chinese Helps Archaic Chinese Processing: Finding and Exploiting the Shared Properties

2014-05-01 · LREC 2014 5 · Yan Song, Fei Xia

Languages change over time and ancient languages have been studied in linguistics and other related fields. A main challenge in this research area is the lack of empirical data; for instance, ancient spoken languages oft…

ArticlesPOSSentiment Analysis

Image and Information

2016-02-03 · Frank Nielsen

A well-known old adage says that {\em "A picture is worth a thousand words!"} (attributed to the Chinese philosopher Confucius ca 500 years BC). But more precisely, what do we mean by information in images? And how can i…