paper-with-me

홈 › Papers

On the Idiosyncrasies of the Mandarin Chinese Classifier System

2019-02-26 · NAACL 2019 6 · Shijia Liu, Hongyuan Mei, Adina Williams, Ryan Cotterell

While idiosyncrasies of the Chinese classifier system have been a richly studied topic among linguists (Adams and Conklin, 1973; Erbaugh, 1986; Lakoff, 1986), not much work has been done to quantify them with statistical methods. In this paper, we introduce an information-theoretic approach to measuring idiosyncrasy; we examine how much the uncertainty in Mandarin Chinese classifiers can be reduced by knowing semantic information about the nouns that the classifiers modify. Using the empirical distribution of classifiers from the parsed Chinese Gigaword corpus (Graff et al., 2005), we compute the mutual information (in bits) between the distribution over classifiers and distributions over other linguistic quantities. We investigate whether semantic classes of nouns and adjectives differ in how much they reduce uncertainty in classifier choice, and find that it is not fully idiosyncratic; while there are no obvious trends for the majority of semantic classes, shape nouns reduce uncertainty in classifier choice the most.

📄 PDF Abstract BibTeX arXiv:1902.10193

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sina Mandarin Alphabetical Words:A Web-driven Code-mixing Lexical Resource

2020-12-01 · Asian Chapter of the Association for Computational Linguistics 2020 · Rong Xiang, Mingyu Wan, Qi Su, Chu-Ren Huang 외

Mandarin Alphabetical Word (MAW) is one indispensable component of Modern Chinese that demonstrates unique code-mixing idiosyncrasies influenced by language exchanges. Yet, this interesting phenomenon has not been proper…

Ensemble Methods to Distinguish Mainland and Taiwan Chinese

2019-06-01 · WS 2019 6 · Hai Hu, Wen Li, He Zhou, Zuoyu Tian 외

This paper describes the IUCL system at VarDial 2019 evaluation campaign for the task of discriminating between Mainland and Taiwan variation of mandarin Chinese. We first build several base classifiers, including a Naiv…

Word Embeddings

Comparing Theories of Speaker Choice Using a Model of Classifier Production in Mandarin Chinese

2018-06-01 · NAACL 2018 6 · Meilin Zhan, Roger Levy

Speakers often have more than one way to express the same meaning. What general principles govern speaker choice in the face of optionality when near semantically invariant alternation exists? Studies have shown that opt…

A Polyphone BERT for Polyphone Disambiguation in Mandarin Chinese

2022-07-01 · Song Zhang, Ken Zheng, Xiaoxu Zhu, Baoxiang Li

Grapheme-to-phoneme (G2P) conversion is an indispensable part of the Chinese Mandarin text-to-speech (TTS) system, and the core of G2P conversion is to solve the problem of polyphone disambiguation, which is to pick up t…

Polyphone disambiguationtext-to-speechText to Speech

CLiMP: A Benchmark for Chinese Language Model Evaluation

2021-01-26 · EACL 2021 2 · Beilei Xiang, Changbing Yang, Yu Li, Alex Warstadt 외

Linguistically informed analyses of language models (LMs) contribute to the understanding and improvement of these models. Here, we introduce the corpus of Chinese linguistic minimal pairs (CLiMP), which can be used to i…

Language Model EvaluationLanguage ModelingLanguage Modellingmodel