paper-with-me

Papers

Native Chinese Reader: A Dataset Towards Native-Level Chinese Machine Reading Comprehension

2021-12-13 · Shusheng Xu, Yichen Liu, Xiaoyu Yi, Siyuan Zhou, Huizi Li, Yi Wu

We present Native Chinese Reader (NCR), a new machine reading comprehension (MRC) dataset with particularly long articles in both modern and classical Chinese. NCR is collected from the exam questions for the Chinese course in China's high schools, which are designed to evaluate the language proficiency of native Chinese youth. Existing Chinese MRC datasets are either domain-specific or focusing on short contexts of a few hundreds of characters in modern Chinese only. By contrast, NCR contains 8390 documents with an average length of 1024 characters covering a wide range of Chinese writing styles, including modern articles, classical literature and classical poetry. A total of 20477 questions on these documents also require strong reasoning abilities and common sense to figure out the correct answers. We implemented multiple baseline models using popular Chinese pre-trained models and additionally launched an online competition using our dataset to examine the limit of current methods. The best model achieves 59% test accuracy while human evaluation shows an average accuracy of 79%, which indicates a significant performance gap between current MRC models and native Chinese speakers. We release the dataset at https://sites.google.com/view/native-chinese-reader/.

📄 PDF Abstract BibTeX arXiv:2112.06494

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesCommon Sense ReasoningMachine Reading ComprehensionReading Comprehension

Similar Papers 제목 키워드 기반

Japanese Lexical Complexity for Non-Native Readers: A New Dataset

2023-06-30 · Yusuke Ide, Masato Mita, Adam Nohejl, Hiroki Ouchi 외

Lexical complexity prediction (LCP) is the task of predicting the complexity of words in a text on a continuous scale. It plays a vital role in simplifying or annotating complex words to assist readers. To study lexical …

Lexical Complexity Prediction

Word Complexity is in the Eye of the Beholder

2021-06-01 · NAACL 2021 4 · Sian Gooding, Ekaterina Kochmar, Seid Muhie Yimam, Chris Biemann

Lexical complexity is a highly subjective notion, yet this factor is often neglected in lexical simplification and readability systems which use a {''}one-size-fits-all{''} approach. In this paper, we investigate which a…

Lexical Simplification

Detection of Chinese Word Usage Errors for Non-Native Chinese Learners with Bidirectional LSTM

2017-07-01 · ACL 2017 7 · Yow-Ting Shiue, Hen-Hsen Huang, Hsin-Hsi Chen

Selecting appropriate words to compose a sentence is one common problem faced by non-native Chinese learners. In this paper, we propose (bidirectional) LSTM sequence labeling models and explore various features to detect…

Grammatical Error DetectionPOSPositionSentence

CSCD-NS: a Chinese Spelling Check Dataset for Native Speakers

2022-11-16 · Yong Hu, Fandong Meng, Jie zhou

In this paper, we present CSCD-NS, the first Chinese spelling check (CSC) dataset designed for native speakers, containing 40,000 samples from a Chinese social platform. Compared with existing CSC datasets aimed at Chine…

Spelling Correction

CCTC: A Cross-Sentence Chinese Text Correction Dataset for Native Speakers

2022-10-01 · COLING 2022 10 · Baoxin Wang, Xingyi Duan, Dayong Wu, Wanxiang Che 외

The Chinese text correction (CTC) focuses on detecting and correcting Chinese spelling errors and grammatical errors. Most existing datasets of Chinese spelling check (CSC) and Chinese grammatical error correction (GEC) …

Grammatical Error CorrectionSentence