paper-with-me

홈 › Papers

Automatic Translating between Ancient Chinese and Contemporary Chinese with Limited Aligned Corpora

2018-03-05 · Zhiyuan Zhang, Wei Li, Qi Su

The Chinese language has evolved a lot during the long-term development. Therefore, native speakers now have trouble in reading sentences written in ancient Chinese. In this paper, we propose to build an end-to-end neural model to automatically translate between ancient and contemporary Chinese. However, the existing ancient-contemporary Chinese parallel corpora are not aligned at the sentence level and sentence-aligned corpora are limited, which makes it difficult to train the model. To build the sentence level parallel training data for the model, we propose an unsupervised algorithm that constructs sentence-aligned ancient-contemporary pairs by using the fact that the aligned sentence pair shares many of the tokens. Based on the aligned corpus, we propose an end-to-end neural model with copying mechanism and local attention to translate between ancient and contemporary Chinese. Experiments show that the proposed unsupervised algorithm achieves 99.4% F1 score for sentence alignment, and the translation model achieves 26.95 BLEU from ancient to contemporary, and 36.34 BLEU from contemporary to ancient.

📄 PDF Abstract BibTeX arXiv:1803.01557

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceTranslation

Similar Papers 제목 키워드 기반

Exploring the Capabilities of ChatGPT in Ancient Chinese Translation and Person Name Recognition

2023-12-23 · Shijing Si, Siqing Zhou, Le Tang, Xiaoqing Cheng 외

ChatGPT's proficiency in handling modern standard languages suggests potential for its use in understanding ancient Chinese. This paper explores ChatGPT's capabilities on ancient Chinese via two tasks: translating ancien…

Translation

Ancient-Modern Chinese Translation with a Large Training Dataset

2018-08-11 · Dayiheng Liu, Jiancheng Lv, Kexin Yang, Qian Qu

Ancient Chinese brings the wisdom and spirit culture of the Chinese nation. Automatic translation from ancient Chinese to modern Chinese helps to inherit and carry forward the quintessence of the ancients. However, the l…

Cultural Vocal Bursts Intensity PredictionMachine TranslationNMTTranslation

BERT 4EVER@EvaHan 2022: Ancient Chinese Word Segmentation and Part-of-Speech Tagging Based on Adversarial Learning and Continual Pre-training

2022-06-01 · LT4HALA (LREC) 2022 6 · Hailin Zhang, Ziyu Yang, Yingwen Fu, Ruoyao Ding

With the development of artificial intelligence (AI) and digital humanities, ancient Chinese resources and language technology have also developed and grown, which have become an increasingly important part to the study …

Chinese Word SegmentationCultural Vocal Bursts Intensity PredictionEnsemble LearningPart-Of-Speech Tagging+3

Kanbun-LM: Reading and Translating Classical Chinese in Japanese Methods by Language Models

2023-05-22 · Hao Wang, Hirofumi Shimizu, Daisuke Kawahara

Recent studies in natural language processing (NLP) have focused on modern languages and achieved state-of-the-art results in many tasks. Meanwhile, little attention has been paid to ancient texts and related tasks. Clas…

Machine Translation

Data Augmentation for Low-resource Word Segmentation and POS Tagging of Ancient Chinese Texts

2022-06-01 · LT4HALA (LREC) 2022 6 · Yutong Shen, Jiahuan Li, ShuJian Huang, Yi Zhou 외

Automatic word segmentation and part-of-speech tagging of ancient books can help relevant researchers to study ancient texts. In recent years, pre-trained language models have achieved significant improvements on text pr…

Data AugmentationLanguage ModelingLanguage ModellingPart-Of-Speech Tagging+2