paper-with-me

홈 › Papers

RoChBert: Towards Robust BERT Fine-tuning for Chinese

2022-10-28 · Zihan Zhang, Jinfeng Li, Ning Shi, Bo Yuan, Xiangyu Liu, Rong Zhang, Hui Xue, Donghong Sun, Chao Zhang

Despite of the superb performance on a wide range of tasks, pre-trained language models (e.g., BERT) have been proved vulnerable to adversarial texts. In this paper, we present RoChBERT, a framework to build more Robust BERT-based models by utilizing a more comprehensive adversarial graph to fuse Chinese phonetic and glyph features into pre-trained representations during fine-tuning. Inspired by curriculum learning, we further propose to augment the training dataset with adversarial texts in combination with intermediate samples. Extensive experiments demonstrate that RoChBERT outperforms previous methods in significant ways: (i) robust -- RoChBERT greatly improves the model robustness without sacrificing accuracy on benign texts. Specifically, the defense lowers the success rates of unlimited and limited attacks by 59.43% and 39.33% respectively, while remaining accuracy of 93.30%; (ii) flexible -- RoChBERT can easily extend to various language models to solve different downstream tasks with excellent performance; and (iii) efficient -- RoChBERT can be directly applied to the fine-tuning stage without pre-training language model from scratch, and the proposed data augmentation method is also low-cost.

📄 PDF Abstract BibTeX arXiv:2210.15944

Code (1)

zzh-z/rochbert 공식 구현

Tasks

Data AugmentationLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Chinese ModernBERT with Whole-Word Masking

2025-10-14 · Zeyu Zhao, Ningtao Wang, Xing Fu, Yu Cheng arxiv

Encoder-only Transformers have advanced along three axes -- architecture, data, and systems -- yielding Pareto gains in accuracy, speed, and memory efficiency. Yet these improvements have not fully transferred to Chinese…

From BERT to LLMs: Comparing and Understanding Chinese Classifier Prediction in Language Models

2025-08-25 · Ziqi Zhang, Jianfei Ma, Emmanuele Chersoni, Jieshun You 외 arxiv

Classifiers are an important and defining feature of the Chinese language, and their correct prediction is key to numerous educational applications. Yet, whether the most popular Large Language Models (LLMs) possess prop…

Entity Enhanced BERT Pre-training for Chinese NER

2020-11-01 · EMNLP 2020 11 · Chen Jia, Yuefeng Shi, Qinrong Yang, Yue Zhang

Character-level BERT pre-trained in Chinese suffers a limitation of lacking lexicon information, which shows effectiveness for Chinese NER. To integrate the lexicon into pre-trained LMs for Chinese NER, we investigate a …

NER

RAC-BERT: Character Radical Enhanced BERT for Ancient Chinese

2023-10-08 · journal 2023 10 · Lifan Han, Xin Wang, Meng Wang, Zhao Li 외

In recent years, Chinese pre-training language models have achieved significant improvements in the fields, such as natural language understanding (NLU) and text generation. However, most of these existing pre-trained la…

Natural Language UnderstandingText Generation

Cross-Dataset Stability of Expert-Informed Skill Prompting and Fine-Tuning for Chinese Metaphor Identification

2026-08-26 · Yufeng Wu, Meichun Liu arxiv

Metaphor-identification performance can change markedly across datasets that differ in text distribution and annotation policy. We examine whether a fixed expert-informed procedure produces a more even cross-dataset prof…