paper-with-me

홈 › Papers

ToxiTrace: Gradient-Aligned Training for Explainable Chinese Toxicity Detection

2026-04-14 · Boyang Li, Hongzhe Shou, Yuanyuan Liang, Jingbin Zhang, Fang Zhou arxiv

Existing Chinese toxic content detection methods mainly target sentence-level classification but often fail to provide readable and contiguous toxic evidence spans. We propose \textbf{ToxiTrace}, an explainability-oriented method for BERT-style encoders with three components: (1) \textbf{CuSA}, which refines encoder-derived saliency cues into fine-grained toxic spans with lightweight LLM guidance; (2) \textbf{GCLoss}, a gradient-constrained objective that concentrates token-level saliency on toxic evidence while suppressing irrelevant activations; and (3) \textbf{ARCL}, which constructs sample-specific contrastive reasoning pairs to sharpen the semantic boundary between toxic and non-toxic content. Experiments show that ToxiTrace improves classification accuracy and toxic span extraction while preserving efficient encoder-based inference and producing more coherent, human-readable explanations. We have released the model at https://huggingface.co/ArdLi/ToxiTrace.

📄 PDF Abstract BibTeX arXiv:2604.12321

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Automatic Translating between Ancient Chinese and Contemporary Chinese with Limited Aligned Corpora

2018-03-05 · Zhiyuan Zhang, Wei Li, Qi Su

The Chinese language has evolved a lot during the long-term development. Therefore, native speakers now have trouble in reading sentences written in ancient Chinese. In this paper, we propose to build an end-to-end neura…

SentenceTranslation

Sociocultural Norm Similarities and Differences via Situational Alignment and Explainable Textual Entailment

2023-05-23 · Sky CH-Wang, Arkadiy Saakyan, Oliver Li, Zhou Yu 외

Designing systems that can reason across cultures requires that they are grounded in the norms of the contexts in which they operate. However, current research on developing computational models of social norms has prima…

DescriptiveIn-Context LearningNatural Language Inference

ChLogic: Evaluating Robustness of Logical Reasoning in Chinese Expressions

2026-06-16 · Peixian Zhou, Yuxu Chen, Chaorui Zhang, Wei Han 외 arxiv

Large language models perform increasingly well on standardized logical reasoning benchmarks, but whether this ability remains robust beyond English is unclear. We introduce ChLogic, an English--Chinese aligned benchmark…

Logical Reasoning

An Aligned French-Chinese corpus of 10K segments from university educational material

2016-12-01 · WS 2016 12 · Ruslan Kalitvianski, Lingxiao Wang, Val{\'e}rie Bellynck, Christian Boitet

This paper describes a corpus of nearly 10K French-Chinese aligned segments, produced by post-editing machine translated computer science courseware. This corpus was built from 2013 to 2016 within the PROJECT{\_}NAME pro…

Machine TranslationTranslation

GCRC: A New Challenging MRC Dataset from Gaokao Chinese for Explainable Evaluation

2021-08-01 · Findings (ACL) 2021 8 · Hongye Tan, Xiaoyue Wang, Yu Ji, Ru Li 외