paper-with-me

Papers

A Robust Self-Learning Framework for Cross-Lingual Text Classification

2019-11-01 · IJCNLP 2019 11 · Xin Dong, Gerard de Melo

Based on massive amounts of data, recent pretrained contextual representation models have made significant strides in advancing a number of different English NLP tasks. However, for other languages, relevant training data may be lacking, while state-of-the-art deep learning methods are known to be data-hungry. In this paper, we present an elegantly simple robust self-learning framework to include unlabeled non-English samples in the fine-tuning process of pretrained multilingual representation models. We leverage a multilingual model{'}s own predictions on unlabeled non-English data in order to obtain additional information that can be used during further fine-tuning. Compared with original multilingual models and other cross-lingual classification models, we observe significant gains in effectiveness on document and sentiment classification for a range of diverse languages.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral ClassificationSelf-LearningSentiment AnalysisSentiment Classificationtext-classificationText Classification

Similar Papers 제목 키워드 기반

HMS-BERT: Hybrid Multi-Task Self-Training for Multilingual and Multi-Label Cyberbullying Detection

2026-03-13 · Zixin Feng, Xinying Cui, Yifan Sun, Zheng Wei 외 arxiv

Cyberbullying on social media is inherently multilingual and multi-faceted, where abusive behaviors often overlap across multiple categories. Existing methods are commonly limited by monolingual assumptions or single-tas…

Borrowing Human Senses: Comment-Aware Self-Training for Social Media Multimodal Classification

2023-03-27 · Chunpu Xu, Jing Li

Social media is daily creating massive multimedia content with paired image and text, presenting the pressing need to automate the vision and language understanding for various multimodal classification tasks. Compared t…

ClassificationHate Speech DetectionRelation ClassificationSarcasm Detection+2

SentiXRL: An advanced large language Model Framework for Multilingual Fine-Grained Emotion Classification in Complex Text Environment

2024-11-27 · Jie Wang, Yichen Wang, Zhilin Zhang, Jianhao Zeng 외

With strong expressive capabilities in Large Language Models(LLMs), generative models effectively capture sentiment structures and deep semantics, however, challenges remain in fine-grained sentiment classification acros…

ClassificationDecision MakingEmotion ClassificationLanguage Modeling+6

AfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African Languages

2022-11-07 · Bonaventure F. P. Dossou, Atnafu Lambebo Tonja, Oreen Yousuf, Salomey Osei 외

In recent years, multilingual pre-trained language models have gained prominence due to their remarkable performance on numerous downstream Natural Language Processing tasks (NLP). However, pre-training these large multi…

Active LearningLanguage ModelingLanguage ModellingNER+3

Leveraging Adversarial Training in Self-Learning for Cross-Lingual Text Classification

2020-07-29 · Xin Dong, Yaxin Zhu, Yupeng Zhang, Zuohui Fu 외

In cross-lingual text classification, one seeks to exploit labeled data from one language to train a text classification model that can then be applied to a completely different language. Recent multilingual representati…

ClassificationGeneral Classificationintent-classificationIntent Classification+3