paper-with-me

Papers XLM-R

“XLM-R” 태그가 달린 논문 221편 · 필터 해제

Cross-Linguistic Transfer in Multilingual NLP: The Role of Language Families and Morphology

2025-05-20 · Ajitesh Bankula, Praney Bankula

Cross-lingual transfer has become a crucial aspect of multilingual NLP, as it allows for models trained on resource-rich languages to be applied to low-resource languages more effectively. Recently massively multilingual…

Cross-Lingual TransferMultilingual NLPXLM-R

Subasa -- Adapting Language Models for Low-resourced Offensive Language Detection in Sinhala

2025-04-02 · Shanilka Haturusinghe, Tharindu Cyril Weerasooriya, Marcos Zampieri, Christopher M. Homan 외

Accurate detection of offensive language is essential for a number of applications related to social media safety. There is a sharp contrast in performance in this task between low and high-resource languages. In this pa…

XLM-R

Multilingual Encoder Knows more than You Realize: Shared Weights Pretraining for Extremely Low-Resource Languages

2025-02-15 · Zeli Su, Ziyin Zhang, Guixian Xu, Jianing Liu 외

While multilingual language models like XLM-R have advanced multilingualism in NLP, they still perform poorly in extremely low-resource languages. This situation is exacerbated by the fact that modern LLMs such as LLaMA …

DecoderText GenerationXLM-R

AmaSQuAD: A Benchmark for Amharic Extractive Question Answering

2025-02-04 · Nebiyou Daniel Hailemariam, Blessed Guda, Tsegazeab Tefferi

This research presents a novel framework for translating extractive question-answering datasets into low-resource languages, as demonstrated by the creation of the AmaSQuAD dataset, a translation of SQuAD 2.0 into Amhari…

Extractive Question-AnsweringQuestion AnsweringXLM-R

Evaluating the Effectiveness of XAI Techniques for Encoder-Based Language Models

2025-01-26 · Melkamu Abay Mersha, Mesay Gemeda Yigezu, Jugal Kalita

The black-box nature of large language models (LLMs) necessitates the development of eXplainable AI (XAI) techniques for transparency and trustworthiness. However, evaluating these techniques remains a challenge. This st…

XLM-R

FuocChuVIP123 at CoMeDi Shared Task: Disagreement Ranking with XLM-Roberta Sentence Embeddings and Deep Neural Regression

2025-01-21 · Phuoc Duong Huy Chu

This paper presents results of our system for CoMeDi Shared Task, focusing on Subtask 2: Disagreement Ranking. Our system leverages sentence embeddings generated by the paraphrase-xlm-r-multilingual-v1 model, combined wi…

SentenceSentence EmbeddingsXLM-R

Comparative Approaches to Sentiment Analysis Using Datasets in Major European and Arabic Languages

2025-01-21 · Mikhail Krasitskii, Olga Kolesnikova, Liliana Chanona Hernandez, Grigori Sidorov 외

This study explores transformer-based models such as BERT, mBERT, and XLM-R for multi-lingual sentiment analysis across diverse linguistic structures. Key contributions include the identification of XLM-R superior adapta…

Sentiment AnalysisSentiment ClassificationXLM-R

Multi-stage Training of Bilingual Islamic LLM for Neural Passage Retrieval

2025-01-17 · Vera Pavlova

This study examines the use of Natural Language Processing (NLP) technology within the Islamic domain, focusing on developing an Islamic neural retrieval model. By leveraging the robust XLM-R model, the research employs …

Data AugmentationDomain AdaptationLanguage ModelingLanguage Modelling+4

BabyLMs for isiXhosa: Data-Efficient Language Modelling in a Low-Resource Context

2025-01-07 · Alexis Matzopoulos, Charl Hendriks, Hishaam Mahomed, Francois Meyer

The BabyLM challenge called on participants to develop sample-efficient language models. Submissions were pretrained on a fixed English corpus, limited to the amount of words children are exposed to in development (<100m…

Language ModellingNERPOSPOS Tagging+1

USTCCTSU at SemEval-2024 Task 1: Reducing Anisotropy for Cross-lingual Semantic Textual Relatedness Task

2024-11-28 · Jianjian Li, Shengwei Liang, Yong Liao, Hongping Deng 외

Cross-lingual semantic textual relatedness task is an important research task that addresses challenges in cross-lingual communication and text understanding. It helps establish semantic connections between different lan…

Information RetrievalMachine TranslationRetrievalSentence+1

Retrofitting Large Language Models with Dynamic Tokenization

2024-11-27 · Darius Feher, Ivan Vulić, Benjamin Minixhofer

Current language models (LMs) use a fixed, static subword tokenizer. This default choice typically results in degraded efficiency and language capabilities, especially in languages other than English. To address this iss…

DecoderFairnessXLM-R

Transformer-Based Contextualized Language Models Joint with Neural Networks for Natural Language Inference in Vietnamese

2024-11-20 · Dat Van-Thanh Nguyen, Tin Van Huynh, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen

Natural Language Inference (NLI) is a task within Natural Language Processing (NLP) that holds value for various AI applications. However, there have been limited studies on Natural Language Inference in Vietnamese that …

Natural Language InferenceXLM-R

From N-grams to Pre-trained Multilingual Models For Language Identification

2024-10-11 · Thapelo Sindane, Vukosi Marivate

In this paper, we investigate the use of N-gram models and Large Pre-trained Multilingual models for Language Identification (LID) across 11 South African languages. For N-gram models, this study shows that effective dat…

Language IdentificationXLM-R

LangSAMP: Language-Script Aware Multilingual Pretraining

2024-09-26 · Yihong Liu, Haotian Ye, Chunlan Ma, Mingyang Wang 외

Recent multilingual pretrained language models (mPLMs) often avoid using language embeddings -- learnable vectors assigned to different languages. These embeddings are discarded for two main reasons: (1) mPLMs are expect…

Continual PretrainingLanguage ModelingLanguage ModellingRepresentation Learning+1

GrEmLIn: A Repository of Green Baseline Embeddings for 87 Low-Resource Languages Injected with Multilingual Graph Knowledge

2024-09-26 · Daniil Gurgurov, Rishu Kumar, Simon Ostermann

Contextualized embeddings based on large language models (LLMs) are available for various languages, but their coverage is often limited for lower resourced languages. Using LLMs for such languages is often difficult due…

Natural Language InferenceSentiment AnalysisTopic ClassificationWord Embeddings+1

mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval

2024-07-29 · Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie 외

We present systematic efforts in building long-context multilingual text representation model (TRM) and reranker from scratch for text retrieval. We first introduce a text encoder (base size) enhanced with RoPE and unpad…

Contrastive LearningRerankingRetrievalText Retrieval+1

The Model Arena for Cross-lingual Sentiment Analysis: A Comparative Study in the Era of Large Language Models

2024-06-27 · Xiliang Zhu, Shayna Gardiner, Tere Roldán, David Rossouw

Sentiment analysis serves as a pivotal component in Natural Language Processing (NLP). Advancements in multilingual pre-trained models such as XLM-R and mT5 have contributed to the increasing interest in cross-lingual se…

Cross-Lingual TransferSentiment AnalysisXLM-R

Medical Spoken Named Entity Recognition

2024-06-19 · Khai Le-Duc, David Thulke, Hung-Phong Tran, Long Vo-Dang 외

Spoken Named Entity Recognition (NER) aims to extract named entities from speech and categorise them into types like person, location, organization, etc. In this work, we present VietMed-NER - the first spoken NER datase…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

Multilingual Large Language Models and Curse of Multilinguality

2024-06-15 · Daniil Gurgurov, Tanja Bäumel, Tatiana Anikina

Multilingual Large Language Models (LLMs) have gained large popularity among Natural Language Processing (NLP) researchers and practitioners. These models, trained on huge datasets, show proficiency across various langua…

DecoderXLM-R

Exploring Alignment in Shared Cross-lingual Spaces

2024-05-23 · Basel Mousi, Nadir Durrani, Fahim Dalvi, Majd Hawasly 외

Despite their remarkable ability to capture linguistic nuances across diverse languages, questions persist regarding the degree of alignment between languages in multilingual embeddings. Drawing inspiration from research…

Machine Translationnamed-entity-recognitionNamed Entity RecognitionSentiment Analysis+1
1–20 / 221 다음 →