paper-with-me

Papers

Enhancing Portuguese Variety Identification with Cross-Domain Approaches

2025-02-20 · Hugo Sousa, Rúben Almeida, Purificação Silvano, Inês Cantante, Ricardo Campos, Alípio Jorge

Recent advances in natural language processing have raised expectations for generative models to produce coherent text across diverse language varieties. In the particular case of the Portuguese language, the predominance of Brazilian Portuguese corpora online introduces linguistic biases in these models, limiting their applicability outside of Brazil. To address this gap and promote the creation of European Portuguese resources, we developed a cross-domain language variety identifier (LVI) to discriminate between European and Brazilian Portuguese. Motivated by the findings of our literature review, we compiled the PtBrVarId corpus, a cross-domain LVI dataset, and study the effectiveness of transformer-based LVI classifiers for cross-domain scenarios. Although this research focuses on two Portuguese varieties, our contribution can be extended to other varieties and languages. We open source the code, corpus, and models to foster further research in this task.

📄 PDF Abstract BibTeX arXiv:2502.14394

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Language Variety Identification with True Labels

2023-03-02 · Marcos Zampieri, Kai North, Tommi Jauhiainen, Mariano Felice 외

Language identification is an important first step in many IR and NLP applications. Most publicly available language identification datasets, however, are compiled under the assumption that the gold label of each instanc…

Language Identification

Including Dialects and Language Varieties in Author Profiling

2017-07-03 · Alina Maria Ciobanu, Marcos Zampieri, Shervin Malmasi, Liviu P. Dinu

This paper presents a computational approach to author profiling taking gender and language variety into account. We apply an ensemble system with the output of multiple linear SVM classifiers trained on character and wo…

Author Profiling

ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs

2026-03-27 · Inês Vieira, Inês Calvo, Iago Paulo, James Furtado 외 arxiv

As Large Language Models (LLMs) expand across multilingual domains, evaluating their performance in under-represented languages becomes increasingly important. European Portuguese (pt-PT) is particularly affected, as exi…

Using Social Networks to Improve Language Variety Identification with Neural Networks

2017-11-01 · IJCNLP 2017 11 · Yasuhide Miura, Tomoki Taniguchi, Motoki Taniguchi, Shotaro Misawa 외

We propose a hierarchical neural network model for language variety identification that integrates information from a social network. Recently, language variety identification has enjoyed heightened popularity as an adva…

Language Identification

P3B3: A Multi-Turn Conversational Benchmark for Measuring European and Brazilian Portuguese Variety Bias in LLMs

2026-06-15 · Rafael Ferreira, Inês Vieira, Inês Calvo, James Furtado 외 arxiv

As Large Language Models (LLMs) become embedded in everyday communication, capturing regional linguistic variation is essential for reliable and equitable language use. In Portuguese, European (pt-PT) and Brazilian (pt-B…