paper-with-me

홈 › Papers

FreCDo: A Large Corpus for French Cross-Domain Dialect Identification

2022-12-15 · Mihaela Gaman, Adrian-Gabriel Chifu, William Domingues, Radu Tudor Ionescu

We present a novel corpus for French dialect identification comprising 413,522 French text samples collected from public news websites in Belgium, Canada, France and Switzerland. To ensure an accurate estimation of the dialect identification performance of models, we designed the corpus to eliminate potential biases related to topic, writing style, and publication source. More precisely, the training, validation and test splits are collected from different news websites, while searching for different keywords (topics). This leads to a French cross-domain (FreCDo) dialect identification task. We conduct experiments with four competitive baselines, a fine-tuned CamemBERT model, an XGBoost based on fine-tuned CamemBERT features, a Support Vector Machines (SVM) classifier based on fine-tuned CamemBERT features, and an SVM based on word n-grams. Aside from presenting quantitative results, we also make an analysis of the most discriminative features learned by CamemBERT. Our corpus is available at https://github.com/MihaelaGaman/FreCDo.

📄 PDF Abstract BibTeX arXiv:2212.07707

Code (1)

mihaelagaman/frecdo 공식 구현 pytorch

Tasks

Dialect Identification

Methods 이 논문이 사용한 방법론

Test 설명 없음
SVM A Support Vector Machine, or SVM, is a non-parametric supervised learning model. For non-linear classification and regression, they utilise the kernel trick to map inputs…

Similar Papers 제목 키워드 기반

HistoriQA-ThirdRepublic: Multi-Hop Question Answering Corpus for Historical Research, Parliamentary Debates from the French Third Republic (1870-1940)

2026-06-30 · Aurélien Pellet, Julien Perez, Marie Puren arxiv

We present HistoriQA-ThirdRepublic: a French-language dataset of multi-hop historical questions derived from parliamentary debates and newspapers of the French Third Republic. Designed in collaboration with a historian, …

Multi-hop Question Answering

SEDAR: a Large Scale French-English Financial Domain Parallel Corpus

2020-05-01 · LREC 2020 5 · Abbas Ghaddar, Phillippe Langlais

This paper describes the acquisition, preprocessing and characteristics of SEDAR, a large scale English-French parallel corpus for the financial domain. Our extensive experiments on machine translation show that SEDAR is…

Domain AdaptationMachine TranslationSentenceTranslation

French Resources for Extraction and Normalization of Temporal Expressions with HeidelTime

2014-05-01 · LREC 2014 5 · V{\'e}ronique Moriceau, Xavier Tannier

In this paper, we describe the development of French resources for the extraction and normalization of temporal expressions with HeidelTime, a open-source multilingual, cross-domain temporal tagger. HeidelTime extracts t…

ArticlesInformation Retrieval

Is Biomedical Specialization Still Worth It? Insights from Domain-Adaptive Language Modelling with a New French Health Corpus

2026-04-08 · Aidan Mannion, Cécile Macaire, Armand Violle, Stéphane Ohayon 외 arxiv

Large language models (LLMs) have demonstrated remarkable capabilities across diverse domains, yet their adaptation to specialized fields remains challenging, particularly for non-English languages. This study investigat…

Language ModellingDomain Adaptation

SENCORPUS: A French-Wolof Parallel Corpus

2020-05-01 · LREC 2020 5 · Elhadji Mamadou Nguer, Alla Lo, Cheikh M. Bamba Dione, Sileye O. Ba 외

In this paper, we report efforts towards the acquisition and construction of a bilingual parallel corpus between French and Wolof, a Niger-Congo language belonging to the Northern branch of the Atlantic group. The corpus…

Machine TranslationTranslation