paper-with-me

Papers

Data-Efficient French Language Modeling with CamemBERTa

2023-06-02 · Wissam Antoun, Benoît Sagot, Djamé Seddah

Recent advances in NLP have significantly improved the performance of language models on a variety of tasks. While these advances are largely driven by the availability of large amounts of data and computational power, they also benefit from the development of better training methods and architectures. In this paper, we introduce CamemBERTa, a French DeBERTa model that builds upon the DeBERTaV3 architecture and training objective. We evaluate our model's performance on a variety of French downstream tasks and datasets, including question answering, part-of-speech tagging, dependency parsing, named entity recognition, and the FLUE benchmark, and compare against CamemBERT, the state-of-the-art monolingual model for French. Our results show that, given the same amount of training tokens, our model outperforms BERT-based models trained with MLM on most tasks. Furthermore, our new model reaches similar or superior performance on downstream tasks compared to CamemBERT, despite being trained on only 30% of its total number of input tokens. In addition to our experimental results, we also publicly release the weights and code implementation of CamemBERTa, making it the first publicly available DeBERTaV3 model outside of the original paper and the first openly available implementation of a DeBERTaV3 training objective. https://gitlab.inria.fr/almanach/CamemBERTa

📄 PDF Abstract BibTeX arXiv:2306.01497

Code (0)

등록된 구현이 없습니다.

Tasks

Dependency ParsingFLUELanguage ModelingLanguage Modellingnamed-entity-recognitionNamed Entity RecognitionPart-Of-Speech TaggingQuestion Answering

Methods 이 논문이 사용한 방법론

How do I file a dispute with Expedia?*DisputeFastService How do I file a dispute with Expedia? To file a dispute with Expedia, call +1(888) (829) (0881) OR +1(805) (330) (4056), or use their Help Center to submit your case with…
DeBERTa DeBERTa is a Transformer-based neural language model that aims to improve the…

Similar Papers 제목 키워드 기반

CamemBERT 2.0: A Smarter French Language Model Aged to Perfection

2024-11-13 · Wissam Antoun, Francis Kulumba, Rian Touchent, Éric de la Clergerie 외

French language models, such as CamemBERT, have been widely adopted across industries for natural language processing (NLP) tasks, with models like CamemBERT seeing over 4 million downloads per month. However, these mode…

Language ModelingLanguage ModellingMasked Language Modeling

ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance

2025-04-11 · Wissam Antoun, Benoît Sagot, Djamé Seddah

Pretrained transformer-encoder models like DeBERTaV3 and ModernBERT introduce architectural advancements aimed at improving efficiency and performance. Although the authors of ModernBERT report improved performance over …

Modeling French Sign Language: a proposal for a semantically compositional system

2018-05-01 · LREC 2018 5 · Mohamed Nassime Hadjadj, Michael Filhol, Annelies Braffort

FQuAD: French Question Answering Dataset

2020-02-14 · Findings of the Association for Computational Linguistics 2020 · Martin d'Hoffschmidt, Wacim Belblidia, Tom Brendlé, Quentin Heinrich 외

Recent advances in the field of language modeling have improved state-of-the-art results on many Natural Language Processing tasks. Among them, Reading Comprehension has made significant progress over the past few years.…

ArticlesFQuADLanguage ModelingLanguage Modelling+3

Two New AZee Production Rules Refining Multiplicity in French Sign Language

2022-06-01 · SignLang (LREC) 2022 6 · Emmanuella Martinod, Claire Danet, Michael Filhol

This paper is a contribution to sign language (SL) modeling. We focus on the hitherto imprecise notion of “Multiplicity”, assumed to express plurality in French Sign Language (LSF), using AZee approach. AZee is a linguis…

Vocal Bursts Valence Prediction