paper-with-me

Papers

Evaluation Of Word Embeddings From Large-Scale French Web Content

2021-05-05 · Hadi Abdine, Christos Xypolopoulos, Moussa Kamal Eddine, Michalis Vazirgiannis

Distributed word representations are popularly used in many tasks in natural language processing. Adding that pretrained word vectors on huge text corpus achieved high performance in many different NLP tasks. This paper introduces multiple high-quality word vectors for the French language where two of them are trained on massive crawled French data during this study and the others are trained on an already existing French corpus. We also evaluate the quality of our proposed word vectors and the existing French word vectors on the French word analogy task. In addition, we do the evaluation on multiple real NLP tasks that shows the important performance enhancement of the pre-trained word vectors compared to the existing and random ones. Finally, we created a demo web application to test and visualize the obtained word embeddings. The produced French word embeddings are available to the public, along with the finetuning code on the NLU tasks and the demo code.

📄 PDF Abstract BibTeX arXiv:2105.01990

Code (1)

hadi-abdine/WordEmbeddingsEvalFLUE 공식 구현 pytorch

Tasks

Word Embeddings

Similar Papers 제목 키워드 기반

BERTrade: Using Contextual Embeddings to Parse Old French

2022-06-01 · LREC 2022 6 · Loïc Grobol, Mathilde Regnault, Pedro Ortiz Suarez, Benoît Sagot 외

The successes of contextual word embeddings learned by training large-scale language models, while remarkable, have mostly occurred for languages where significant amounts of raw texts are available and where annotated d…

Dependency ParsingPOSPositionPOS Tagging+1

Neural Networks approaches focused on French Spoken Language Understanding: application to the MEDIA Evaluation Task

2020-12-01 · COLING 2020 8 · Sahar Ghannay, Christophe Servan, Sophie Rosset

In this paper, we present a study on a French Spoken Language Understanding (SLU) task: the MEDIA task. Many works and studies have been proposed for many tasks, but most of them are focused on English language and tasks…

Spoken Language UnderstandingWord Embeddings

Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM

2026-09-15 · Adam Zachary Wasserman, David Beauchemin arxiv

We submit MéTRON-FR, a 125M GPT-2 pretrained on 92.47M words of French, to the BabyLM 2026 Strict track. It scores 85.97 +/- 0.17% on QFrBLiMP (a native Quebec-French benchmark of grammatical minimal pairs) and 62.80% on…

Adapted Sentiment Similarity Seed Words For French Tweets' Polarity Classification

2018-05-01 · JEPTALNRECITAL 2018 5 · Amal Htait

We present, in this paper, our contribution in DEFT 2018 task 2 : {``}Global polarity{''}, determining the overall polarity (Positive, Negative, Neutral or MixPosNeg) of tweets regarding public transport, in French langu…

General ClassificationTask 2Word Embeddings

JeuxDeLiens: Word Embeddings and Path-Based Similarity for Entity Linking using the French JeuxDeMots Lexical Semantic Network

2018-05-01 · JEPTALNRECITAL 2018 5 · Julien Plu, Kevin Cousot, Mathieu Lafourcade, Rapha{\"e}l Troncy 외

Entity linking systems typically rely on encyclopedic knowledge bases such as DBpedia or Freebase. In this paper, we use, instead, a French lexical-semantic network named JeuxDeMots to jointly type and link entities. Our…

Entity LinkingWord Embeddings