paper-with-me

Papers

Quantifying French Document Complexity

2022-08-27 · Vincent Primpied, David Beauchemin, Richard Khoury

Measuring a document's complexity level is an open challenge, particularly when one is working on a diverse corpus of documents rather than comparing several documents on a similar topic or working on a language other than English. In this paper, we define a methodology to measure the complexity of French documents, using a new general and diversified corpus of texts, the "French Canadian complexity level corpus", and a wide range of metrics. We compare different learning algorithms to this task and contrast their performances and their observations on which characteristics of the texts are more significant to their complexity. Our results show that our methodology gives a general-purpose measurement of text complexity in French.

📄 PDF Abstract BibTeX arXiv:2208.12924

Code (2)

graal-research/fcclc 공식 구현
graal-research/textcomplexitycomputer 공식 구현

Similar Papers 제목 키워드 기반

Automatic identification of document sections for designing a French clinical corpus (Identification automatique de zones dans des documents pour la constitution d'un corpus m\'edical en fran\ccais) [in French]

2014-07-01 · JEPTALNRECITAL 2014 7 · Louise Del{\'e}ger, Aur{\'e}lie N{\'e}v{\'e}ol

Dating Ancient texts: an Approach for Noisy French Documents

2020-05-01 · LREC 2020 5 · Ana{\"e}lle Baledent, Nicolas Hiebel, Ga{\"e}l Lejeune

Automatic dating of ancient documents is a very important area of research for digital humanities applications. Many documents available via digital libraries do not have any dating or dating that is uncertain. Document …

Document DatingPOS

Automated Drug-Related Information Extraction from French Clinical Documents: ReLyfe Approach

2021-11-29 · Azzam Alwan, Maayane Attias, Larry Rubin, Adnan El Bakri

Structuring medical data in France remains a challenge mainly because of the lack of medical data due to privacy concerns and the lack of methods and approaches on processing the French language. One of these challenges …

Management

French Resources for Extraction and Normalization of Temporal Expressions with HeidelTime

2014-05-01 · LREC 2014 5 · V{\'e}ronique Moriceau, Xavier Tannier

In this paper, we describe the development of French resources for the extraction and normalization of temporal expressions with HeidelTime, a open-source multilingual, cross-domain temporal tagger. HeidelTime extracts t…

ArticlesInformation Retrieval

Detection of Text Reuse in French Medical Corpora

2016-12-01 · WS 2016 12 · Eva D{'}hondt, Cyril Grouin, Aur{\'e}lie N{\'e}v{\'e}ol, Efstathios Stamatatos 외

Electronic Health Records (EHRs) are increasingly available in modern health care institutions either through the direct creation of electronic documents in hospitals{'} health information systems, or through the digitiz…

De-identificationOptical Character Recognition (OCR)