Quantifying French Document Complexity
Measuring a document's complexity level is an open challenge, particularly when one is working on a diverse corpus of documents rather than comparing several documents on a similar topic or working on a language other than English. In this paper, we define a methodology to measure the complexity of French documents, using a new general and diversified corpus of texts, the "French Canadian complexity level corpus", and a wide range of metrics. We compare different learning algorithms to this task and contrast their performances and their observations on which characteristics of the texts are more significant to their complexity. Our results show that our methodology gives a general-purpose measurement of text complexity in French.
Code (2)
Similar Papers 제목 키워드 기반
Automatic identification of document sections for designing a French clinical corpus (Identification automatique de zones dans des documents pour la constitution d'un corpus m\'edical en fran\ccais) [in French]
Dating Ancient texts: an Approach for Noisy French Documents
Automatic dating of ancient documents is a very important area of research for digital humanities applications. Many documents available via digital libraries do not have any dating or dating that is uncertain. Document …
Document DatingPOSAutomated Drug-Related Information Extraction from French Clinical Documents: ReLyfe Approach
Structuring medical data in France remains a challenge mainly because of the lack of medical data due to privacy concerns and the lack of methods and approaches on processing the French language. One of these challenges …
ManagementFrench Resources for Extraction and Normalization of Temporal Expressions with HeidelTime
In this paper, we describe the development of French resources for the extraction and normalization of temporal expressions with HeidelTime, a open-source multilingual, cross-domain temporal tagger. HeidelTime extracts t…
ArticlesInformation RetrievalDetection of Text Reuse in French Medical Corpora
Electronic Health Records (EHRs) are increasingly available in modern health care institutions either through the direct creation of electronic documents in hospitals{'} health information systems, or through the digitiz…
De-identificationOptical Character Recognition (OCR)