paper-with-me

홈 › Papers

NILC-Metrix: assessing the complexity of written and spoken language in Brazilian Portuguese

2021-12-17 · Sidney Evaldo Leal, Magali Sanches Duran, Carolina Evaristo Scarton, Nathan Siegle Hartmann, Sandra Maria Aluísio

This paper presents and makes publicly available the NILC-Metrix, a computational system comprising 200 metrics proposed in studies on discourse, psycholinguistics, cognitive and computational linguistics, to assess textual complexity in Brazilian Portuguese (BP). These metrics are relevant for descriptive analysis and the creation of computational models and can be used to extract information from various linguistic levels of written and spoken language. The metrics in NILC-Metrix were developed during the last 13 years, starting in 2008 with Coh-Metrix-Port, a tool developed within the scope of the PorSimples project. Coh-Metrix-Port adapted some metrics to BP from the Coh-Metrix tool that computes metrics related to cohesion and coherence of texts in English. After the end of PorSimples in 2010, new metrics were added to the initial 48 metrics of Coh-Metrix-Port. Given the large number of metrics, we present them following an organisation similar to the metrics of Coh-Metrix v3.0 to facilitate comparisons made with metrics in Portuguese and English. In this paper, we illustrate the potential of NILC-Metrix by presenting three applications: (i) a descriptive analysis of the differences between children's film subtitles and texts written for Elementary School I and II (Final Years); (ii) a new predictor of textual complexity for the corpus of original and simplified texts of the PorSimples project; (iii) a complexity prediction model for school grades, using transcripts of children's story narratives told by teenagers. For each application, we evaluate which groups of metrics are more discriminative, showing their contribution for each task.

📄 PDF Abstract BibTeX arXiv:2201.03445

Code (0)

등록된 구현이 없습니다.

Tasks

Descriptive

Similar Papers 제목 키워드 기반

Coh-Metrix-Esp: A Complexity Analysis Tool for Documents Written in Spanish

2016-05-01 · LREC 2016 5 · Andre Quispesaravia, Walter Perez, Marco Sobrevilla Cabezudo, Fern Alva-Manchego 외

Text Complexity Analysis is an useful task in Education. For example, it can help teachers select appropriate texts for their students according to their educational level. This task requires the analysis of several text…

Cross-corpora experiments of automatic proficiency assessment and error detection for spoken English

2022-07-01 · NAACL (BEA) 2022 7 · Stefano Bannò, Marco Matassoni

The growing demand for learning English as a second language has led to an increasing interest in automatic approaches for assessing spoken language proficiency. One of the most significant challenges in this field is th…

PUCP-Metrix: An Open-source and Comprehensive Toolkit for Linguistic Analysis of Spanish Texts

2025-11-21 · Javier Alonso Villegas Luis, Marco Antonio Sobrevilla Cabezudo arxiv

Linguistic features remain essential for interpretability and tasks that involve style, structure, and readability, but existing Spanish tools offer limited coverage. We present PUCP-Metrix, an open-source and comprehens…

Text Detection

Assessing Human Translations from French to Bambara for Machine Learning: a Pilot Study

2020-03-31 · Michael Leventhal, Allahsera Tapo, Sarah Luger, Marcos Zampieri 외

We present novel methods for assessing the quality of human-translated aligned texts for learning machine translation models of under-resourced languages. Malian university students translated French texts, producing eit…

BIG-bench Machine LearningMachine TranslationTranslation

Open corpora and toolkit for assessing text readability in French

2022-06-01 · READI (LREC) 2022 6 · Nicolas Hernandez, Nabil Oulbaz, Tristan Faine

Measuring the linguistic complexity or assessing the readability of spoken or written productions has been the concern of several researchers in pedagogy and (foreign) language teaching for decades. Researchers study for…

Text Simplification