paper-with-me

홈 › Papers

UMUTextStats: A linguistic feature extraction tool for Spanish

2022-06-01 · LREC 2022 6 · José Antonio García-Díaz, Pedro José Vivancos-Vicente, Ángela Almela, Rafael Valencia-García

Feature Engineering consists in the application of domain knowledge to select and transform relevant features to build efficient machine learning models. In the Natural Language Processing field, the state of the art concerning automatic document classification tasks relies on word and sentence embeddings built upon deep learning models based on transformers that have outperformed the competition in several tasks. However, the models built from these embeddings are usually difficult to interpret. On the contrary, linguistic features are easy to understand, they result in simpler models, and they usually achieve encouraging results. Moreover, both linguistic features and embeddings can be combined with different strategies which result in more reliable machine-learning models. The de facto tool for extracting linguistic features in Spanish is LIWC. However, this software does not consider specific linguistic phenomena of Spanish such as grammatical gender and lacks certain verb tenses. In order to solve these drawbacks, we have developed UMUTextStats, a linguistic extraction tool designed from scratch for Spanish. Furthermore, this tool has been validated to conduct different experiments in areas such as infodemiology, hate-speech detection, author profiling, authorship verification, humour or irony detection, among others. The results indicate that the combination of linguistic features and embeddings based on transformers are beneficial in automatic document classification.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Author ProfilingAuthorship VerificationDocument ClassificationFeature EngineeringHate Speech DetectionSentenceSentence Embeddings

Similar Papers 제목 키워드 기반

PUCP-Metrix: An Open-source and Comprehensive Toolkit for Linguistic Analysis of Spanish Texts

2025-11-21 · Javier Alonso Villegas Luis, Marco Antonio Sobrevilla Cabezudo arxiv

Linguistic features remain essential for interpretability and tasks that involve style, structure, and readability, but existing Spanish tools offer limited coverage. We present PUCP-Metrix, an open-source and comprehens…

Text Detection

Linguistically Informed Relation Extraction and Neural Architectures for Nested Named Entity Recognition in BioNLP-OST 2019

2019-10-08 · WS 2019 11 · Usama Yaseen, Pankaj Gupta, Hinrich Schütze

Named Entity Recognition (NER) and Relation Extraction (RE) are essential tools in distilling knowledge from biomedical literature. This paper presents our findings from participating in BioNLP Shared Tasks 2019. We addr…

Binary Relation Extractionnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+4

CNER: A tool Classifier of Named-Entity Relationships

2024-05-17 · Jefferson A. Peña Torres, Raúl E. Gutiérrez De Piñerez

We introduce CNER, an ensemble of capable tools for extraction of semantic relationships between named entities in Spanish language. Built upon a container-based architecture, CNER integrates different Named entity recog…

named-entity-recognitionNamed Entity RecognitionRelation Extraction

Automatic Detection of Offensive Language in Social Media: Defining Linguistic Criteria to build a Mexican Spanish Dataset

2020-05-01 · LREC 2020 5 · Mar{\'\i}a Jos{\'e} D{\'\i}az-Torres, Paulina Alej Mor{\'a}n-M{\'e}ndez, ra, Luis Villasenor-Pineda 외

Phenomena such as bullying, homophobia, sexism and racism have transcended to social networks, motivating the development of tools for their automatic detection. The challenge becomes greater for languages rich in popula…

Abusive Language

MUST&P-SRL: Multi-lingual and Unified Syllabification in Text and Phonetic Domains for Speech Representation Learning

2023-10-17 · Noé Tits

In this paper, we present a methodology for linguistic feature extraction, focusing particularly on automatically syllabifying words in multiple languages, with a design to be compatible with a forced-alignment tool, the…

DisentanglementRepresentation LearningSpeech Representation Learning