paper-with-me

홈 › Papers

Development of email classifier in Brazilian Portuguese using feature selection for automatic response

2019-07-08 · Rogerio Bonatti, Arthur Gola de Paula

Automatic email categorization is an important application of text classification. We study the automatic reply of email business messages in Brazilian Portuguese. We present a novel corpus containing messages from a real application, and baseline categorization experiments using Naive Bayes and support Vector Machines. We then discuss the effect of lemmatization and the role of part-of-speech tagging filtering on precision and recall. Support Vector Machines classification coupled with nonlemmatized selection of verbs, nouns and adjectives was the best approach, with 87.3% maximum accuracy. Straightforward lemmatization in Portuguese led to the lowest classification results in the group, with 85.3% and 81.7% precision in SVM and Naive Bayes respectively. Thus, while lemmatization reduced precision and recall, part-of-speech filtering improved overall results.

📄 PDF Abstract BibTeX arXiv:1907.04905

Code (0)

등록된 구현이 없습니다.

Tasks

Classificationfeature selectionGeneral ClassificationLemmatizationPart-Of-Speech Taggingtext-classificationText Classification

Methods 이 논문이 사용한 방법론

SVM A Support Vector Machine, or SVM, is a non-parametric supervised learning model. For non-linear classification and regression, they utilise the kernel trick to map inputs…

Similar Papers 제목 키워드 기반

Image captioning for Brazilian Portuguese using GRIT model

2024-02-07 · Rafael Silva de Alencar, William Alberto Cruz Castañeda, Marcellus Amadeus

This work presents the early development of a model of image captioning for the Brazilian Portuguese language. We used the GRIT (Grid - and Region-based Image captioning Transformer) model to accomplish this work. GRIT i…

Image Captioningmodel

From Brazilian Portuguese to European Portuguese

2024-08-14 · João Sanches, Rui Ribeiro, Luísa Coheur

Brazilian Portuguese and European Portuguese are two varieties of the same language and, despite their close similarities, they exhibit several differences. However, there is a significant disproportion in the availabili…

SentenceTranslation

PublicHearingBR: A Brazilian Portuguese Dataset of Public Hearing Transcripts for Summarization of Long Documents

2024-10-10 · Leandro Carísio Fernandes, Guilherme Zeferino Rodrigues Dobins, Roberto Lotufo, Jayr Alencar Pereira

This paper introduces PublicHearingBR, a Brazilian Portuguese dataset designed for summarizing long documents. The dataset consists of transcripts of public hearings held by the Brazilian Chamber of Deputies, paired with…

ArticlesDocument SummarizationHallucinationNatural Language Inference

Descri\cc\~ao e modelagem de constru\cc\~oes interrogativas QU- em Portugu\^es Brasileiro para o desenvolvimento de um chatbot (Description and modeling of interrogative constructs QU- in Brazilian Portuguese for the development of a chatbot)[In Portuguese]

2017-10-01 · WS 2017 10 · Nat{\'a}lia Duarte Mar{\c{c}}{\~a}o, Tiago Timponi Torrent, Ely Edison da Silva Matos
Chatbot

Enhancing Portuguese Variety Identification with Cross-Domain Approaches

2025-02-20 · Hugo Sousa, Rúben Almeida, Purificação Silvano, Inês Cantante 외

Recent advances in natural language processing have raised expectations for generative models to produce coherent text across diverse language varieties. In the particular case of the Portuguese language, the predominanc…