paper-with-me

홈 › Papers

Comparing the Performance of Feature Representations for the Categorization of the Easy-to-Read Variety vs Standard Language

2019-09-01 · WS (NoDaLiDa) 2019 9 · Marina Santini, Benjamin Danielsson, Arne Jönsson

We explore the effectiveness of four feature representations – bag-of-words, word embeddings, principal components and autoencoders – for the binary categorization of the easy-to-read variety vs standard language. Standard language refers to the ordinary language variety used by a population as a whole or by a community, while the “easy-to-read” variety is a simpler (or a simplified) version of the standard language. We test the efficiency of these feature representations on three corpora, which differ in size, class balance, unit of analysis, language and topic. We rely on supervised and unsupervised machine learning algorithms. Results show that bag-of-words is a robust and straightforward feature representation for this task and performs well in many experimental settings. Its performance is equivalent or equal to the performance achieved with principal components and autoencorders, whose preprocessing is however more time-consuming. Word embeddings are less accurate than the other feature representations for this classification task.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Word Embeddings

Similar Papers 제목 키워드 기반

Image Categorization and Search via a GAT Autoencoder and Representative Models

2025-10-18 · Duygu Sap, Martin Lotz, Connor Mattinson arxiv

We propose a method for image categorization and retrieval that leverages graphs and a graph attention network (GAT)-based autoencoder. Our approach is representative-centric, that is, we execute the categorization and r…

Learning Deep Parsimonious Representations

2016-12-01 · NeurIPS 2016 12 · Renjie Liao, Alex Schwing, Richard Zemel, Raquel Urtasun

In this paper we aim at facilitating generalization for deep networks while supporting interpretability of the learned representations. Towards this goal, we propose a clustering based regularization that encourages pars…

ClusteringFew-Shot Image ClassificationGeneral ClassificationZero-Shot Learning

Multiple Granularity Descriptors for Fine-Grained Categorization

2015-12-01 · ICCV 2015 12 · Dequan Wang, Zhiqiang Shen, Jie Shao, Wei zhang 외

Fine-grained categorization, which aims to distinguish subordinate-level categories such as bird species or dog breeds, is an extremely challenging task. This is due to two main issues: how to localize discriminative reg…

Protein Conformational States: A First Principles Bayesian Method

2020-08-05 · David M. Rogers

Automated identification of protein conformational states from simulation of an ensemble of structures is a hard problem because it requires teaching a computer to recognize shapes. We adapt the naive Bayes classifier fr…

Vision CNNs trained to estimate spatial latents learned similar ventral-stream-aligned representations

2024-12-12 · Yudi Xie, Weichen Huang, Esther Alter, Jeremy Schwartz 외

Studies of the functional role of the primate ventral visual stream have traditionally focused on object categorization, often ignoring -- despite much prior evidence -- its role in estimating "spatial" latents such as o…

Object Categorization