Diseño de un espacio semántico sobre la base de la Wikipedia. Una propuesta de análisis de la semántica latente para el idioma español
Latent Semantic Analysis (LSA) was initially conceived by the cognitive psychology at the 90s decade. Since its emergence, the LSA has been used to model cognitive processes, pointing out academic texts, compare literature works and analyse political speeches, among other applications. Taking as starting point multivariate method for dimensionality reduction, this paper propose a semantic space for Spanish language. Out results include a document text matrix with dimensions 1.3 x10^6 and 5.9x10^6, which later is decomposed into singular values. Those singular values are used to semantically words or text.
Code (0)
등록된 구현이 없습니다.
Tasks
Dimensionality ReductionSimilar Papers 제목 키워드 기반
Nota Sobre Algumas Interpretacoes da Teoria de Tributacao Otima
This note discusses some aspects of interpretations of the theory of optimal taxation presented in recent works on the Brazilian tax system.
Investiga\cc\~ao Preliminar Sobre a Pros\'odia Sem\^antica de Verbos de Elocu\cc\~ao: o Caso do Verbo ``Confessar'' (A Preliminar Investigation About the Semantic Prosody of Elocution Verbs: the Case for the Verb ``Confess'')[In Portuguese]
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modelling techniques to audio data. However, tr…
DecoderLanguage ModellingHas Anti-corruption Efforts lowered Enterprises Innovation Efficiency? -An Empirical Analysis from China
This study adopts the fixed effects panel model and provincial panel data on anticorruption and the innovation efficiency of high-level technology and new technology enterprises in China from 2005 to 2014, to estimate th…
Auditing and Robustifying COVID-19 Misinformation Datasets via Anticontent Sampling
This paper makes two key contributions. First, it argues that highly specialized rare content classifiers trained on small data typically have limited exposure to the richness and topical diversity of the negative class …
Active LearningDiversityMisinformation