KeyVec: Key-semantics Preserving Document Representations
Previous studies have demonstrated the empirical success of word embeddings in various applications. In this paper, we investigate the problem of learning distributed representations for text documents which many machine learning algorithms take as input for a number of NLP tasks. We propose a neural network model, KeyVec, which learns document representations with the goal of preserving key semantics of the input text. It enables the learned low-dimensional vectors to retain the topics and important information from the documents that will flow to downstream tasks. Our empirical evaluations show the superior quality of KeyVec representations in two different document understanding tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine Learningdocument understandingWord EmbeddingsSimilar Papers 제목 키워드 기반
Structure and Semantics Preserving Document Representations
Retrieving relevant documents from a corpus is typically based on the semantic similarity between the document content and query text. The inclusion of structural relationship between documents can benefit the retrieval …
Metric LearningRetrievalSemantic SimilaritySemantic Textual SimilarityAMR Beyond the Sentence: the Multi-sentence AMR corpus
There are few corpora that endeavor to represent the semantic content of entire documents. We present a corpus that accomplishes one way of capturing document level semantics, by annotating coreference and similar phenom…
Question AnsweringSentenceCascaded Semantic and Positional Self-Attention Network for Document Classification
Transformers have shown great success in learning representations for language modelling. However, an open challenge still remains on how to systematically aggregate semantic information (word embedding) with positional …
ClassificationDocument ClassificationGeneral ClassificationLanguage ModellingTop2Vec: Distributed Representations of Topics
Topic modeling is used for discovering latent semantic structure, usually referred to as topics, in a large collection of documents. The most widely used methods are Latent Dirichlet Allocation and Probabilistic Latent S…
LemmatizationSemantic SimilaritySemantic Textual SimilarityTopic ModelsSemantics-aware Motion Retargeting with Vision-Language Models
Capturing and preserving motion semantics is essential to motion retargeting between animation characters. However, most of the previous works neglect the semantic information or rely on human-designed joint-level repres…
Language ModelingLanguage Modellingmotion retargeting