paper-with-me

Papers

SimDoc: Topic Sequence Alignment based Document Similarity Framework

2016-11-15 · Gaurav Maheshwari, Priyansh Trivedi, Harshita Sahijwani, Kunal Jha, Sourish Dasgupta, Jens Lehmann

Document similarity is the problem of estimating the degree to which a given pair of documents has similar semantic content. An accurate document similarity measure can improve several enterprise relevant tasks such as document clustering, text mining, and question-answering. In this paper, we show that a document's thematic flow, which is often disregarded by bag-of-word techniques, is pivotal in estimating their similarity. To this end, we propose a novel semantic document similarity framework, called SimDoc. We model documents as topic-sequences, where topics represent latent generative clusters of related words. Then, we use a sequence alignment algorithm to estimate their semantic similarity. We further conceptualize a novel mechanism to compute topic-topic similarity to fine tune our system. In our experiments, we show that SimDoc outperforms many contemporary bag-of-words techniques in accurately computing document similarity, and on practical applications such as document clustering.

📄 PDF Abstract BibTeX arXiv:1611.04822

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringQuestion AnsweringSemantic SimilaritySemantic Textual Similarity

Similar Papers 제목 키워드 기반

Building and Aligning Comparable Corpora

2025-08-04 · Motaz Saad, David Langlois, Kamel Smaili arxiv

Comparable corpus is a set of topic aligned documents in multiple languages, which are not necessarily translations of each other. These documents are useful for multilingual natural language processing when there is no …

Simple is not Enough: Document-level Text Simplification using Readability and Coherence

2024-12-24 · Laura Vásquez-Rodríguez, Nhung T. H. Nguyen, Piotr Przybyła, Matthew Shardlow 외

In this paper, we present the SimDoc system, a simplification model considering simplicity, readability, and discourse aspects, such as coherence. In the past decade, the progress of the Text Simplification (TS) field ha…

SentenceText Simplification

Topic Similarity Networks: Visual Analytics for Large Document Sets

2014-09-26 · Arun S. Maiya, Robert M. Rolfe

We investigate ways in which to improve the interpretability of LDA topic models by better analyzing and visualizing their outputs. We focus on examining what we refer to as topic similarity networks: graphs in which nod…

Topic Models

Calculating Semantic Similarity between Academic Articles using Topic Event and Ontology

2017-11-30 · Ming Liu, Bo Lang, Zepeng Gu

Determining semantic similarity between academic documents is crucial to many tasks such as plagiarism detection, automatic technical survey and semantic search. Current studies mostly focus on semantic similarity betwee…

Articlesdocument understandingSemantic SimilaritySemantic Textual Similarity

Hierarchical thematic classification of major conference proceedings

2024-06-21 · Arsentii Kuzmin, Alexander Aduenko, Vadim Strijov

In this paper, we develop a decision support system for the hierarchical text classification. We consider text collections with a fixed hierarchical structure of topics given by experts in the form of a tree. The system …

Bayesian InferenceClassificationtext-classificationText Classification