Semantic Matching of Documents from Heterogeneous Collections: A Simple and Transparent Method for Practical Applications
We present a very simple, unsupervised method for the pairwise matching of documents from heterogeneous collections. We demonstrate our method with the Concept-Project matching task, which is a binary classification task involving pairs of documents from heterogeneous collections. Although our method only employs standard resources without any domain- or task-specific modifications, it clearly outperforms the more complex system of the original authors. In addition, our method is transparent, because it provides explicit information about how a similarity score was computed, and efficient, because it is based on the aggregation of (pre-computable) word-level similarities.
Code (1)
Tasks
Binary ClassificationGeneral ClassificationSimilar Papers 제목 키워드 기반
Semantic Matching of Documents from Heterogeneous Collections: A Simple and Transparent Method for Practical Applications
We present a very simple, unsupervised method for the pairwise matching of documents from heterogeneous collections. We demonstrate our method with the Concept-Project matching task, which is a binary classification task…
Binary ClassificationFast, Small, and Simple Document Listing on Repetitive Text Collections
Document listing on string collections is the task of finding all documents where a pattern appears. It is regarded as the most fundamental document retrieval problem, and is useful in various applications. Many of the f…
RetrievalFinding Salient Context based on Semantic Matching for Relevance Ranking
In this paper, we propose a salient-context based semantic matching method to improve relevance ranking in information retrieval. We first propose a new notion of salient context and then define how to measure it. Then w…
Information RetrievalRetrievalSemantic SimilaritySemantic Textual SimilarityExploratory Analysis of Highly Heterogeneous Document Collections
We present an effective multifaceted system for exploratory analysis of highly heterogeneous document collections. Our system is based on intelligently tagging individual documents in a purely automated fashion and explo…
ArticlesKeyword ExtractionCross-Modal Entity Matching for Visually Rich Documents
Visually rich documents (e.g. leaflets, banners, magazine articles) are physical or digital documents that utilize visual cues to augment their semantics. Information contained in these documents are ad-hoc and often inc…
Articles