Measuring Similarity: Computationally Reproducing the Scholar's Interests
Computerized document classification already orders the news articles that Apple's "News" app or Google's "personalized search" feature groups together to match a reader's interests. The invisible and therefore illegible decisions that go into these tailored searches have been the subject of a critique by scholars who emphasize that our intelligence about documents is only as good as our ability to understand the criteria of search. This article will attempt to unpack the procedures used in computational classification of texts, translating them into term legible to humanists, and examining opportunities to render the computational text classification process subject to expert critique and improvement.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesClassificationDocument ClassificationGeneral Classificationtext-classificationText ClassificationSimilar Papers 제목 키워드 기반
Dataset Search In Biodiversity Research: Do Metadata In Data Repositories Reflect Scholarly Information Needs?
The increasing amount of research data provides the opportunity to link and integrate data to create novel hypotheses, to repeat experiments or to compare recent data to data collected at a different time or place. Howev…
RetrievalMeasuring dissimilarity with diffeomorphism invariance
Measures of similarity (or dissimilarity) are a key ingredient to many machine learning algorithms. We introduce DID, a pairwise dissimilarity measure applicable to a wide range of data spaces, which leverages the data's…
A Kernelized Stein Discrepancy for Goodness-of-fit Tests and Model Evaluation
We derive a new discrepancy statistic for measuring differences between two probability distributions based on combining Stein's identity with the reproducing kernel Hilbert space theory. We apply our result to test how …
Metrics for Inter-Dataset Similarity with Example Applications in Synthetic Data and Feature Selection Evaluation -- Extended Version
Measuring inter-dataset similarity is an important task in machine learning and data mining with various use cases and applications. Existing methods for measuring inter-dataset similarity are computationally expensive, …
feature selectionOnline Asymmetric Similarity Learning for Cross-Modal Retrieval
Cross-modal retrieval has attracted intensive attention in recent years. Measuring the semantic similarity between heterogeneous data objects is an essential yet challenging problem in cross-modal retrieval. In this pape…
Cross-Modal RetrievalRetrievalSemantic SimilaritySemantic Textual Similarity