paper-with-me

홈 › Papers

Self-supervised similarity search for large scientific datasets

2021-10-25 · George Stein, Peter Harrington, Jacqueline Blaum, Tomislav Medan, Zarija Lukic

We present the use of self-supervised learning to explore and exploit large unlabeled datasets. Focusing on 42 million galaxy images from the latest data release of the Dark Energy Spectroscopic Instrument (DESI) Legacy Imaging Surveys, we first train a self-supervised model to distill low-dimensional representations that are robust to symmetries, uncertainties, and noise in each image. We then use the representations to construct and publicly release an interactive semantic similarity search tool. We demonstrate how our tool can be used to rapidly discover rare objects given only a single example, increase the speed of crowd-sourcing campaigns, and construct and improve training sets for supervised applications. While we focus on images from sky surveys, the technique is straightforward to apply to any scientific dataset of any dimensionality. The similarity search web app can be found at https://github.com/georgestein/galaxy_search

📄 PDF Abstract BibTeX arXiv:2110.13151

Code (2)

georgestein/galaxy_search 공식 구현
georgestein/ssl-legacysurvey pytorch

Tasks

Self-Supervised LearningSemantic SimilaritySemantic Textual Similarity

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Similarity Analysis of Self-Supervised Speech Representations

2020-10-22 · Yu-An Chung, Yonatan Belinkov, James Glass

Self-supervised speech representation learning has recently been a prosperous research topic. Many algorithms have been proposed for learning useful representations from large-scale unlabeled data, and their applications…

Representation LearningSpeech Representation Learning

The similarity index of scientific publications with equations and formulas, identification of self-plagiarism, and testing of the iThenticate system

2021-12-21 · Andrei D. Polyanin, Inna K. Shingareva

The problems of estimating the similarity index of mathematical and other scientific publications containing equations and formulas are discussed for the first time. It is shown that the presence of equations and formula…

SciConceptMiner: A system for large-scale scientific concept discovery

2021-08-01 · ACL 2021 5 · Zhihong Shen, Chieh-Han Wu, Li Ma, Chien-Pang Chen 외

Scientific knowledge is evolving at an unprecedented rate of speed, with new concepts constantly being introduced from millions of academic articles published every month. In this paper, we introduce a self-supervised en…

Articles

Multi-objective Representation Learning for Scientific Document Retrieval

2022-10-01 · sdp (COLING) 2022 10 · Mathias Parisot, Jakub Zavrel

Existing dense retrieval models for scientific documents have been optimized for either retrieval by short queries, or for document similarity, but usually not for both. In this paper, we explore the space of combining m…

Representation LearningRetrievalSentence

Musical Audio Similarity with Self-supervised Convolutional Neural Networks

2022-02-04 · Carl Thomé, Sebastian Piwell, Oscar Utterbäck

We have built a music similarity search engine that lets video producers search by listenable music excerpts, as a complement to traditional full-text search. Our system suggests similar sounding track segments in a larg…

Triplet