paper-with-me

홈 › Papers

Multifaceted Domain-Specific Document Embeddings

2021-06-01 · NAACL 2021 4 · Julian Risch, Philipp Hager, Ralf Krestel

Current document embeddings require large training corpora but fail to learn high-quality representations when confronted with a small number of domain-specific documents and rare terms. Further, they transform each document into a single embedding vector, making it hard to capture different notions of document similarity or explain why two documents are considered similar. In this work, we propose our Faceted Domain Encoder, a novel approach to learn multifaceted embeddings for domain-specific documents. It is based on a Siamese neural network architecture and leverages knowledge graphs to further enhance the embeddings even if only a few training samples are available. The model identifies different types of domain knowledge and encodes them into separate dimensions of the embedding, thereby enabling multiple ways of finding and comparing related documents in the vector space. We evaluate our approach on two benchmark datasets and find that it achieves the same embedding quality as state-of-the-art models while requiring only a tiny fraction of their training data. An interactive demo, our source code, and the evaluation datasets are available online: https://hpi.de/naumann/s/multifaceted-embeddings and a screencast is available on YouTube: https://youtu.be/HHcsX2clEwg

📄 PDF Abstract BibTeX

Code (1)

philipphager/faceted-domain-encoder 공식 구현 pytorch

Tasks

Document EmbeddingKnowledge Graphs

Methods 이 논문이 사용한 방법론

Siamese Network 설명 없음

Similar Papers 제목 키워드 기반

SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain Mapping

2025-10-13 · Marc Brinner, Sina Zarrieß arxiv

We propose SemCSE-Multi, a novel unsupervised framework for generating multifaceted embeddings of scientific abstracts, evaluated in the domains of invasion biology and medicine. These embeddings capture distinct, indivi…

Rethinking ANN-based Retrieval: Multifaceted Learnable Index for Large-scale Recommendation System

2026-02-18 · Jiang Zhang, Yubo Wang, Wei Chang, Lu Han 외 arxiv

Approximate nearest neighbor (ANN) search is widely used in the retrieval stage of large-scale recommendation systems. In this stage, candidate items are indexed using their learned embedding vectors, and ANN search is e…

Recommendation Systems

MULTI3IR: A Benchmark for Multi-perspective Multi-domain Multi-modal Information Retrieval

2026-08-31 · Seokwon Song, Sohyeon Kim, Gunhee Kim hf

Information retrieval (IR) increasingly targets open-ended queries that admit diverse perspectives. Existing IR benchmarks, however, focus primarily on closed-ended queries, while even open-ended benchmarks largely consi…

Information Retrieval

Hybrid Improved Document-level Embedding (HIDE)

2020-06-01 · Satanik Mitra, Mamata Jenamani

In recent times, word embeddings are taking a significant role in sentiment analysis. As the generation of word embeddings needs huge corpora, many applications use pretrained embeddings. In spite of the success, word em…

Sentiment AnalysisWord Embeddings

Learning Word Embeddings for Data Sparse and Sentiment Rich Data Sets

2018-06-01 · NAACL 2018 6 · Prathusha Kameswara Sarma

This research proposal describes two algorithms that are aimed at learning word embeddings for data sparse and sentiment rich data sets. The goal is to use word embeddings adapted for domain specific data sets in downstr…

General ClassificationLearning Word EmbeddingsSentiment AnalysisSentiment Classification+2