Papers Document Embedding
“Document Embedding” 태그가 달린 논문 78편 · 필터 해제
AC-LoRA: (Almost) Training-Free Access Control-Aware Multi-Modal LLMs
Corporate LLMs are gaining traction for efficient knowledge dissemination and management within organizations. However, as current LLMs are vulnerable to leaking sensitive information, it has proven difficult to apply th…
Document EmbeddingBERTopic for Topic Modeling of Hindi Short Texts: A Comparative Study
As short text data in native languages like Hindi increasingly appear in modern media, robust methods for topic modeling on such data have gained importance. This study investigates the performance of BERTopic in modelin…
Document EmbeddingTopic ModelsTrajectories of Change: Approaches for Tracking Knowledge Evolution
We explore local vs. global evolution of knowledge systems through the framework of socio-epistemic networks (SEN), applying two complementary methods to a corpus of scientific texts. The framework comprises three interc…
Document EmbeddingEnhancing Question Answering Precision with Optimized Vector Retrieval and Instructions
Question-answering (QA) is an important application of Information Retrieval (IR) and language models, and the latest trend is toward pre-trained large neural networks with embedding parameters. Augmenting QA performance…
Document EmbeddingInformation RetrievalQuestion AnsweringRetrieval+2Contextual Document Embeddings
Dense document embeddings are central to neural retrieval. The dominant paradigm is to train and construct embeddings by running encoders directly on individual documents. In this work, we argue that these embeddings, wh…
Contrastive LearningDocument EmbeddingGPUMTEB Benchmark+2QAEncoder: Towards Aligned Representation Learning in Question Answering System
Modern QA systems entail retrieval-augmented generation (RAG) for accurate and trustworthy responses. However, the inherent gap between user queries and relevant documents hinders precise matching. Motivated by our conic…
Document EmbeddingQuestion AnsweringRAGRepresentation Learning+2DocNet: Semantic Structure in Inductive Bias Detection Models
News will be biased so long as people have opinions. As social media becomes the primary entry point for news and partisan differences increase, it is increasingly important for informed citizens to be able to recognize …
ArticlesBias DetectionDocument EmbeddingInductive Biasrollama: An R package for using generative large language models through Ollama
rollama is an R package that wraps the Ollama API, which allows you to run different Generative Large Language Models (GLLM) locally. The package and learning material focus on making it easy to use Ollama for annotating…
Document EmbeddingARAGOG: Advanced RAG Output Grading
Retrieval-Augmented Generation (RAG) is essential for integrating external knowledge into Large Language Model (LLM) outputs. While the literature on RAG is growing, it primarily focuses on systematic reviews and compari…
Document EmbeddingLanguage ModelingLanguage ModellingLarge Language Model+5HILL: Hierarchy-aware Information Lossless Contrastive Learning for Hierarchical Text Classification
Existing self-supervised methods in natural language processing (NLP), especially hierarchical text classification (HTC), mainly focus on self-supervised contrastive learning, extremely relying on human-designed augmenta…
Contrastive LearningDocument EmbeddingHierarchical Multi-label ClassificationRepresentation Learning+2Word Embeddings Revisited: Do LLMs Offer Something New?
Learning meaningful word embeddings is key to training a robust language model. The recent rise of Large Language Models (LLMs) has provided us with many new word/sentence/document embedding models. Although LLMs have sh…
Document EmbeddingLanguage ModelingLanguage ModellingSentence+1KeyGen2Vec: Learning Document Embedding via Multi-label Keyword Generation in Question-Answering
Representing documents into high dimensional embedding space while preserving the structural similarity between document sources has been an ultimate goal for many works on text representation learning. Current embedding…
Document EmbeddingKeyphrase GenerationQuestion AnsweringRepresentation LearningFacilitating Interdisciplinary Knowledge Transfer with Research Paper Recommender Systems
In the extensive recommender systems literature, novelty and diversity have been identified as key properties of useful recommendations. However, these properties have received limited attention in the specific sub-field…
DiversityDocument EmbeddingRecommendation SystemsTransfer LearningA Novel Method of Fuzzy Topic Modeling based on Transformer Processing
Topic modeling is admittedly a convenient way to monitor markets trend. Conventionally, Latent Dirichlet Allocation, LDA, is considered a must-do model to gain this type of information. By given the merit of deducing key…
Document EmbeddingShuffle & Divide: Contrastive Learning for Long Text
We propose a self-supervised learning method for long text documents based on contrastive learning. A key to our method is Shuffle and Divide (SaD), a simple text augmentation algorithm that sets up a pretext task requir…
Contrastive LearningDocument EmbeddingSelf-Supervised LearningText Augmentation+3Caching Historical Embeddings in Conversational Search
Rapid response, namely low latency, is fundamental in search applications; it is particularly so in interactive search sessions, such as those encountered in conversational settings. An observation with a potential to re…
Conversational SearchDocument EmbeddingRetrievalLEA: Meta Knowledge-Driven Self-Attentive Document Embedding for Few-Shot Text Classification
Text classification has achieved great success with the prosperity of deep learning and pre-trained language models. However, we often encounter labeled data deficiency problems in real-world text-classification tasks. T…
ClassificationDocument EmbeddingFew-Shot LearningFew-Shot Text Classification+3Assessing the trade-off between prediction accuracy and interpretability for topic modeling on energetic materials corpora
As the amount and variety of energetics research increases, machine aware topic identification is necessary to streamline future research pipelines. The makeup of an automatic topic identification process consists of cre…
Document EmbeddingPredictionApproach to Predicting News -- A Precise Multi-LSTM Network With BERT
Varieties of Democracy (V-Dem) is a new approach to conceptualizing and measuring democracy and politics. It has information for 200 countries and is one of the biggest databases for political science. According to the V…
Document EmbeddingSentenceSentence EmbeddingsWord EmbeddingsAcademic Resource Text Level Multi-label Classification based on Attention
Hierarchical multi-label academic text classification (HMTC) is to assign academic texts into a hierarchically structured labeling system. We propose an attention-based hierarchical multi-label classification algorithm o…
ClassificationDocument EmbeddingHierarchical Multi-label ClassificationMulti-Label Classification+3