Papers Sentence Embedding
“Sentence Embedding” 태그가 달린 논문 407편 · 필터 해제
Intrinsic vs. Extrinsic Evaluation of Czech Sentence Embeddings: Semantic Relevance Doesn't Help with MT Evaluation
In this paper, we compare Czech-specific and multilingual sentence embedding models through intrinsic and extrinsic evaluation paradigms. For intrinsic evaluation, we employ Costra, a complex sentence transformation data…
Machine TranslationSemantic SimilaritySemantic Textual SimilaritySentence+5Factors affecting the in-context learning abilities of LLMs for dialogue state tracking
This study explores the application of in-context learning (ICL) to the dialogue state tracking (DST) problem and investigates the factors that influence its effectiveness. We use a sentence embedding based k-nearest nei…
Dialogue State TrackingIn-Context LearningSentenceSentence Embedding+1Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models
Ensuring safety of large language models (LLMs) is important. Red teaming--a systematic approach to identifying adversarial prompts that elicit harmful responses from target LLMs--has emerged as a crucial safety evaluati…
DiversityRed TeamingSentence EmbeddingSentence-EmbeddingMechanistic Decomposition of Sentence Representations
Sentence embeddings are central to modern NLP and AI systems, yet little is known about their internal structure. While we can compare these embeddings using measures such as cosine similarity, the contributing features …
Dictionary LearningSentenceSentence EmbeddingSentence-Embedding+1Rethinking the Understanding Ability across LLMs through Mutual Information
Recent advances in large language models (LLMs) have revolutionized natural language processing, yet evaluating their intrinsic linguistic understanding remains challenging. Moving beyond specialized evaluation tasks, we…
SentenceSentence EmbeddingSentence-EmbeddingSentence EmbeddingsContrastive Prompting Enhances Sentence Embeddings in LLMs through Inference-Time Steering
Extracting sentence embeddings from large language models (LLMs) is a practical direction, as it requires neither additional data nor fine-tuning. Previous studies usually focus on prompt engineering to guide LLMs to enc…
Prompt EngineeringSemantic Textual SimilaritySentenceSentence Embedding+3Semantic Probabilistic Control of Language Models
Semantic control entails steering LM generations towards satisfying subtle non-lexical constraints, e.g., toxicity, sentiment, or politeness, attributes that can be captured by a sequence-level verifier. It can thus be v…
AttributeSentenceSentence EmbeddingSentence-EmbeddingScalable Unit Harmonization in Medical Informatics Using Bi-directional Transformers and Bayesian-Optimized BM25 and Sentence Embedding Retrieval
Objective: To develop and evaluate a scalable methodology for harmonizing inconsistent units in large-scale clinical datasets, addressing a key barrier to data interoperability. Materials and Methods: We designed a novel…
Bayesian OptimizationInformation RetrievalRe-RankingRetrieval+4Information Leakage of Sentence Embeddings via Generative Embedding Inversion Attacks
Text data are often encoded as dense vectors, known as embeddings, which capture semantic, syntactic, contextual, and domain-specific information. These embeddings, widely adopted in various applications, inherently cont…
SentenceSentence EmbeddingSentence-EmbeddingSentence EmbeddingssEEG-based Encoding for Sentence Retrieval: A Contrastive Learning Approach to Brain-Language Alignment
Interpreting neural activity through meaningful latent representations remains a complex and evolving challenge at the intersection of neuroscience and artificial intelligence. We investigate the potential of multimodal …
Contrastive LearningSentenceSentence EmbeddingSentence-Embedding+1MultiClaimNet: A Massively Multilingual Dataset of Fact-Checked Claim Clusters
In the context of fact-checking, claims are often repeated across various platforms and in different languages, which can benefit from a process that reduces this redundancy. While retrieving previously fact-checked clai…
ClusteringFact CheckingRetrievalSentence+2CASE -- Condition-Aware Sentence Embeddings for Conditional Semantic Textual Similarity Measurement
The meaning conveyed by a sentence often depends on the context in which it appears. Despite the progress of sentence embedding methods, it remains unclear how to best modify a sentence embedding conditioned on its conte…
Dimensionality ReductionLanguage ModelingLanguage ModellingLarge Language Model+7IPCGRL: Language-Instructed Reinforcement Learning for Procedural Level Generation
Recent research has highlighted the significance of natural language in enhancing the controllability of generative models. While various efforts have been made to leverage natural language for content generation, resear…
Deep Reinforcement Learningreinforcement-learningReinforcement LearningSentence+2FanChuan: A Multilingual and Graph-Structured Benchmark For Parody Detection and Analysis
Parody is an emerging phenomenon on social media, where individuals imitate a role or position opposite to their own, often for humor, provocation, or controversy. Detecting and analyzing parody can be challenging and is…
SentenceSentence EmbeddingSentence-EmbeddingSentiment AnalysisExploring RWKV for Sentence Embeddings: Layer-wise Analysis and Baseline Comparison for Semantic Similarity
This paper investigates the efficacy of RWKV, a novel language model architecture known for its linear attention mechanism, for generating sentence embeddings in a zero-shot setting. I conduct a layer-wise analysis to ev…
GPULanguage ModelingLanguage ModellingMRPC+6Evolutionary Algorithms Approach For Search Based On Semantic Document Similarity
Advancements in cloud computing and distributed computing have fostered research activities in Computer science. As a result, researchers have made significant progress in Neural Networks, Evolutionary Computing Algorith…
Cloud ComputingDistributed ComputingEvolutionary AlgorithmsSemantic Similarity+5Task-agnostic Prompt Compression with Context-aware Sentence Embedding and Reward-guided Task Descriptor
The rise of Large Language Models (LLMs) has led to significant interest in prompt compression, a technique aimed at reducing the length of input prompts while preserving critical information. However, the prominent appr…
SentenceSentence EmbeddingSentence-EmbeddingRefining Sentence Embedding Model through Ranking Sentences Generation with Large Language Models
Sentence embedding is essential for many NLP tasks, with contrastive learning methods achieving strong performance using annotated datasets like NLI. Yet, the reliance on manual labels limits scalability. Recent studies …
Contrastive LearningSentenceSentence EmbeddingSentence-EmbeddingPerformance Evaluation of Sentiment Analysis on Text and Emoji Data Using End-to-End, Transfer Learning, Distributed and Explainable AI Models
Emojis are being frequently used in todays digital world to express from simple to complex thoughts more than ever before. Hence, they are also being used in sentiment analysis and targeted marketing campaigns. In this w…
MarketingSentenceSentence EmbeddingSentence-Embedding+4Optimizing Sentence Embedding with Pseudo-Labeling and Model Ensembles: A Hierarchical Framework for Enhanced NLP Tasks
Sentence embedding tasks are important in natural language processing (NLP), but improving their performance while keeping them reliable is still hard. This paper presents a framework that combines pseudo-label generatio…
Data AugmentationPseudo LabelSentenceSentence Embedding+2