Papers Sentence Similarity
“Sentence Similarity” 태그가 달린 논문 194편 · 필터 해제
EL4NER: Ensemble Learning for Named Entity Recognition via Multiple Small-Parameter Large Language Models
In-Context Learning (ICL) technique based on Large Language Models (LLMs) has gained prominence in Named Entity Recognition (NER) tasks for its lower computing resource consumption, less manual labeling overhead, and str…
Ensemble LearningIn-Context Learningnamed-entity-recognitionNamed Entity Recognition+3Can Vision-Language Models Understand and Interpret Dynamic Gestures from Pedestrians? Pilot Datasets and Exploration Towards Instructive Nonverbal Commands for Cooperative Autonomous Vehicles
In autonomous driving, it is crucial to correctly interpret traffic gestures (TGs), such as those of an authority figure providing orders or instructions, or a pedestrian signaling the driver, to ensure a safe and pleasa…
Autonomous DrivingAutonomous VehiclesSentence SimilarityCoarse-to-Fine Semantic Communication Systems for Text Transmission
Achieving more powerful semantic representations and semantic understanding is one of the key problems in improving the performance of semantic communication systems. This work focuses on enhancing the semantic understan…
Semantic CommunicationSentenceSentence SimilarityHow does a Multilingual LM Handle Multiple Languages?
Multilingual language models have significantly advanced due to rapid progress in natural language processing. Models like BLOOM 1.7B, trained on diverse multilingual datasets, aim to bridge linguistic gaps. However, the…
Multilingual NLPMultilingual Word Embeddingsnamed-entity-recognitionNamed Entity Recognition+9Can linguists better understand DNA?
Multilingual transfer ability, which reflects how well models fine-tuned on one source language can be applied to other languages, has been well studied in multilingual pre-trained models. However, the existence of such …
ClassificationSentenceSentence-Pair ClassificationSentence Similarity3D Spatial Understanding in MLLMs: Disambiguation and Evaluation
Multimodal Large Language Models (MLLMs) have made significant progress in tasks such as image captioning and question answering. However, while these models can generate realistic captions, they often struggle with prov…
3D dense captioning3D visual groundingDense CaptioningImage Captioning+5A Novel Word Pair-based Gaussian Sentence Similarity Algorithm For Bengali Extractive Text Summarization
Extractive Text Summarization is the process of selecting the most representative parts of a larger text without losing any key information. Recent attempts at extractive text summarization in Bengali, either relied on s…
ArticlesExtractive SummarizationExtractive Text SummarizationSentence+2Toeing the Party Line: Election Manifestos as a Key to Understand Political Discourse on Twitter
Political discourse on Twitter is a moving target: politicians continuously make statements about their positions. It is therefore crucial to track their discourse on social media to understand their ideological position…
Political evalutationSemantic Textual SimilaritySentence SimilarityTowards Quantifying The Privacy Of Redacted Text
In this paper we propose use of a k-anonymity-like approach for evaluating the privacy of redacted text. Given a piece of redacted text we use a state of the art transformer-based deep learning network to reconstruct the…
DiversitySentenceSentence SimilarityNo Dataset Needed for Downstream Knowledge Benchmarking: Response Dispersion Inversely Correlates with Accuracy on Domain-specific QA
This research seeks to obviate the need for creating QA datasets and grading (chatbot) LLM responses when comparing LLMs' knowledge in specific topic domains. This is done in an entirely end-user centric way without need…
BenchmarkingChatbotSentence SimilarityEnhancing Semantic Similarity Understanding in Arabic NLP with Nested Embedding Learning
This work presents a novel framework for training Arabic nested embedding models through Matryoshka Embedding Learning, leveraging multilingual, Arabic-specific, and English-based models, to highlight the power of nested…
Natural Language InferenceSemantic SimilaritySemantic Textual SimilaritySentence+2Word Embedding Dimension Reduction via Weakly-Supervised Feature Selection
As a fundamental task in natural language processing, word embedding converts each word into a representation in a vector space. A challenge with word embedding is that as the vocabulary grows, the vector space's dimensi…
Dimensionality Reductionfeature selectionMulti-class ClassificationSentence+1SuperGLEBer: German Language Understanding Evaluation Benchmark
We assemble a broad Natural Language Understanding benchmark suite for the German language and consequently evaluate a wide array of existing German-capable models in order to create a better understanding of the current…
Document ClassificationNatural Language UnderstandingQuestion AnsweringSentence+1OTTAWA: Optimal TransporT Adaptive Word Aligner for Hallucination and Omission Translation Errors Detection
Recently, there has been considerable attention on detecting hallucinations and omissions in Machine Translation (MT) systems. The two dominant approaches to tackle this task involve analyzing the MT system's internal st…
HallucinationMachine TranslationSentenceSentence SimilarityMTEB-French: Resources for French Sentence Embedding Evaluation and Analysis
Recently, numerous embedding models have been made available and widely used for various NLP tasks. The Massive Text Embedding Benchmark (MTEB) has primarily simplified the process of choosing a model that performs well …
SentenceSentence EmbeddingSentence-EmbeddingSentence Embeddings+1Data Augmentation Techniques for Process Extraction from Scientific Publications
We present data augmentation techniques for process extraction tasks in scientific publications. We cast the process extraction task as a sequence labeling task where we identify all the entities in a sentence and label …
Data AugmentationSentenceSentence SimilarityEvaluation of large language model performance on the Biomedical Language Understanding and Reasoning Benchmark
Background The ability of large language models (LLMs) to interpret and generate human-like text has been accompanied with speculation about their application in medicine and clinical research. There is limited data avai…
Document ClassificationLanguage ModelingLanguage ModellingLarge Language Model+8Span-Aggregatable, Contextualized Word Embeddings for Effective Phrase Mining
Dense vector representations for sentences made significant progress in recent years as can be seen on sentence similarity tasks. Real-world phrase retrieval applications, on the other hand, still encounter challenges fo…
RetrievalSentenceSentence EmbeddingsSentence Similarity+3From News to Summaries: Building a Hungarian Corpus for Extractive and Abstractive Summarization
Training summarization models requires substantial amounts of training data. However for less resourceful languages like Hungarian, openly available models and datasets are notably scarce. To address this gap our paper i…
Abstractive Text SummarizationExtractive SummarizationSentenceSentence SimilarityContrastive Learning and Mixture of Experts Enables Precise Vector Embeddings
The advancement of transformer neural networks has significantly elevated the capabilities of sentence similarity models, but they still struggle with highly discriminative tasks and may produce sub-optimal representatio…
Contrastive LearningDescriptiveMixture-of-ExpertsRepresentation Learning+4