paper-with-me

Papers Sentence Similarity

“Sentence Similarity” 태그가 달린 논문 194편 · 필터 해제

EL4NER: Ensemble Learning for Named Entity Recognition via Multiple Small-Parameter Large Language Models

2025-05-29 · Yuzhen Xiao, Jiahe Song, Yongxin Xu, Ruizhe Zhang 외

In-Context Learning (ICL) technique based on Large Language Models (LLMs) has gained prominence in Named Entity Recognition (NER) tasks for its lower computing resource consumption, less manual labeling overhead, and str…

Ensemble LearningIn-Context Learningnamed-entity-recognitionNamed Entity Recognition+3

Can Vision-Language Models Understand and Interpret Dynamic Gestures from Pedestrians? Pilot Datasets and Exploration Towards Instructive Nonverbal Commands for Cooperative Autonomous Vehicles

2025-04-15 · Tonko E. W. Bossen, Andreas Møgelmose, Ross Greer

In autonomous driving, it is crucial to correctly interpret traffic gestures (TGs), such as those of an authority figure providing orders or instructions, or a pedestrian signaling the driver, to ensure a safe and pleasa…

Autonomous DrivingAutonomous VehiclesSentence Similarity

Coarse-to-Fine Semantic Communication Systems for Text Transmission

2025-04-02 · Mengli Tao, Jiancun Fan, Jie Luo, Huiqiang Xie

Achieving more powerful semantic representations and semantic understanding is one of the key problems in improving the performance of semantic communication systems. This work focuses on enhancing the semantic understan…

Semantic CommunicationSentenceSentence Similarity

How does a Multilingual LM Handle Multiple Languages?

2025-02-06 · Santhosh Kakarla, Gautama Shastry Bulusu Venkata, Aishwarya Gaddam

Multilingual language models have significantly advanced due to rapid progress in natural language processing. Models like BLOOM 1.7B, trained on diverse multilingual datasets, aim to bridge linguistic gaps. However, the…

Multilingual NLPMultilingual Word Embeddingsnamed-entity-recognitionNamed Entity Recognition+9

Can linguists better understand DNA?

2024-12-10 · Wang Liang

Multilingual transfer ability, which reflects how well models fine-tuned on one source language can be applied to other languages, has been well studied in multilingual pre-trained models. However, the existence of such …

ClassificationSentenceSentence-Pair ClassificationSentence Similarity

3D Spatial Understanding in MLLMs: Disambiguation and Evaluation

2024-12-09 · Chun-Peng Chang, Alain Pagani, Didier Stricker

Multimodal Large Language Models (MLLMs) have made significant progress in tasks such as image captioning and question answering. However, while these models can generate realistic captions, they often struggle with prov…

3D dense captioning3D visual groundingDense CaptioningImage Captioning+5

A Novel Word Pair-based Gaussian Sentence Similarity Algorithm For Bengali Extractive Text Summarization

2024-11-26 · Fahim Morshed, Md. Abdur Rahman, Sumon Ahmed

Extractive Text Summarization is the process of selecting the most representative parts of a larger text without losing any key information. Recent attempts at extractive text summarization in Bengali, either relied on s…

ArticlesExtractive SummarizationExtractive Text SummarizationSentence+2

Toeing the Party Line: Election Manifestos as a Key to Understand Political Discourse on Twitter

2024-10-21 · Maximilian Maurer, Tanise Ceron, Sebastian Padó, Gabriella Lapesa

Political discourse on Twitter is a moving target: politicians continuously make statements about their positions. It is therefore crucial to track their discourse on social media to understand their ideological position…

Political evalutationSemantic Textual SimilaritySentence Similarity

Towards Quantifying The Privacy Of Redacted Text

2024-10-10 · Vaibhav Gusain, Douglas Leith

In this paper we propose use of a k-anonymity-like approach for evaluating the privacy of redacted text. Given a piece of redacted text we use a state of the art transformer-based deep learning network to reconstruct the…

DiversitySentenceSentence Similarity

No Dataset Needed for Downstream Knowledge Benchmarking: Response Dispersion Inversely Correlates with Accuracy on Domain-specific QA

2024-08-24 · Robert L Simione II

This research seeks to obviate the need for creating QA datasets and grading (chatbot) LLM responses when comparing LLMs' knowledge in specific topic domains. This is done in an entirely end-user centric way without need…

BenchmarkingChatbotSentence Similarity

Enhancing Semantic Similarity Understanding in Arabic NLP with Nested Embedding Learning

2024-07-30 · Omer Nacar, Anis Koubaa

This work presents a novel framework for training Arabic nested embedding models through Matryoshka Embedding Learning, leveraging multilingual, Arabic-specific, and English-based models, to highlight the power of nested…

Natural Language InferenceSemantic SimilaritySemantic Textual SimilaritySentence+2

Word Embedding Dimension Reduction via Weakly-Supervised Feature Selection

2024-07-17 · Jintang Xue, Yun-Cheng Wang, Chengwei Wei, C. -C. Jay Kuo

As a fundamental task in natural language processing, word embedding converts each word into a representation in a vector space. A challenge with word embedding is that as the vocabulary grows, the vector space's dimensi…

Dimensionality Reductionfeature selectionMulti-class ClassificationSentence+1

SuperGLEBer: German Language Understanding Evaluation Benchmark

2024-06-20 · NAACL 2024 6 · Jan Pfister, Andreas Hotho

We assemble a broad Natural Language Understanding benchmark suite for the German language and consequently evaluate a wide array of existing German-capable models in order to create a better understanding of the current…

Document ClassificationNatural Language UnderstandingQuestion AnsweringSentence+1

OTTAWA: Optimal TransporT Adaptive Word Aligner for Hallucination and Omission Translation Errors Detection

2024-06-04 · Chenyang Huang, Abbas Ghaddar, Ivan Kobyzev, Mehdi Rezagholizadeh 외

Recently, there has been considerable attention on detecting hallucinations and omissions in Machine Translation (MT) systems. The two dominant approaches to tackle this task involve analyzing the MT system's internal st…

HallucinationMachine TranslationSentenceSentence Similarity

MTEB-French: Resources for French Sentence Embedding Evaluation and Analysis

2024-05-30 · Mathieu Ciancone, Imene Kerboua, Marion Schaeffer, Wissam Siblini

Recently, numerous embedding models have been made available and widely used for various NLP tasks. The Massive Text Embedding Benchmark (MTEB) has primarily simplified the process of choosing a model that performs well …

SentenceSentence EmbeddingSentence-EmbeddingSentence Embeddings+1

Data Augmentation Techniques for Process Extraction from Scientific Publications

2024-05-23 · Yuni Susanti

We present data augmentation techniques for process extraction tasks in scientific publications. We cast the process extraction task as a sequence labeling task where we identify all the entities in a sentence and label …

Data AugmentationSentenceSentence Similarity

Evaluation of large language model performance on the Biomedical Language Understanding and Reasoning Benchmark

2024-05-17 · medRxiv 2024 5 · Hui Feng, Francesco Ronzano, Jude LaFleur, Matthew Garber 외

Background The ability of large language models (LLMs) to interpret and generate human-like text has been accompanied with speculation about their application in medicine and clinical research. There is limited data avai…

Document ClassificationLanguage ModelingLanguage ModellingLarge Language Model+8

Span-Aggregatable, Contextualized Word Embeddings for Effective Phrase Mining

2024-05-12 · Eyal Orbach, Lev Haikin, Nelly David, Avi Faizakof

Dense vector representations for sentences made significant progress in recent years as can be seen on sentence similarity tasks. Real-world phrase retrieval applications, on the other hand, still encounter challenges fo…

RetrievalSentenceSentence EmbeddingsSentence Similarity+3

From News to Summaries: Building a Hungarian Corpus for Extractive and Abstractive Summarization

2024-04-04 · Botond Barta, Dorina Lakatos, Attila Nagy, Milán Konor Nyist 외

Training summarization models requires substantial amounts of training data. However for less resourceful languages like Hungarian, openly available models and datasets are notably scarce. To address this gap our paper i…

Abstractive Text SummarizationExtractive SummarizationSentenceSentence Similarity

Contrastive Learning and Mixture of Experts Enables Precise Vector Embeddings

2024-01-28 · Logan Hallee, Rohan Kapur, Arjun Patel, Jason P. Gleghorn 외

The advancement of transformer neural networks has significantly elevated the capabilities of sentence similarity models, but they still struggle with highly discriminative tasks and may produce sub-optimal representatio…

Contrastive LearningDescriptiveMixture-of-ExpertsRepresentation Learning+4
1–20 / 194 다음 →