paper-with-me

Semantic Textual Similarity

13개 벤치마크 · 논문 2,406편 · 이 태스크의 논문 보기 →

Benchmarks

STS Benchmark

결과 66개

MRPC

결과 45개

MTEB

결과 30개

SICK

결과 22개

STS13

결과 22개

STS14

결과 21개

STS12

결과 20개

STS15

결과 20개

STS16

결과 20개

SentEval

결과 6개

CxC

결과 4개

MRPC Dev

결과 2개

SICK-R

결과 2개

Most implemented

Papers

The Embedder's Dilemma: LLMs Are Better, but at What Cost?

2026-08-13 · Adnan El Assadi, Niklas Muennighoff, Jinhyuk Lee arxiv

Should you replace your text-embedding pipeline with a large language model? We answer this with a controlled, cost-aware comparison of ten LLMs across six families and 26 embedding models (118M to 14B parameters) on 37 …

Semantic Textual Similarity

Mitigating Scoring Bias in LLM-as-a-Judge via Random Number Generation

2026-08-06 · Yuma Asato, Kiyoaki Shirai, Natthawut Kertkeidkachorn arxiv

Large Language Models (LLMs) are often used as evaluators of text quality, known as LLM-as-a-Judge, which can outperform conventional automatic evaluation metrics that rely on reference texts. However, LLM evaluators ten…

Semantic Textual Similarity

MTEB-BR: A Text Embedding Benchmark for Brazilian Portuguese

2026-07-06 · Tardelli Ronan Coelho Stekel arxiv

Text embeddings for Portuguese have no dedicated benchmark: evaluation rests on translated corpora such as English MS MARCO or on thin multilingual coverage, with native tasks scattered and unconsolidated. We introduce M…

Semantic Textual Similarity

Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders

2026-07-05 · Lucas Hideki Takeuchi Okamura, Alexandre Alcoforado, Anna Helena Reali Costa arxiv

Portuguese remains underrepresented in text embedding evaluation, despite being one of the most widely spoken languages in the world. As a result, embedding models are often selected based on English or multilingual metr…

Semantic Textual SimilarityRepresentation Learning

DySem: Uncovering Dynamic Semantic Components of Large Language Models for Calculating Semantic Textual Similarity

2026-05-28 · Kaijie Zheng, Weiqin Wang, Yile Wang, Hui Huang arxiv

Calculating semantic textual similarity is a foundational task in natural language processing. Current large language models (LLMs) based methods typically rely on extracting last-layer hidden states with fixed dimension…

Semantic Textual SimilaritySemantic SimilarityGeneral Knowledge

MATCHA: Matching Text via Contrastive Semantic Alignment

2026-05-26 · Siran Li, Ece Sena Etoglu, Carsten Eickhoff, Seyed Ali Bahrainian arxiv

Reliable evaluation is essential for understanding large language model (LLM) performance, yet today's go-to metrics, namely token-overlap scores (e.g., ROUGE) and embedding-based measures (e.g., BERTScore), often misjud…

Semantic Textual SimilarityNatural Language InferenceSemantic Similarity

전체 2,406편 보기 →