Semantic Similarity
7개 벤치마크 · 논문 2,122편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
ERNIE: Enhanced Representation through Knowledge Integration
Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks
Improving Language Understanding by Generative Pre-Training
Language-agnostic BERT Sentence Embedding
MedSTS: A Resource for Clinical Semantic Textual Similarity
Papers
Tracing Generated Samples to Training-Data Clusters in Flow-Matching Models
Understanding which training samples influence a generated image is an important problem in generative modeling. In flow matching, training samples influence the generated image through the velocity field along the gener…
Semantic SimilarityBEACON: Behavior-Anchored Cross-Source Knowledge Graph Construction for Cyber Threat Intelligence
Cyber threat intelligence (CTI) is foundational to modern cyber defense, yet much of it resides in unstructured reports whose volume and heterogeneity far exceed manual analysis, motivating research on automatically cons…
Semantic SimilarityKnowledge GraphsFrom Documents to Reasoning: A Validated Synthetic Data Pipeline and Semantic-Aware Fine-Tuning for Financial Numerical Reasoning
Financial question answering (QA) has emerged as a key benchmark for evaluating the performance of Large Language Models (LLMs) on domain-specific tasks involving complex data formats such as tables, charts, and rich tex…
Synthetic Data GenerationSemantic SimilarityQuestion AnsweringAnswer GenerationGenerative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study
Automatic Speech Recognition (ASR) is typically evaluated using Word Error Rate (WER), which poorly reflects semantic similarity. While embedding-based metrics correlate better with human judgments, the respective roles …
Semantic SimilaritySpeech RecognitionA Storage-Retrieval Gap in Parametric Knowledge Graph Memory
Graph retrieval-augmented generation places retrieved subgraphs into the model's context window at query time, paying a recurring token cost and exposing source data on every call. We study an alternative: compiling a kn…
Semantic SimilarityAraDetox: A Multi-Dialect Arabic Detoxification Dataset
Arabic harmful-language detection has received considerable attention, yet Arabic text detoxification remains underexplored. We introduce AraDetox, a multi-dialect Arabic detoxification dataset comprising 10,500 harmful …
Semantic SimilarityText Generation