paper-with-me

Semantic Similarity

7개 벤치마크 · 논문 2,122편 · 이 태스크의 논문 보기 →

Benchmarks

SICK

결과 15개

BIOSSES

결과 9개

CHIP-STS

결과 3개

ClinicalSTS

결과 3개

MedSTS

결과 3개

Most implemented

Language-agnostic BERT Sentence Embedding

2020-07-03 · 구현 6개

Papers

Tracing Generated Samples to Training-Data Clusters in Flow-Matching Models

2026-08-30 · Rania Briq, Ohad Fried, Michael Kamp, Stefan Kesselheim arxiv

Understanding which training samples influence a generated image is an important problem in generative modeling. In flow matching, training samples influence the generated image through the velocity field along the gener…

Semantic Similarity

BEACON: Behavior-Anchored Cross-Source Knowledge Graph Construction for Cyber Threat Intelligence

2026-08-28 · Changze Li, Yutong Cheng, Tsania Camila Finnisa, Qian Cui 외 arxiv

Cyber threat intelligence (CTI) is foundational to modern cyber defense, yet much of it resides in unstructured reports whose volume and heterogeneity far exceed manual analysis, motivating research on automatically cons…

Semantic SimilarityKnowledge Graphs

From Documents to Reasoning: A Validated Synthetic Data Pipeline and Semantic-Aware Fine-Tuning for Financial Numerical Reasoning

2026-08-28 · Lokendra Birla, Milind Savagaonkar, Visnu Srinivasan, Sowmya Rasipuram 외 arxiv

Financial question answering (QA) has emerged as a key benchmark for evaluating the performance of Large Language Models (LLMs) on domain-specific tasks involving complex data formats such as tables, charts, and rich tex…

Synthetic Data GenerationSemantic SimilarityQuestion AnsweringAnswer Generation

Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study

2026-08-26 · Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil, Sergio Burdisso 외 arxiv

Automatic Speech Recognition (ASR) is typically evaluated using Word Error Rate (WER), which poorly reflects semantic similarity. While embedding-based metrics correlate better with human judgments, the respective roles …

Semantic SimilaritySpeech Recognition

A Storage-Retrieval Gap in Parametric Knowledge Graph Memory

2026-08-26 · Martino M. L. Pulici, Cuong Xuan Chu, Evgeny Kharlamov, Volker Tresp arxiv

Graph retrieval-augmented generation places retrieved subgraphs into the model's context window at query time, paying a recurring token cost and exposing source data on every call. We study an alternative: compiling a kn…

Semantic Similarity

AraDetox: A Multi-Dialect Arabic Detoxification Dataset

2026-08-24 · Mo El-Haj arxiv

Arabic harmful-language detection has received considerable attention, yet Arabic text detoxification remains underexplored. We introduce AraDetox, a multi-dialect Arabic detoxification dataset comprising 10,500 harmful …

Semantic SimilarityText Generation

전체 2,122편 보기 →