paper-with-me

홈 › Papers

Cross-Level Semantic Similarity for Serbian Newswire Texts

2022-06-01 · LREC 2022 6 · Vuk Batanović, Maja Miličević Petrović

Cross-Level Semantic Similarity (CLSS) is a measure of the level of semantic overlap between texts of different lengths. Although this problem was formulated almost a decade ago, research on it has been sparse, and limited exclusively to the English language. In this paper, we present the first CLSS dataset in another language, in the form of CLSS.news.sr – a corpus of 1000 phrase-sentence and 1000 sentence-paragraph newswire text pairs in Serbian, manually annotated with fine-grained semantic similarity scores using a 0–4 similarity scale. We describe the methodology of data collection and annotation, and compare the resulting corpus to its preexisting counterpart in English, SemEval CLSS, following up with a preliminary linguistic analysis of the newly created dataset. State-of-the-art pre-trained language models are then fine-tuned and evaluated on the CLSS task in Serbian using the produced data, and their settings and results are discussed. The CLSS.news.sr corpus and the guidelines used in its creation are made publicly available.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SimilaritySemantic Textual SimilaritySentence

Similar Papers 제목 키워드 기반

Fine-grained Semantic Textual Similarity for Serbian

2018-05-01 · LREC 2018 5 · Vuk Batanovi{\'c}, Milo{\v{s}} Cvetanovi{\'c}, Bo{\v{s}}ko Nikoli{\'c}
Information RetrievalMachine TranslationNatural Language InferenceQuestion Answering+1

A Massive Scale Semantic Similarity Dataset of Historical English

2023-06-30 · NeurIPS 2023 11

A diversity of tasks use language models trained on semantic similarity data. While there are a variety of datasets that capture semantic similarity, they are either constructed from modern web data or are relatively sma…

ArticlesSemantic SimilaritySemantic Textual Similarity

One Language, Two Scripts: Probing Script-Invariance in LLM Concept Representations

2026-03-09 · Sripad Karne arxiv

Do the features learned by Sparse Autoencoders (SAEs) represent abstract meaning, or are they tied to how text is written? We investigate this question using Serbian digraphia as a controlled testbed: Serbian is written …

How Far Do Auto-Interpretation Labels Generalize: A Controlled Study Across Languages, Scripts, and Rewordings

2026-05-29 · Sripad Karne arxiv

Sparse autoencoder (SAE) features are increasingly used to interpret language models, with auto-generated natural-language labels serving as the primary interface for understanding what each feature represents. We ask wh…

Multi-word Expressions for Abusive Speech Detection in Serbian

2020-12-01 · COLING (MWE) 2020 12 · Ranka Stanković, Jelena Mitrović, Danka Jokić, Cvetana Krstev

This paper presents our work on the refinement and improvement of the Serbian language part of Hurtlex, a multilingual lexicon of words to hurt. We pay special attention to adding Multi-word expressions that can be seen …

Abusive Language