paper-with-me

홈 › Papers

SwissGov-RSD: A Human-annotated, Cross-lingual Benchmark for Token-level Recognition of Semantic Differences Between Related Documents

2025-12-08 · Michelle Wastl, Jannis Vamvas, Rico Sennrich arxiv

Recognizing semantic differences across documents is crucial for text generation evaluation and content alignment, especially in cross-lingual settings. However, as a standalone task, it has received little attention. We address this by introducing SwissGov-RSD, the first naturalistic, document-level, cross-lingual dataset for semantic difference recognition. It encompasses a total of 224 multi-parallel documents in English--German, English--French, and English--Italian with token-level difference annotations by human annotators. We evaluate a variety of open-source and closed-source large language models as well as encoder models across different fine-tuning settings on this new benchmark. Our results show that current automatic approaches perform poorly compared to their performance on monolingual, sentence-level, and synthetic benchmarks, revealing a considerable gap for both LLMs and encoder models. We make our code and dataset publicly available.

📄 PDF Abstract BibTeX arXiv:2512.07538

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

MTG: A Benchmarking Suite for Multilingual Text Generation

2021-10-16 · ACL ARR October 2021 10 · Anonymous

We introduce MTG, a new benchmark suite for training and evaluating multilingual text generation. It is the first and largest multilingual multiway text generation benchmark with 400k human-annotated data for four tasks …

BenchmarkingQuestion GenerationQuestion-GenerationStory Generation+3

MTG: A Benchmark Suite for Multilingual Text Generation

2021-08-13 · Findings (NAACL) 2022 7 · Yiran Chen, Zhenqiao Song, Xianze Wu, Danqing Wang 외

We introduce MTG, a new benchmark suite for training and evaluating multilingual text generation. It is the first-proposed multilingual multiway text generation dataset with the largest human-annotated data (400k). It in…

Question GenerationQuestion-GenerationStory GenerationText Generation+2

BCWS: Bilingual Contextual Word Similarity

2018-10-21 · Ta-Chung Chi, Ching-Yen Shih, Yun-Nung Chen

This paper introduces the first dataset for evaluating English-Chinese Bilingual Contextual Word Similarity, namely BCWS (https://github.com/MiuLab/BCWS). The dataset consists of 2,091 English-Chinese word pairs with the…

Word Similarity

Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset

2022-05-25 · Ashish V. Thapliyal, Jordi Pont-Tuset, Xi Chen, Radu Soricut

Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets. In this paper we present the Crossmodal-3600 dataset (XM3600 in short), a geographically diver…

Image CaptioningImage RetrievalImage-text RetrievalImage-to-Text Retrieval+3

Revisiting Cross-Lingual Summarization: A Corpus-based Study and A New Benchmark with Improved Annotation

2023-07-08 · Yulong Chen, Huajian Zhang, Yijie Zhou, Xuefeng Bai 외

Most existing cross-lingual summarization (CLS) work constructs CLS corpora by simply and directly translating pre-annotated summaries from one language to another, which can contain errors from both summarization and tr…

Conversation Summarization