paper-with-me

Papers

WAGS: A Beautiful English-Italian Benchmark Supporting Word Alignment Evaluation on Rare Words

2016-05-01 · LREC 2016 5 · Luisa Bentivogli, Mauro Cettolo, M. Amin Farajian, Marcello Federico

This paper presents WAGS (Word Alignment Gold Standard), a novel benchmark which allows extensive evaluation of WA tools on out-of-vocabulary (OOV) and rare words. WAGS is a subset of the Common Test section of the Europarl English-Italian parallel corpus, and is specifically tailored to OOV and rare words. WAGS is composed of 6,715 sentence pairs containing 11,958 occurrences of OOV and rare words up to frequency 15 in the Europarl Training set (5,080 English words and 6,878 Italian words), representing almost 3{\%} of the whole text. Since WAGS is focused on OOV/rare words, manual alignments are provided for these words only, and not for the whole sentences. Two off-the-shelf word aligners have been evaluated on WAGS, and results have been compared to those obtained on an existing benchmark tailored to full text alignment. The results obtained confirm that WAGS is a valuable resource, which allows a statistically sound evaluation of WA systems{'} performance on OOV and rare words, as well as extensive data analyses. WAGS is publicly released under a Creative Commons Attribution license.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceWord Alignment

Similar Papers 제목 키워드 기반

Annotating Information Structure in Italian: Characteristics and Cross-Linguistic Applicability of a QUD-Based Approach

2019-08-01 · WS 2019 8 · Kordula De Kuthy, Lisa Brunetti, Marta Berardi

We present a discourse annotation study, in which an annotation method based on Questions under Discussion (QuD) is applied to Italian data. The results of our inter-annotator agreement analysis show that the QUD-based a…

Culturally Grounded Physical Commonsense Reasoning in Italian and English: A Submission to the MRL 2025 Shared Task

2025-10-26 · Marco De Santis, Lisa Alazraki arxiv

This paper presents our submission to the MRL 2025 Shared Task on Multilingual Physical Reasoning Datasets. The objective of the shared task is to create manually-annotated evaluation data in the physical commonsense rea…

Physical Commonsense Reasoning

HATE-ITA: New Baselines for Hate Speech Detection in Italian

2022-07-01 · NAACL (WOAH) 2022 7 · Debora Nozza, Federico Bianchi, Giuseppe Attanasio

Online hate speech is a dangerous phenomenon that can (and should) be promptly counteracted properly. While Natural Language Processing supplies appropriate algorithms for trying to reach this objective, all research eff…

BenchmarkingHate Speech DetectionLanguage ModelingLanguage Modelling

DIETA: A Decoder-only transformer-based model for Italian-English machine TrAnslation

2026-01-25 · Pranav Kasela, Marco Braga, Alessandro Ghiotto, Andrea Pilzer 외 arxiv

In this paper, we present DIETA, a small, decoder-only Transformer model with 0.5 billion parameters, specifically designed and trained for Italian-English machine translation. We collect and curate a large parallel corp…

Machine Translation

Women Are Beautiful, Men Are Leaders: Gender Stereotypes in Machine Translation and Language Modeling

2023-11-30 · Matúš Pikuliak, Andrea Hrckova, Stefan Oresko, Marián Šimko

We present GEST -- a new manually created dataset designed to measure gender-stereotypical reasoning in language models and machine translation systems. GEST contains samples for 16 gender stereotypes about men and women…

Language ModelingLanguage ModellingMachine TranslationTranslation