paper-with-me

홈 › Papers

Can current NLI systems handle German word order? Investigating language model performance on a new German challenge set of minimal pairs

2023-06-07 · Ines Reinig, Katja Markert

Compared to English, German word order is freer and therefore poses additional challenges for natural language inference (NLI). We create WOGLI (Word Order in German Language Inference), the first adversarial NLI dataset for German word order that has the following properties: (i) each premise has an entailed and a non-entailed hypothesis; (ii) premise and hypotheses differ only in word order and necessary morphological changes to mark case and number. In particular, each premise andits two hypotheses contain exactly the same lemmata. Our adversarial examples require the model to use morphological markers in order to recognise or reject entailment. We show that current German autoencoding models fine-tuned on translated NLI data can struggle on this challenge set, reflecting the fact that translated NLI datasets will not mirror all necessary language phenomena in the target language. We also examine performance after data augmentation as well as on related word order phenomena derived from WOGLI. Our datasets are publically available at https://github.com/ireinig/wogli.

📄 PDF Abstract BibTeX arXiv:2306.04523

Code (1)

ireinig/wogli 공식 구현

Tasks

Data AugmentationLanguage ModelingLanguage ModellingNatural Language Inference

Similar Papers 제목 키워드 기반

Continuous Space Reordering Models for Phrase-based MT

2018-01-25 · IWSLT 2017 12 · Nadir Durrani, Fahim Dalvi

Bilingual sequence models improve phrase-based translation and reordering by overcoming phrasal independence assumption and handling long range reordering. However, due to data sparsity, these models often fall back to v…

DecoderPOS

Graph-to-Sequence Neural Machine Translation

2020-09-16 · Sufeng Duan, Hai Zhao, Rui Wang

Neural machine translation (NMT) usually works in a seq2seq learning way by viewing either source or target sentence as a linear sequence of words, which can be regarded as a special case of graph, taking words in the se…

Graph-to-SequenceMachine TranslationNMTSentence+1

Adapting and evaluating a generic term extraction tool

2012-05-01 · LREC 2012 5 · Anita Gojun, Ulrich Heid, Bernd Wei{\ss}bach, Carola Loth 외

We present techniques for monolingual term candidate extraction which are being developed in the EU project TTC. We designed an application for German and English data that serves as a first evaluation of the methods for…

LemmatizationTerm Extraction

An Unsupervised Approach for Mapping between Vector Spaces

2017-11-15 · Syed Sarfaraz Akhtar, Arihant Gupta, Avijit Vajpayee, Arjit Srivastava 외

We present a language independent, unsupervised approach for transforming word embeddings from source language to target language using a transformation matrix. Our model handles the problem of data scarcity which is fac…

Word EmbeddingsWord Similarity

What Makes Word-level Neural Machine Translation Hard: A Case Study on English-German Translation

2016-12-01 · COLING 2016 12 · Fabian Hirschmann, Jinseok Nam, Johannes F{\"u}rnkranz

Traditional machine translation systems often require heavy feature engineering and the combination of multiple techniques for solving different subproblems. In recent years, several end-to-end learning architectures bas…

Feature EngineeringMachine TranslationNMTTranslation