paper-with-me

Papers

SIGMORPHON 2021 Shared Task on Morphological Reinflection: Generalization Across Languages

2021-08-01 · ACL (SIGMORPHON) 2021 8 · Tiago Pimentel, Maria Ryskina, Sabrina J. Mielke, Shijie Wu, Eleanor Chodroff, Brian Leonard, Garrett Nicolai, Yustinus Ghanggo Ate, Salam Khalifa, Nizar Habash, Charbel El-Khaissi, Omer Goldman, Michael Gasser, William Lane, Matt Coler, Arturo Oncevay, Jaime Rafael Montoya Samame, Gema Celeste Silva Villegas, Adam Ek, Jean-Philippe Bernardy, Andrey Shcherbakov, Aziyana Bayyr-ool, Karina Sheifer, Sofya Ganieva, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Andrew Krizhanovsky, Natalia Krizhanovsky, Clara Vania, Sardana Ivanova, Aelita Salchak, Christopher Straughn, Zoey Liu, Jonathan North Washington, Duygu Ataman, Witold Kieraś, Marcin Woliński, Totok Suhardijanto, Niklas Stoehr, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo M. Ponti, Grant Aiton, Richard J. Hatcher, Emily Prud'hommeaux, Ritesh Kumar, Mans Hulden, Botond Barta, Dorina Lakatos, Gábor Szolnok, Judit Ács, Mohit Raj, David Yarowsky, Ryan Cotterell, Ben Ambridge, Ekaterina Vylomova

This year's iteration of the SIGMORPHON Shared Task on morphological reinflection focuses on typological diversity and cross-lingual variation of morphosyntactic features. In terms of the task, we enrich UniMorph with new data for 32 languages from 13 language families, with most of them being under-resourced: Kunwinjku, Classical Syriac, Arabic (Modern Standard, Egyptian, Gulf), Hebrew, Amharic, Aymara, Magahi, Braj, Kurdish (Central, Northern, Southern), Polish, Karelian, Livvi, Ludic, Veps, Võro, Evenki, Xibe, Tuvan, Sakha, Turkish, Indonesian, Kodi, Seneca, Asháninka, Yanesha, Chukchi, Itelmen, Eibela. We evaluate six systems on the new data and conduct an extensive error analysis of the systems' predictions. Transformer-based models generally demonstrate superior performance on the majority of languages, achieving >90% accuracy on 65% of them. The languages on which systems yielded low accuracy are mainly under-resourced, with a limited amount of data. Most errors made by the systems are due to allomorphy, honorificity, and form variation. In addition, we observe that systems especially struggle to inflect multiword lemmas. The systems also produce misspelled forms or end up in repetitive loops (e.g., RNN-based models). Finally, we report a large drop in systems' performance on previously unseen lemmas.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Diversity

Similar Papers 제목 키워드 기반

The SIGMORPHON 2016 Shared Task---Morphological Reinflection

2016-08-01 · WS 2016 8 · Ryan Cotterell, Christo Kirov, John Sylak-Glassman, David Yarowsky 외
Morphological Analysis

ISI at the SIGMORPHON 2017 Shared Task on Morphological Reinflection

2017-08-01 · CONLL 2017 8 · Abhisek Chakrabarty, Utpal Garain

UZH at CoNLL--SIGMORPHON 2018 Shared Task on Universal Morphological Reinflection

2018-10-01 · CONLL 2018 10 · Peter Makarov, Simon Clematide
Imitation LearningMorphological Inflection

MED: The LMU System for the SIGMORPHON 2016 Shared Task on Morphological Reinflection

2016-08-01 · WS 2016 8 · Katharina Kann, Hinrich Sch{\"u}tze
Machine TranslationMorphological InflectionQuestion Answering

Proceedings of the CoNLL SIGMORPHON 2017 Shared Task: Universal Morphological Reinflection

2017-08-01 · CONLL 2017 8 ·