paper-with-me

홈 › Papers

TAPE: Assessing Few-shot Russian Language Understanding

2022-10-23 · Ekaterina Taktasheva, Tatiana Shavrina, Alena Fenogenova, Denis Shevelev, Nadezhda Katricheva, Maria Tikhonova, Albina Akhmetgareeva, Oleg Zinkevich, Anastasiia Bashmakova, Svetlana Iordanskaia, Alena Spiridonova, Valentina Kurenshchikova, Ekaterina Artemova, Vladislav Mikhailov

Recent advances in zero-shot and few-shot learning have shown promise for a scope of research and practical purposes. However, this fast-growing area lacks standardized evaluation suites for non-English languages, hindering progress outside the Anglo-centric paradigm. To address this line of research, we propose TAPE (Text Attack and Perturbation Evaluation), a novel benchmark that includes six more complex NLU tasks for Russian, covering multi-hop reasoning, ethical concepts, logic and commonsense knowledge. The TAPE's design focuses on systematic zero-shot and few-shot NLU evaluation: (i) linguistic-oriented adversarial attacks and perturbations for analyzing robustness, and (ii) subpopulations for nuanced interpretation. The detailed analysis of testing the autoregressive baselines indicates that simple spelling-based perturbations affect the performance the most, while paraphrasing the input has a more negligible effect. At the same time, the results demonstrate a significant gap between the neural and human baselines for most tasks. We publicly release TAPE (tape-benchmark.com) to foster research on robust LMs that can generalize to new tasks when little to no supervision is available.

📄 PDF Abstract BibTeX arXiv:2210.12813

Code (1)

RussianNLP/TAPE 공식 구현

Tasks

Adversarial AttackAdversarial TextEthicsFew-Shot LearningLogical ReasoningQuestion AnsweringZero-Shot Learning

Similar Papers 제목 키워드 기반

RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark

2020-10-29 · EMNLP 2020 11 · Tatiana Shavrina, Alena Fenogenova, Anton Emelyanov, Denis Shevelev 외

In this paper, we introduce an advanced Russian general language understanding evaluation benchmark -- RussianGLUE. Recent advances in the field of universal language models and transformers require the development of a …

Common Sense ReasoningDiagnosticLexical EntailmentLogical Reasoning Question Answering+5

Comparing LLM Text Annotation Skills: A Study on Human Rights Violations in Social Media Data

2025-05-15 · Poli Apollinaire Nemkova, Solomon Ubani, Mark V. Albert

In the era of increasingly sophisticated natural language processing (NLP) systems, large language models (LLMs) have demonstrated remarkable potential for diverse applications, including tasks requiring nuanced textual …

Binary Classificationtext annotation

The Russian-focused embedders' exploration: ruMTEB benchmark and Russian embedding model design

2024-08-22 · Artem Snegirev, Maria Tikhonova, Anna Maksimova, Alena Fenogenova 외

Embedding models play a crucial role in Natural Language Processing (NLP) by creating text embeddings used in various tasks such as information retrieval and assessing semantic text similarity. This paper focuses on rese…

Information RetrievalRerankingRetrievalSemantic Textual Similarity+3

Cross-Cultural Simulation of Citizen Emotional Responses to Bureaucratic Red Tape Using LLM Agents

2026-04-14 · Wanchun Ni, Jiugeng Sun, Yixian Liu, Mennatallah El-Assady arxiv

Improving policymaking is a central concern in public administration. Prior human subject studies reveal substantial cross-cultural differences in citizens' emotional responses to red tape during policy implementation. W…

KoBigBird-large: Transformation of Transformer for Korean Language Understanding

2023-09-19 · Kisu Yang, Yoonna Jang, Taewoo Lee, Jinwoo Seong 외

This work presents KoBigBird-large, a large size of Korean BigBird that achieves state-of-the-art performance and allows long sequence processing for Korean language understanding. Without further pretraining, we only tr…

Document ClassificationQuestion Answering