Common Sense Reasoning on PARus
Accuracy
- 2020-10-22 — MT5 Large: Accuracy 0.504
- 2020-10-29 — Human Benchmark: Accuracy 0.982
| Rank | Model | Accuracy | Paper | Code | Year |
|---|---|---|---|---|---|
| 1 | Human Benchmark | 0.982 | RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark | RussianNLP/RussianSuperGLUE · RussianNLP/MOROCCO | 2020 |
| 2 | Golden Transformer | 0.908 | |||
| 3 | YaLM 1.0B few-shot | 0.766 | |||
| 4 | RuGPT3XL few-shot | 0.676 | |||
| 5 | ruT5-large-finetune | 0.66 | |||
| 6 | RuGPT3Medium | 0.598 | |||
| 7 | RuGPT3Large | 0.584 | |||
| 8 | RuBERT plain | 0.574 | |||
| 9 | RuGPT3Small | 0.562 | |||
| 10 | ruT5-base-finetune | 0.554 | |||
| 11 | Multilingual Bert | 0.528 | |||
| 12 | ruRoberta-large finetune | 0.508 | |||
| 12 | RuBERT conversational | 0.508 | |||
| 14 | MT5 Large | 0.504 | mT5: A massively multilingual pre-trained text-to-text transformer | huggingface/transformers · google-research/multilingual-t5 · google-research/byt5 · +5 | 2020 |
| 15 | SBERT_Large_mt_ru_finetuning | 0.498 | |||
| 15 | SBERT_Large | 0.498 | |||
| 15 | majority_class | 0.498 | Unreasonable Effectiveness of Rule-Based Heuristics in Solving Russian SuperGLUE Tasks | 2021 | |
| 18 | ruBert-large finetune | 0.492 | |||
| 19 | Baseline TF-IDF1.1 | 0.486 | RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark | RussianNLP/RussianSuperGLUE · RussianNLP/MOROCCO | 2020 |
| 20 | Random weighted | 0.48 | Unreasonable Effectiveness of Rule-Based Heuristics in Solving Russian SuperGLUE Tasks | 2021 |