paper-with-me

Common Sense Reasoning 벤치마크

Common Sense Reasoning on PARus

22개 결과 · ⬇ CSV · JSON

Accuracy

0.478 0.604 0.73 0.856 0.982 2020-10 2026-09 MT5 Large — 0.504 (2020-10-22) Human Benchmark — 0.982 (2020-10-29) Baseline TF-IDF1.1 — 0.486 (2020-10-29) majority_class — 0.498 (2021-05-03) Random weighted — 0.48 (2021-05-03) heuristic majority — 0.478 (2021-05-03) MT5 Large — 0.504 (2020-10-22) Human Benchmark — 0.982 (2020-10-29)
RankModel Accuracy PaperCodeYear
1 Human Benchmark 0.982 RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark RussianNLP/RussianSuperGLUE · RussianNLP/MOROCCO 2020
2 Golden Transformer 0.908
3 YaLM 1.0B few-shot 0.766
4 RuGPT3XL few-shot 0.676
5 ruT5-large-finetune 0.66
6 RuGPT3Medium 0.598
7 RuGPT3Large 0.584
8 RuBERT plain 0.574
9 RuGPT3Small 0.562
10 ruT5-base-finetune 0.554
11 Multilingual Bert 0.528
12 ruRoberta-large finetune 0.508
12 RuBERT conversational 0.508
14 MT5 Large 0.504 mT5: A massively multilingual pre-trained text-to-text transformer huggingface/transformers · google-research/multilingual-t5 · google-research/byt5 · +5 2020
15 SBERT_Large_mt_ru_finetuning 0.498
15 SBERT_Large 0.498
15 majority_class 0.498 Unreasonable Effectiveness of Rule-Based Heuristics in Solving Russian SuperGLUE Tasks 2021
18 ruBert-large finetune 0.492
19 Baseline TF-IDF1.1 0.486 RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark RussianNLP/RussianSuperGLUE · RussianNLP/MOROCCO 2020
20 Random weighted 0.48 Unreasonable Effectiveness of Rule-Based Heuristics in Solving Russian SuperGLUE Tasks 2021
1–20 / 22 다음 → 페이지당 10 20 50 100