Logical Reasoning 벤치마크
Logical Reasoning on BIG-bench (Logic Grid Puzzle)
Accuracy
- 2021-12-08 — Gopher-280B (few-shot, k=5): Accuracy 35.1
- 2022-03-29 — Chinchilla-70B (few-shot, k=5): Accuracy 44.0
| Rank | Model | Accuracy | Paper | Code | Year |
|---|---|---|---|---|---|
| 1 | Chinchilla-70B (few-shot, k=5) | 44 | Training Compute-Optimal Large Language Models | karpathy/llama2.c · nkluge-correa/teenytinyllama | 2022 |
| 2 | PaLM-540B (few-shot, k=5) | 42.4 | PaLM 2 Technical Report | eternityyw/tram-benchmark | 2023 |
| 3 | PaLM-62B (few-shot, k=5) | 36.5 | PaLM 2 Technical Report | eternityyw/tram-benchmark | 2023 |
| 4 | Gopher-280B (few-shot, k=5) | 35.1 | Scaling Language Models: Methods, Analysis & Insights from Training Gopher | allenai/dolma · rvlopes/gloria · bramiozo/PubScience | 2021 |