paper-with-me

Logical Reasoning 벤치마크

Logical Reasoning on BIG-bench (Logic Grid Puzzle)

4개 결과 · ⬇ CSV · JSON

Accuracy

35.1 37.33 39.55 41.77 44 2021-12 2026-09 Gopher-280B (few-shot, k=5) — 35.1 (2021-12-08) Chinchilla-70B (few-shot, k=5) — 44.0 (2022-03-29) PaLM-540B (few-shot, k=5) — 42.4 (2023-05-17) PaLM-62B (few-shot, k=5) — 36.5 (2023-05-17) Gopher-280B (few-shot, k=5) — 35.1 (2021-12-08) Chinchilla-70B (few-shot, k=5) — 44.0 (2022-03-29)
RankModel Accuracy PaperCodeYear
1 Chinchilla-70B (few-shot, k=5) 44 Training Compute-Optimal Large Language Models karpathy/llama2.c · nkluge-correa/teenytinyllama 2022
2 PaLM-540B (few-shot, k=5) 42.4 PaLM 2 Technical Report eternityyw/tram-benchmark 2023
3 PaLM-62B (few-shot, k=5) 36.5 PaLM 2 Technical Report eternityyw/tram-benchmark 2023
4 Gopher-280B (few-shot, k=5) 35.1 Scaling Language Models: Methods, Analysis & Insights from Training Gopher allenai/dolma · rvlopes/gloria · bramiozo/PubScience 2021
1–4 / 4 페이지당 10 20 50 100