paper-with-me

Logical Reasoning 벤치마크

Logical Reasoning on BIG-bench (Temporal Sequences)

9개 결과 · ⬇ CSV · JSON

Accuracy

19 39.25 59.5 79.75 100 2021-12 2026-09 Gopher-280B (few-shot, k=5) — 19.0 (2021-12-08) Chinchilla-70B (few-shot, k=5) — 32.0 (2022-03-29) PaLM 540B (few-shot, k=3) — 39.6 (2023-03-30) BLOOM 176B (few-shot, k=3) — 36.8 (2023-03-30) Bloomberg GPT (few-shot, k=3) — 29.2 (2023-03-30) OPT 66B (few-shot, k=3) — 23.6 (2023-03-30) GPT-NeoX (few-shot, k=3) — 21.2 (2023-03-30) PaLM 2 (few-shot, k=3, CoT) — 100.0 (2023-05-17) PaLM 2 (few-shot, k=3, Direct) — 96.4 (2023-05-17) Gopher-280B (few-shot, k=5) — 19.0 (2021-12-08) Chinchilla-70B (few-shot, k=5) — 32.0 (2022-03-29) PaLM 540B (few-shot, k=3) — 39.6 (2023-03-30) PaLM 2 (few-shot, k=3, CoT) — 100.0 (2023-05-17)
RankModel Accuracy PaperCodeYear
1 PaLM 2 (few-shot, k=3, CoT) 100 PaLM 2 Technical Report eternityyw/tram-benchmark 2023
2 PaLM 2 (few-shot, k=3, Direct) 96.4 PaLM 2 Technical Report eternityyw/tram-benchmark 2023
3 PaLM 540B (few-shot, k=3) 39.6 BloombergGPT: A Large Language Model for Finance yangletliu/finlora · open-finance-lab/finlora 2023
4 BLOOM 176B (few-shot, k=3) 36.8 BloombergGPT: A Large Language Model for Finance yangletliu/finlora · open-finance-lab/finlora 2023
5 Chinchilla-70B (few-shot, k=5) 32.0 Training Compute-Optimal Large Language Models karpathy/llama2.c · nkluge-correa/teenytinyllama 2022
6 Bloomberg GPT (few-shot, k=3) 29.2 BloombergGPT: A Large Language Model for Finance yangletliu/finlora · open-finance-lab/finlora 2023
7 OPT 66B (few-shot, k=3) 23.6 BloombergGPT: A Large Language Model for Finance yangletliu/finlora · open-finance-lab/finlora 2023
8 GPT-NeoX (few-shot, k=3) 21.2 BloombergGPT: A Large Language Model for Finance yangletliu/finlora · open-finance-lab/finlora 2023
9 Gopher-280B (few-shot, k=5) 19.0 Scaling Language Models: Methods, Analysis & Insights from Training Gopher allenai/dolma · rvlopes/gloria · bramiozo/PubScience 2021
1–9 / 9 페이지당 10 20 50 100