paper-with-me

Multi-task Language Understanding 벤치마크

Multi-task Language Understanding on BBH-alg

14개 결과 · ⬇ CSV · JSON

Average (%)

38.3 47.2 56.1 65 73.9 2021-07 2026-09 code-davinci-002 175B (CoT) — 73.9 (2021-07-07) code-davinci-002 175B (CoT) — 73.9 (2021-07-07) Flan-PaLM 540B (3-shot, fine-tuned, CoT + SC) — 66.5 (2022-10-20) PaLM 540B (CoT + self-consistency) — 62.2 (2022-10-20) Flan-PaLM 540B (3-shot, fine-tuned, CoT) — 61.3 (2022-10-20) PaLM 540B (CoT) — 57.6 (2022-10-20) Flan-PaLM 540B (3-shot, fine-tuned) — 48.2 (2022-10-20) PaLM 540B — 38.3 (2022-10-20) Flan-PaLM 540B (3-shot, fine-tuned, CoT + SC) — 66.5 (2022-10-20) PaLM 540B (CoT + self-consistency) — 62.2 (2022-10-20) Flan-PaLM 540B (3-shot, fine-tuned, CoT) — 61.3 (2022-10-20) PaLM 540B (CoT) — 57.6 (2022-10-20) Flan-PaLM 540B (3-shot, fine-tuned) — 48.2 (2022-10-20) PaLM 540B — 38.3 (2022-10-20) code-davinci-002 175B (CoT) — 73.9 (2021-07-07)
RankModel Average (%) PaperCodeYear
1 code-davinci-002 175B (CoT) 73.9 Evaluating Large Language Models Trained on Code THUDM/CodeGeeX · ncoop57/gpt-code-clippy · codedotal/gpt-code-clippy · +10 2021
2 Flan-PaLM 540B (3-shot, fine-tuned, CoT + SC) 66.5 Scaling Instruction-Finetuned Language Models google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite · +6 2022
3 PaLM 540B (CoT + self-consistency) 62.2 Scaling Instruction-Finetuned Language Models google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite · +6 2022
4 Flan-PaLM 540B (3-shot, fine-tuned, CoT) 61.3 Scaling Instruction-Finetuned Language Models google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite · +6 2022
5 PaLM 540B (CoT) 57.6 Scaling Instruction-Finetuned Language Models google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite · +6 2022
6 Flan-PaLM 540B (3-shot, fine-tuned) 48.2 Scaling Instruction-Finetuned Language Models google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite · +6 2022
7 PaLM 540B 38.3 Scaling Instruction-Finetuned Language Models google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite · +6 2022
8 code-davinci-002 175B (CoT) 73.9 Evaluating Large Language Models Trained on Code THUDM/CodeGeeX · ncoop57/gpt-code-clippy · codedotal/gpt-code-clippy · +10 2021
9 Flan-PaLM 540B (3-shot, fine-tuned, CoT + SC) 66.5 Scaling Instruction-Finetuned Language Models google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite · +6 2022
10 PaLM 540B (CoT + self-consistency) 62.2 Scaling Instruction-Finetuned Language Models google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite · +6 2022
11 Flan-PaLM 540B (3-shot, fine-tuned, CoT) 61.3 Scaling Instruction-Finetuned Language Models google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite · +6 2022
12 PaLM 540B (CoT) 57.6 Scaling Instruction-Finetuned Language Models google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite · +6 2022
13 Flan-PaLM 540B (3-shot, fine-tuned) 48.2 Scaling Instruction-Finetuned Language Models google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite · +6 2022
14 PaLM 540B 38.3 Scaling Instruction-Finetuned Language Models google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite · +6 2022
1–14 / 14 페이지당 10 20 50 100