| Rank | Model |
Average (%) |
Paper | Code | Year |
| 1 |
code-davinci-002 175B (CoT) |
73.9 |
Evaluating Large Language Models Trained on Code
|
THUDM/CodeGeeX · ncoop57/gpt-code-clippy · codedotal/gpt-code-clippy
· +10 |
2021 |
| 2 |
Flan-PaLM 540B (3-shot, fine-tuned, CoT + SC) |
66.5 |
Scaling Instruction-Finetuned Language Models
|
google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite
· +6 |
2022 |
| 3 |
PaLM 540B (CoT + self-consistency) |
62.2 |
Scaling Instruction-Finetuned Language Models
|
google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite
· +6 |
2022 |
| 4 |
Flan-PaLM 540B (3-shot, fine-tuned, CoT) |
61.3 |
Scaling Instruction-Finetuned Language Models
|
google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite
· +6 |
2022 |
| 5 |
PaLM 540B (CoT) |
57.6 |
Scaling Instruction-Finetuned Language Models
|
google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite
· +6 |
2022 |
| 6 |
Flan-PaLM 540B (3-shot, fine-tuned) |
48.2 |
Scaling Instruction-Finetuned Language Models
|
google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite
· +6 |
2022 |
| 7 |
PaLM 540B |
38.3 |
Scaling Instruction-Finetuned Language Models
|
google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite
· +6 |
2022 |
| 8 |
code-davinci-002 175B (CoT) |
73.9 |
Evaluating Large Language Models Trained on Code
|
THUDM/CodeGeeX · ncoop57/gpt-code-clippy · codedotal/gpt-code-clippy
· +10 |
2021 |
| 9 |
Flan-PaLM 540B (3-shot, fine-tuned, CoT + SC) |
66.5 |
Scaling Instruction-Finetuned Language Models
|
google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite
· +6 |
2022 |
| 10 |
PaLM 540B (CoT + self-consistency) |
62.2 |
Scaling Instruction-Finetuned Language Models
|
google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite
· +6 |
2022 |
| 11 |
Flan-PaLM 540B (3-shot, fine-tuned, CoT) |
61.3 |
Scaling Instruction-Finetuned Language Models
|
google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite
· +6 |
2022 |
| 12 |
PaLM 540B (CoT) |
57.6 |
Scaling Instruction-Finetuned Language Models
|
google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite
· +6 |
2022 |
| 13 |
Flan-PaLM 540B (3-shot, fine-tuned) |
48.2 |
Scaling Instruction-Finetuned Language Models
|
google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite
· +6 |
2022 |
| 14 |
PaLM 540B |
38.3 |
Scaling Instruction-Finetuned Language Models
|
google-research/flan · declare-lab/flan-alpaca · formulamonks/llm-benchmarker-suite
· +6 |
2022 |