Code Generation 벤치마크
Code Generation on MBPP
Accuracy
- 2022-04-05 — PaLM Coder 540B: Accuracy 47.0
- 2022-04-25 — code-davinci-001 175B + MBR-Exec: Accuracy 58.2
- 2022-07-21 — code-davinci-002 175B + CodeT: Accuracy 67.7
- 2023-02-16 — code-davinci-002 175B + LEVER: Accuracy 68.9
- 2023-04-11 — GPT-4 (Self-Debugging with unit tests + trace): Accuracy 80.2
- 2023-07-24 — GPT-4 (ChatGPT Plus): Accuracy 87.5
- 2023-12-20 — GPT-4 + AgentCoder: Accuracy 91.8
- 2024-05-18 — o1-mini + MapCoder (Hamming.ai): Accuracy 93.2
- 2025-01-20 — QualityFlow (Sonnet-3.5): Accuracy 94.2
- 2025-06-12 — EG-CFG (DeepSeek-V3-0324): Accuracy 96.6