paper-with-me

Code Generation 벤치마크

Code Generation on Turbulence

5개 결과 · ⬇ CSV · JSON

CorrSc

0.063 0.2592 0.4555 0.6517 0.848 2023-12 2026-09 GPT-4 — 0.848 (2023-12-22) GPT-3.5-Turbo — 0.617 (2023-12-22) CodeLlama:13B-4bit-quantised — 0.327 (2023-12-22) CodeLlama:7B-4bit-quantised — 0.289 (2023-12-22) Command — 0.063 (2023-12-22) GPT-4 — 0.848 (2023-12-22)
RankModel CorrSc PaperCodeYear
1 GPT-4 0.848 Turbulence: Systematically and Automatically Testing Instruction-Tuned Large Language Models for Code shahinhonarvar/turbulence-benchmark 2023
2 GPT-3.5-Turbo 0.617 Turbulence: Systematically and Automatically Testing Instruction-Tuned Large Language Models for Code shahinhonarvar/turbulence-benchmark 2023
3 CodeLlama:13B-4bit-quantised 0.327 Turbulence: Systematically and Automatically Testing Instruction-Tuned Large Language Models for Code shahinhonarvar/turbulence-benchmark 2023
4 CodeLlama:7B-4bit-quantised 0.289 Turbulence: Systematically and Automatically Testing Instruction-Tuned Large Language Models for Code shahinhonarvar/turbulence-benchmark 2023
5 Command 0.063 Turbulence: Systematically and Automatically Testing Instruction-Tuned Large Language Models for Code shahinhonarvar/turbulence-benchmark 2023
1–5 / 5 페이지당 10 20 50 100