paper-with-me

홈 › Papers

A Systematic Evaluation of Large Language Models of Code

2022-02-26 · Frank F. Xu, Uri Alon, Graham Neubig, Vincent J. Hellendoorn

Large language models (LMs) of code have recently shown tremendous promise in completing code and synthesizing code from natural language descriptions. However, the current state-of-the-art code LMs (e.g., Codex (Chen et al., 2021)) are not publicly available, leaving many questions about their model and data design decisions. We aim to fill in some of these blanks through a systematic evaluation of the largest existing models: Codex, GPT-J, GPT-Neo, GPT-NeoX-20B, and CodeParrot, across various programming languages. Although Codex itself is not open-source, we find that existing open-source models do achieve close results in some programming languages, although targeted mainly for natural language modeling. We further identify an important missing piece in the form of a large open-source model trained exclusively on a multi-lingual corpus of code. We release a new model, PolyCoder, with 2.7B parameters based on the GPT-2 architecture, which was trained on 249GB of code across 12 programming languages on a single machine. In the C programming language, PolyCoder outperforms all models including Codex. Our trained models are open-source and publicly available at https://github.com/VHellendoorn/Code-LMs, which enables future research and application in this area.

📄 PDF Abstract BibTeX arXiv:2202.13169

Code (3)

eleutherai/gpt-neox 공식 구현 pytorch
vhellendoorn/code-lms 공식 구현
scientific-computing-lab-nrcn/tokompiler pytorch

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Weight Decay 설명 없음
Adam 설명 없음

Similar Papers 제목 키워드 기반

InfiBench: Evaluating the Question-Answering Capabilities of Code Large Language Models

2024-03-11 · Linyi Li, Shijie Geng, Zhenwen Li, Yibo He 외

Large Language Models for code (code LLMs) have witnessed tremendous progress in recent years. With the rapid development of code LLMs, many popular evaluation benchmarks, such as HumanEval, DS-1000, and MBPP, have emerg…

Code GenerationHumanEvalmbppQuestion Answering

Enterprise Benchmarks for Large Language Model Evaluation

2024-10-11 · Bing Zhang, Mikio Takeuchi, Ryo Kawahara, Shubhi Asthana 외

The advancement of large language models (LLMs) has led to a greater challenge of having a rigorous and systematic evaluation of complex tasks performed, especially in enterprise applications. Therefore, LLMs need to be …

BenchmarkingLanguage Model EvaluationLanguage ModelingLanguage Modelling+2

How to Trick Your AI TA: A Systematic Study of Academic Jailbreaking in LLM Code Evaluation

2025-12-11 · Devanshu Sahoo, Vasudev Majhi, Arjun Neekhra, Yash Sinha 외 arxiv

The use of Large Language Models (LLMs) as automatic judges for code evaluation is becoming increasingly prevalent in academic environments. But their reliability can be compromised by students who may employ adversarial…

A Taxonomy of Programming Languages for Code Generation

2026-03-31 · Nishat Raihan, Christian Newman, Marcos Zampieri arxiv

The world's 7,000+ languages vary widely in the availability of resources for NLP, motivating efforts to systematically categorize them by their degree of resourcefulness (Joshi et al., 2020). A similar disparity exists …

Code Generation

Gender Disambiguation in Machine Translation: Diagnostic Evaluation in Decoder-Only Architectures

2026-03-18 · Chiara Manna, Hosein Mohebbi, Afra Alishahi, Frédéric Blain 외 arxiv

While Large Language Models achieve state-of-the-art results across a wide range of NLP tasks, they remain prone to systematic biases. Among these, gender bias is particularly salient in MT, due to systematic differences…

Machine Translation