Analyzing CodeBERT's Performance on Natural Language Code Search
Large language models such as CodeBERT perform very well on tasks such as natural language code search. We show that this is most likely due to the high token overlap and similarity between the queries and the code in datasets obtained from large codebases, rather than any deeper understanding of the syntax or semantics of the query or code.
Code (0)
등록된 구현이 없습니다.
Tasks
Code SearchMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
CodeBERT: A Pre-Trained Model for Programming and Natural Languages
We present CodeBERT, a bimodal pre-trained model for programming language (PL) and nat-ural language (NL). CodeBERT learns general-purpose representations that support downstream NL-PL applications such as natural langua…
Code Documentation GenerationNaturalness of Attention: Revisiting Attention in Code Language Models
Language models for code such as CodeBERT offer the capability to learn advanced source code representation, but their opacity poses barriers to understanding of captured properties. Recent attention analysis studies pro…
CodeBERTScore: Evaluating Code Generation with Pretrained Models of Code
Since the rise of neural natural-language-to-code models (NL->Code) that can generate long expressions and statements rather than a single next-token, one of the major problems has been reliably evaluating their generate…
Code GenerationOn the Limitations of Embedding Based Methods for Measuring Functional Correctness for Code Generation
The task of code generation from natural language (NL2Code) has become extremely popular, especially with the advent of Large Language Models (LLMs). However, efforts to quantify and track this progress have suffered due…
Code GenerationHumanEvalProbing Semantic Grounding in Language Models of Code with Representational Similarity Analysis
Representational Similarity Analysis is a method from cognitive neuroscience, which helps in comparing representations from two different sources of data. In this paper, we propose using Representational Similarity Analy…