paper-with-me

홈 › Papers

A Taxonomy of Programming Languages for Code Generation

2026-03-31 · Nishat Raihan, Christian Newman, Marcos Zampieri arxiv

The world's 7,000+ languages vary widely in the availability of resources for NLP, motivating efforts to systematically categorize them by their degree of resourcefulness (Joshi et al., 2020). A similar disparity exists among programming languages (PLs); however, no resource-tier taxonomy has been established for code. As large language models (LLMs) grow increasingly capable of generating code, such a taxonomy becomes essential. To fill this gap, we present the first reproducible PL resource classification, grouping 646 languages into four tiers. We show that only 1.9% of languages (Tier 3, High) account for 74.6% of all tokens in seven major corpora, while 71.7% of languages (Tier 0, Scarce) contribute just 1.0%. Statistical analyses of within-tier inequality, dispersion, and distributional skew confirm that this imbalance is both extreme and systematic. Our results provide a principled framework for dataset curation and tier-aware evaluation of multilingual LLMs.

📄 PDF Abstract BibTeX arXiv:2604.00239

Code (0)

등록된 구현이 없습니다.

Tasks

Code Generation

Similar Papers 제목 키워드 기반

Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages

2024-09-01 · William Zhang, Maria Leon, Ryan Xu, Adrian Cardenas 외

Node-based programming languages are increasingly popular in media arts coding domains. These languages are designed to be accessible to users with limited coding experience, allowing them to achieve creative output with…

BenchmarkingCode Generation

GenCodeSearchNet: A Benchmark Test Suite for Evaluating Generalization in Programming Language Understanding

2023-11-16 · Andor Diera, Abdelhalim Dahou, Lukas Galke, Fabian Karl 외

Language models can serve as a valuable tool for software developers to increase productivity. Large generative models can be used for code generation and code completion, while smaller encoder-only models are capable of…

Code CompletionCode GenerationCode SearchDiversity

A Survey of Machine Learning for Big Code and Naturalness

2017-09-18 · Miltiadis Allamanis, Earl T. Barr, Premkumar Devanbu, Charles Sutton

Research at the intersection of machine learning, programming languages, and software engineering has recently taken important steps in proposing learnable probabilistic models of source code that exploit code's abundanc…

BIG-bench Machine LearningNavigateSurvey

MERA Code: A Unified Framework for Evaluating Code Generation Across Tasks

2025-07-16 · Artem Chervyakov, Alexander Kharitonov, Pavel Zadorozhny, Adamenko Pavel 외

Advancements in LLMs have enhanced task automation in software engineering; however, current evaluations primarily focus on natural language tasks, overlooking code quality. Most benchmarks prioritize high-level reasonin…

Code Generation

MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation

2022-08-17 · Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen 외

Large language models have demonstrated the ability to generate both natural language and programming language text. Such models open up the possibility of multi-language code generation: could code generation models gen…

BenchmarkingCode GenerationHumanEvalmbpp