paper-with-me

Code Generation

31개 벤치마크 · 논문 3,062편 · 이 태스크의 논문 보기 →

Benchmarks

Terminal-Bench Hard

결과 432개

LiveCodeBench

결과 350개

Terminal-Bench v2.1

결과 241개

SciCode

결과 173개

MBPP

결과 101개

APPS

결과 19개

CoNaLa

결과 14개

HumanEval

결과 12개

CodeContests

결과 11개

Django

결과 11개

WikiSQL

결과 10개

RES-Q

결과 9개

PECC

결과 8개

WebApp1K-React

결과 8개

CoNaLa-Ext

결과 6개

WebApp1k-Duo-React

결과 6개

DSEval-LeetCode

결과 5개

Turbulence

결과 5개

VerilogEval

결과 5개

Livecodebench

결과 3개

Shellcode_IA32

결과 3개

TACO-BAAI

결과 3개

BigCodeBench-Complete

결과 2개

BigCodeBench-Instruct

결과 2개

CONCODE

결과 2개

HumanEval-ET

결과 2개

MBPP-ET

결과 2개

Android Repos

결과 1개

FloCo

결과 1개

Most implemented

GPT-4 Technical Report

2023-03-15 · 구현 11개

Papers

Retrofitting Code Using LLMs to Support Exceptional Behavior

2026-09-09 · Linghan Zhong, Jiyang Zhang, Jayanth Srinivasa, Junyi Jessy Li 외 arxiv

Exception Related Code (ERC), which includes throw statements, conditions (if statements) that guard those throw statements, and try/catch blocks, is an essential component of software systems, allowing developers to det…

Code Generation

Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?

2026-09-09 · Leilei Ding, Shumin Wang, Yuting Huang, Fanqi Wan 외 arxiv

Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the prospect that they could assist in developing and optimizing the very infrastructure that powers them. …

Code Generation

RISE: Recursive Improvement via Self-Extrapolating Policy Distillation

2026-09-04 · Yang Li, Semih Yavuz, Shafiq Joty arxiv

On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its effectiveness is bottlenecked by teacher quality: external teachers suffer from distribution mismatch, while se…

Mathematical ReasoningCode Generation

Substrate-Aware AI Agents: Execution Context as a First-Class Input

2026-09-04 · Manu Agrawal arxiv

Autonomous AI agents increasingly select actions in environments whose memory, execution-time, runtime, compute, and operational constraints determine what counts as a suitable plan. We call the absence of this execution…

Code Generation

AutoLR: Automating the Path from Research to Launch Review in Industrial Recommender Systems

2026-09-04 · Qi Zhang, Yanlin Chen, Wenchao Xiao arxiv

Improving an industrial recommender is an iterative research-and-engineering process rather than a direct path from idea to deployment. In \textbf{DASHEN, NetEase's gaming-community app}, algorithm engineers typically id…

Code Generation

Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle

2026-09-04 · Happy Bhati arxiv

AI coding systems are moving from autocomplete and chat toward agents that can inspect repositories, edit multiple files, run tools, write tests, open pull requests, and work for long periods with limited supervision. Th…

Code Generation

전체 3,062편 보기 →