paper-with-me

Papers

Secure Code Generation at Scale with Reflexion

2025-11-05 · Arup Datta, Ahmed Aljohani, Hyunsook Do arxiv

Large language models (LLMs) are now widely used to draft and refactor code, but code that works is not necessarily secure. We evaluate secure code generation using the Instruct Prime, which eliminated compliance-required prompts and cue contamination, and evaluate five instruction-tuned code LLMs using a zero-shot baseline and a three-round reflexion prompting approach. Security is measured using the Insecure Code Detector (ICD), and results are reported by measuring Repair, Regression, and NetGain metrics, considering the programming language and CWE family. Our findings show that insecurity remains common at the first round: roughly 25-33% of programs are insecure at a zero-shot baseline (t0 ). Weak cryptography/config-dependent bugs are the hardest to avoid while templated ones like XSS, code injection, and hard-coded secrets are handled more reliably. Python yields the highest secure rates; C and C# are the lowest, with Java, JS, PHP, and C++ in the middle. Reflexion prompting improves security for all models, improving average accuracy from 70.74% at t0 to 79.43% at t3 , with the largest gains in the first round followed by diminishing returns. The trends with Repair, Regression, and NetGain metrics show that applying one to two rounds produces most of the benefits. A replication package is available at https://doi.org/10.5281/zenodo.17065846.

📄 PDF Abstract BibTeX arXiv:2511.03898

Code (0)

등록된 구현이 없습니다.

Tasks

Code Generation

Similar Papers 제목 키워드 기반

Narrative-Centered Emotional Reflection: Scaffolding Autonomous Emotional Literacy with AI

2025-04-29 · Shou-Tzu Han

Reflexion is an AI-powered platform designed to enable structured emotional self-reflection at scale. By integrating real-time emotion detection, layered reflective prompting, and metaphorical storytelling generation, Re…

Providing Self-Aware Systems with Reflexivity

2017-07-27 · Alessandro Valitutti, Giuseppe Trautteur

We propose a new type of self-aware systems inspired by ideas from higher-order theories of consciousness. First, we discussed the crucial distinction between introspection and reflexion. Then, we focus on computational …

DocSync: Agentic Documentation Maintenance via Critic-Guided Reflexion

2026-05-04 · Sidhesh Badrinarayan, Adithya Parthasarathy arxiv

Software documentation frequently drifts from executable logic as codebases evolve, creating technical debt that degrades maintainability and causes downstream API misuse. While static analysis tools can detect the absen…

Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks

2025-07-31 · Jianghui Wang, Vinay Joshi, Saptarshi Majumder, Xu Chao 외 arxiv

The demand for AI-generated GPU kernels is rapidly growing, influenced by the need for scalable, hardware-optimized solutions in both industry and academia. As deep learning workloads grow in complexity and diversity, it…

Code Generation

SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents

2024-03-23 · Feng Lin, Dong Jae Kim, Tse-Husn, Chen

Software process models are essential to facilitate collaboration and communication among software teams to solve complex development tasks. Inspired by these software engineering practices, we present FlowGen - a code g…

Code GenerationHumanEvalLanguage ModelingLanguage Modelling+2