paper-with-me

홈 › Papers

Risk Assessment Framework for Code LLMs via Leveraging Internal States

2025-04-20 · Yuheng Huang, Lei Ma, Keizaburo Nishikino, Takumi Akazaki

The pre-training paradigm plays a key role in the success of Large Language Models (LLMs), which have been recognized as one of the most significant advancements of AI recently. Building on these breakthroughs, code LLMs with advanced coding capabilities bring huge impacts on software engineering, showing the tendency to become an essential part of developers' daily routines. However, the current code LLMs still face serious challenges related to trustworthiness, as they can generate incorrect, insecure, or unreliable code. Recent exploratory studies find that it can be promising to detect such risky outputs by analyzing LLMs' internal states, akin to how the human brain unconsciously recognizes its own mistakes. Yet, most of these approaches are limited to narrow sub-domains of LLM operations and fall short of achieving industry-level scalability and practicability. To address these challenges, in this paper, we propose PtTrust, a two-stage risk assessment framework for code LLM based on internal state pre-training, designed to integrate seamlessly with the existing infrastructure of software companies. The core idea is that the risk assessment framework could also undergo a pre-training process similar to LLMs. Specifically, PtTrust first performs unsupervised pre-training on large-scale unlabeled source code to learn general representations of LLM states. Then, it uses a small, labeled dataset to train a risk predictor. We demonstrate the effectiveness of PtTrust through fine-grained, code line-level risk assessment and demonstrate that it generalizes across tasks and different programming languages. Further experiments also reveal that PtTrust provides highly intuitive and interpretable features, fostering greater user trust. We believe PtTrust makes a promising step toward scalable and trustworthy assurance for code LLMs.

📄 PDF Abstract BibTeX arXiv:2504.14640

Code (0)

등록된 구현이 없습니다.

Tasks

Unsupervised Pre-training

Similar Papers 제목 키워드 기반

Generative LLM Powered Conversational AI Application for Personalized Risk Assessment: A Case Study in COVID-19

2024-09-23 · Mohammad Amin Roshani, Xiangyu Zhou, Yao Qiang, Srinivasan Suresh 외

Large language models (LLMs) have shown remarkable capabilities in various natural language tasks and are increasingly being applied in healthcare domains. This work demonstrates a new LLM-powered disease risk assessment…

Feature Importance

Leveraging Large Language Models for Risk Assessment in Hyperconnected Logistic Hub Network Deployment

2025-03-27 · Yinzhu Quan, Yujia Xu, Guanlin Chen, Frederick Benaben 외

The growing emphasis on energy efficiency and environmental sustainability in global supply chains introduces new challenges in the deployment of hyperconnected logistic hub networks. In current volatile, uncertain, comp…

Decision MakingLarge Language Model

Leveraging Large Language Models for Cybersecurity Risk Assessment -- A Case from Forestry Cyber-Physical Systems

2025-10-07 · Fikret Mert Gultekin, Oscar Lilja, Ranim Khojah, Rebekka Wohlrab 외 arxiv

In safety-critical software systems, cybersecurity activities become essential, with risk assessment being one of the most critical. In many software teams, cybersecurity experts are either entirely absent or represented…

SandboxEval: Towards Securing Test Environment for Untrusted Code

2025-03-27 · Rafiqul Rabin, Jesse Hostetler, Sean McGregor, Brett Weir 외

While large language models (LLMs) are powerful assistants in programming tasks, they may also produce malicious code. Testing LLM-generated code therefore poses significant risks to assessment infrastructure tasked with…

WildfireVLM: AI-powered Analysis for Early Wildfire Detection and Risk Assessment Using Satellite Imagery

2026-02-09 · Aydin Ayanzadeh, Prakhar Dixit, Sadia Kamal, Milton Halem arxiv

Wildfires are a growing threat to ecosystems, human lives, and infrastructure, with their frequency and intensity rising due to climate change and human activities. Early detection is critical, yet satellite-based monito…