paper-with-me

홈 › Papers

EARL: Entropy-Aware RL Alignment of LLMs for Reliable RTL Code Generation

2025-11-15 · Jiahe Shi, Zhengqi Gao, Ching-Yun Ko, Duane Boning arxiv

Recent advances in large language models (LLMs) have demonstrated significant potential in hardware design automation, particularly in using natural language to synthesize Register-Transfer Level (RTL) code. Despite this progress, a gap remains between model capability and the demands of real-world RTL design, including syntax errors, functional hallucinations, and weak alignment to designer intent. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising approach to bridge this gap, as hardware provides executable and formally checkable signals that can be used to further align model outputs with design intent. However, in long, structured RTL code sequences, not all tokens contribute equally to functional correctness, and naïvely spreading gradients across all tokens dilutes learning signals. A key insight from our entropy analysis in RTL generation is that only a small fraction of tokens (e.g., always, if, assign, posedge) exhibit high uncertainty and largely influence control flow and module structure. To address these challenges, we present EARL, an Entropy-Aware Reinforcement Learning framework for Verilog generation. EARL performs policy optimization using verifiable reward signals and introduces entropy-guided selective updates that gate policy gradients to high-entropy tokens. This approach preserves training stability and concentrates gradient updates on functionally important regions of code. Our experiments on VerilogEval and RTLLM show that EARL improves functional pass rates over prior LLM baselines by up to 14.7%, while reducing unnecessary updates and improving training stability. These results indicate that focusing RL on critical, high-uncertainty tokens enables more reliable and targeted policy improvement for structured RTL code generation.

📄 PDF Abstract BibTeX arXiv:2511.12033

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningCode Generation

Similar Papers 제목 키워드 기반

INTRYGUE: Induction-Aware Entropy Gating for Reliable RAG Uncertainty Estimation

2026-03-23 · Alexandra Bazarova, Andrei Volodichev, Daria Kotova, Alexey Zaytsev arxiv

While retrieval-augmented generation (RAG) enhances LLM performance, it does not eliminate hallucinations, making accurate detection essential. Uncertainty-based methods are attractive for this purpose because they can b…

Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding

2026-06-20 · Xuanming Zhang, Sining Zhoubian, Yuxuan Chen, Tianyi Tang 외 arxiv

Autoregressive generation in large language models (LLMs) conventionally decodes from the final layer, assuming that deeper representations yield more reliable next-token predictions. We revisit this assumption by reveal…

Thinking, Faithful and Stable: Mitigating Hallucinations in LLMs

2025-11-19 · Chelsea Zou, Yiheng Yao, Basant Khalil arxiv

This project develops a self correcting framework for large language models (LLMs) that detects and mitigates hallucinations during multi-step reasoning. Rather than relying solely on final answer correctness, our approa…

Reinforcement Learning

Uncertainty Aware Learning for Language Model Alignment

2024-06-07 · Yikun Wang, Rui Zheng, Liang Ding, Qi Zhang 외

As instruction-tuned large language models (LLMs) evolve, aligning pretrained foundation models presents increasing challenges. Existing alignment strategies, which typically leverage diverse and high-quality data source…

GSM8KLanguage ModelingLanguage Modellingmodel

Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment

2025-11-17 · Jea Kwon, Luiz Felipe Vecchietti, Sungwon Park, Meeyoung Cha arxiv

Humans display significant uncertainty when confronted with moral dilemmas, yet the extent of such uncertainty in machines and AI agents remains underexplored. Recent studies have confirmed the overly confident tendencie…