paper-with-me

홈 › Papers

Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations

2026-03-18 · Haozheng Luo, Yimin Wang, Jiahao Yu, Binghui Wang, Yan Chen arxiv

We propose CRAFT, a red-teaming alignment framework that leverages model reasoning capabilities and hidden representations to improve robustness against jailbreak attacks. Unlike prior defenses that operate primarily at the output level, CRAFT aligns large reasoning models to generate safety-aware reasoning traces by explicitly optimizing objectives defined over the hidden state space. Methodologically, CRAFT integrates contrastive representation learning with reinforcement learning to separate safe and unsafe reasoning trajectories, yielding a latent-space geometry that supports robust, reasoning-level safety alignment. Theoretically, we show that incorporating latent-textual consistency into GRPO eliminates superficially aligned policies by ruling them out as local optima. Empirically, we evaluate CRAFT on multiple safety benchmarks using two strong reasoning models, Qwen3-4B-Thinking and R1-Distill-Llama-8B, where it consistently outperforms state-of-the-art defenses such as IPO and SafeKey. Notably, CRAFT delivers an average 79.0% improvement in reasoning safety and 87.7% improvement in final-response safety over the base models, demonstrating the effectiveness of hidden-space reasoning alignment.

📄 PDF Abstract BibTeX arXiv:2603.17305

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningReinforcement Learning

Similar Papers 제목 키워드 기반

CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization

2026-05-09 · Ziyang Ding, Linjian Meng, Yiming Wu, Yuhan Li 외 arxiv

Due to the potential for exploratory reasoning of Latent Visual Reasoning, recent works tend to enable MLLMs (Multimodal Large Language Models) to perform visual reasoning by propagating continuous hidden states instead …

Reinforcement LearningVisual Reasoning

MechELK: A Mechanistic Interpretability Framework for Eliciting Latent Knowledge in Large Language Models

2026-04-07 · Ji-jun Park, Soo-joon Choi, Jiwon Jeong, Taeyang Yoon 외 arxiv

Large language models (LLMs) frequently encode factual and reasoning knowledge in their internal representations that is not faithfully reflected in their surface-level outputs -- a phenomenon known as \emph{latent knowl…

Expediting and Elevating Large Language Model Reasoning via Hidden Chain-of-Thought Decoding

2024-09-13 · Tianqiao Liu, Zui Chen, Zitao Liu, Mi Tian 외

Large language models (LLMs) have demonstrated remarkable capabilities in tasks requiring reasoning and multi-step problem-solving through the use of chain-of-thought (CoT) prompting. However, generating the full CoT pro…

Contrastive LearningLanguage ModelingLanguage ModellingLarge Language Model+3

Latent-Aligned Reasoning for Multimodal Recommendation

2026-09-04 · Jiarui Jin, Anyang Ji arxiv

Multimodal Vision-Language Models (VLMs) have demonstrated remarkable capabilities in cross-modal understanding, yet a fundamental challenge persists when applying them to recommendation: as representations propagate thr…

Multimodal RecommendationContrastive Learning

CRE-T1 Preview Technical Report: Beyond Contrastive Learning for Reasoning-Intensive Retrieval

2026-03-18 · Guangzhi Wang, Yinghao Jiao, Zhi Liu arxiv

The central challenge of reasoning-intensive retrieval lies in identifying implicitreasoning relationships between queries and documents, rather than superficial se-mantic or lexical similarity. The contrastive learning …

Reinforcement LearningContrastive Learning