paper-with-me

Papers

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data

2026-05-23 · Junlan Feng, Fanyu Meng, Chong Long, Pengyu Cong, Duqing Wang, Yan Zheng, Yuyao Zhang, Xuanchang Gao, Ye Yuan, Yunfei Ma, Zhijie Ren, Fan Yang, Na Wu, Di Jin, Chao Deng arxiv

We introduce JT-Safe-V2, a large language model designed to advance the safety and trustworthiness of foundation models, extending our previous JT-Safe model toward a more comprehensive safety-by-design paradigm. JT-Safe-V2 emphasizes the joint optimization of general intelligence and safety-by-design through several key innovations: enriching pre-training data with contextual world knowledge, high-certainty pre-training procedures, and safety strengthening post-training mechanisms for enterprise-oriented agentic capabilities. Building on these safety-enhanced foundation models, we propose Safe-MoMA (Safe Mixture of Models and Agents), a framework that enables traceable and efficient inference through the orchestrated deployment of multiple models and agents. Extensive evaluations demonstrate that JT-Safe-V2 achieves state-of-the-art performance across both general intelligence and safety benchmarks. Moreover, Safe-MoMA reduces inference costs by more than 30\% compared to using the largest standalone model baseline while maintaining comparable performance. To facilitate future research on safety-by-design foundation models, we publicly release the post-trained JT-Safe-V2-35B model checkpoint.

📄 PDF Abstract BibTeX arXiv:2605.24414

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Modular Safety Guardrails Are Necessary for Foundation-Model-Enabled Robots in the Real World

2026-02-03 · Joonkyung Kim, Wenxi Chen, Davood Soleymanzadeh, Yi Ding 외 arxiv

The integration of foundation models (FMs) into robotics has accelerated real-world deployment, while introducing new safety challenges arising from open-ended semantic reasoning and embodied physical action. These chall…

Understanding Real-World Traffic Safety through RoadSafe365 Benchmark

2026-02-06 · Xinyu Liu, Darryl C. Jacob, Yuxin Liu, Xinsong Du 외 arxiv

Although recent traffic benchmarks have advanced multimodal data analysis, they generally lack systematic evaluation aligned with official safety standards. To fill this gap, we introduce RoadSafe365, a large-scale visio…

Engineering Safety in Machine Learning

2016-01-16 · Kush R. Varshney

Machine learning algorithms are increasingly influencing our decisions and interacting with us in all parts of our daily lives. Therefore, just like for power plants, highways, and myriad other engineered sociotechnical …

BIG-bench Machine Learning

TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts

2026-04-17 · Hua-Rong Chu, Kuan-Chun Wang, Yao-Te Huang arxiv

Safety guardrails have become an active area of research in AI safety, aimed at ensuring the appropriate behavior of large language models (LLMs). However, existing research lacks consideration of nuances across linguist…

Clear, Compelling Arguments: Rethinking the Foundations of Frontier AI Safety Cases

2026-03-08 · Shaun Feakins, Ibrahim Habli, Phillip Morgan arxiv

This paper contributes to the nascent debate around safety cases for frontier AI systems. Safety cases are structured, defensible arguments that a system is acceptably safe to deploy in a given context. Historically, the…