paper-with-me

홈 › Papers

TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis

2025-05-30 · Xiaorui Wu, Xiaofeng Mao, Fei Li, Xin Zhang, Xuanhong Li, Chong Teng, Donghong Ji, Zhuang Li

Large Language Models (LLMs) excel in various natural language processing tasks but remain vulnerable to generating harmful content or being exploited for malicious purposes. Although safety alignment datasets have been introduced to mitigate such risks through supervised fine-tuning (SFT), these datasets often lack comprehensive risk coverage. Most existing datasets focus primarily on lexical diversity while neglecting other critical dimensions. To address this limitation, we propose a novel analysis framework to systematically measure the risk coverage of alignment datasets across three essential dimensions: Lexical Diversity, Malicious Intent, and Jailbreak Tactics. We further introduce TRIDENT, an automated pipeline that leverages persona-based, zero-shot LLM generation to produce diverse and comprehensive instructions spanning these dimensions. Each harmful instruction is paired with an ethically aligned response, resulting in two datasets: TRIDENT-Core, comprising 26,311 examples, and TRIDENT-Edge, with 18,773 examples. Fine-tuning Llama 3.1-8B on TRIDENT-Edge demonstrates substantial improvements, achieving an average 14.29% reduction in Harm Score, and a 20% decrease in Attack Success Rate compared to the best-performing baseline model fine-tuned on the WildBreak dataset.

📄 PDF Abstract BibTeX arXiv:2505.24672

Code (1)

fisht0ucher/trident 공식 구현

Tasks

DiversityLanguage ModelingLanguage ModellingLarge Language ModelRed TeamingSafety Alignment

Methods 이 논문이 사용한 방법론

Focus 설명 없음
LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…

Similar Papers 제목 키워드 기반

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

2025-07-22 · Zheng Hui, Yijiang River Dong, Ehsan Shareghi, Nigel Collier arxiv

As large language models (LLMs) are increasingly deployed in high-risk domains such as law, finance, and medicine, systematically evaluating their domain-specific safety and compliance becomes critical. While prior work …

Trilobite tridents: hydrodynamic lift and stability mechanisms for queue formation

2025-06-18 · Hugh A. Trenchard, Carlton E. Brett, Matjaz Perc

The bizarre trident-like cephalic projections of Walliserops trifurcatus have previously been interpreted as sexually selected weapons for intraspecific combat. We propose an alternative hypothesis grounded in biomechani…

TRIDENT: Breaking the Hybrid-Safety-Physics Coupling for Provably Safe Multi-Agent Reinforcement Learning

2026-06-16 · Zijie Meng, Ziwei Li, Yufei Liu, Zhiyu Li 외 arxiv

Safe coordination in networked cyber-physical systems forces learning algorithms to simultaneously handle hybrid discrete-continuous actions, hard training-time safety constraints, and physics-governed dynamics. We show …

Multi-agent Reinforcement Learning

TRIDENT: The Nonlinear Trilogy for Implicit Neural Representations

2023-11-21 · Zhenda Shen, Yanqi Cheng, Raymond H. Chan, Pietro Liò 외

Implicit neural representations (INRs) have garnered significant interest recently for their ability to model complex, high-dimensional data without explicit parameterisation. In this work, we introduce TRIDENT, a novel …

Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation

2024-11-14 · Yuheng Shi, Minjing Dong, Chang Xu

While Contrastive Language-Image Pre-training (CLIP) has advanced open-vocabulary predictions, its performance on semantic segmentation remains suboptimal. This shortfall primarily stems from its spatial-invariant semant…

SegmentationSemantic SegmentationUnsupervised Semantic Segmentation with Language-image Pre-training