paper-with-me

홈 › Papers

Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models

2024-10-07 · Fei Wang, Ninareh Mehrabi, Palash Goyal, Rahul Gupta, Kai-Wei Chang, Aram Galstyan

Data is a crucial element in large language model (LLM) alignment. Recent studies have explored using LLMs for efficient data collection. However, LLM-generated data often suffers from quality issues, with underrepresented or absent aspects and low-quality datapoints. To address these problems, we propose Data Advisor, an enhanced LLM-based method for generating data that takes into account the characteristics of the desired dataset. Starting from a set of pre-defined principles in hand, Data Advisor monitors the status of the generated data, identifies weaknesses in the current dataset, and advises the next iteration of data generation accordingly. Data Advisor can be easily integrated into existing data generation methods to enhance data quality and coverage. Experiments on safety alignment of three representative LLMs (i.e., Mistral, Llama2, and Falcon) demonstrate the effectiveness of Data Advisor in enhancing model safety against various fine-grained safety issues without sacrificing model utility.

📄 PDF Abstract BibTeX arXiv:2410.05269

Code (1)

feiwang96/Data-Advisor 공식 구현

Tasks

Language ModelingLanguage ModellingLarge Language ModelSafety Alignment

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

ROSA: Roundabout Optimized Speed Advisory with Multi-Agent Trajectory Prediction in Multimodal Traffic

2026-02-16 · Anna-Lena Schlamp, Jeremias Gerner, Klaus Bogenberger, Werner Huber 외 arxiv

We present ROSA -- Roundabout Optimized Speed Advisory -- a system that combines multi-agent trajectory prediction with coordinated speed guidance for multimodal, mixed traffic at roundabouts. Using a Transformer-based m…

Trajectory Prediction

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs

2026-04-08 · Yue Huang, Haomin Zhuang, Jiayi Ye, Han Bao 외 arxiv

Hard-gated safety checkers often over-refuse and misalign with a vendor's model spec; prevailing taxonomies also neglect robustness and honesty, yielding safer-on-paper yet less useful systems. This work introduces Guard…

Usage Governance Advisor: From Intent to AI Governance

2024-12-02 · Elizabeth M. Daly, Sean Rooney, Seshu Tirupathi, Luis Garces-Erice 외

Evaluating the safety of AI Systems is a pressing concern for organizations deploying them. In addition to the societal damage done by the lack of fairness of those systems, deployers are concerned about the legal reperc…

Fairness

Rule-based High-Level Coaching for Goal-Conditioned Reinforcement Learning in Search-and-Rescue UAV Missions Under Limited-Simulation Training

2026-04-29 · Mahya Ramezani, Holger Voos arxiv

This paper presents a hierarchical decision-making framework for unmanned aerial vehicle (UAV) missions motivated by search-and-rescue (SAR) scenarios under limited simulation training. The framework combines a fixed rul…

Reinforcement Learning

Pedestrian Behavior Maps for Safety Advisories: CHAMP Framework and Real-World Data Analysis

2023-05-08 · Ross Greer, Samveed Desai, Lulua Rakla, Akshay Gopalkrishnan 외

It is critical for vehicles to prevent any collisions with pedestrians. Current methods for pedestrian collision prevention focus on integrating visual pedestrian detectors with Automatic Emergency Braking (AEB) systems …

Pedestrian Detection