paper-with-me

Papers

Suppressing Pink Elephants with Direct Principle Feedback

2024-02-12 · Louis Castricato, Nathan Lile, Suraj Anand, Hailey Schoelkopf, Siddharth Verma, Stella Biderman

Existing methods for controlling language models, such as RLHF and Constitutional AI, involve determining which LLM behaviors are desirable and training them into a language model. However, in many cases, it is desirable for LLMs to be controllable at inference time, so that they can be used in multiple contexts with diverse needs. We illustrate this with the Pink Elephant Problem: instructing an LLM to avoid discussing a certain entity (a `Pink Elephant''), and instead discuss a preferred entity (`Grey Elephant''). We apply a novel simplification of Constitutional AI, Direct Principle Feedback, which skips the ranking of responses and uses DPO directly on critiques and revisions. Our results show that after DPF fine-tuning on our synthetic Pink Elephants dataset, our 13B fine-tuned LLaMA 2 model significantly outperforms Llama-2-13B-Chat and a prompted baseline, and performs as well as GPT-4 in on our curated test set assessing the Pink Elephant Problem.

📄 PDF Abstract BibTeX arXiv:2402.07896

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

DPO 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Position-Wise Feed-Forward Layer 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…

Similar Papers 제목 키워드 기반

ElephantBook: A Semi-Automated Human-in-the-Loop System for Elephant Re-Identification

2021-06-29 · Peter Kulits, Jake Wall, Anka Bedetti, Michelle Henley 외

African elephants are vital to their ecosystems, but their populations are threatened by a rise in human-elephant conflict and poaching. Monitoring population dynamics is essential in conservation efforts; however, track…

Attribute

A simple model for pink noise from amplitude modulations

2023-01-26 · Masahiro Morikawa, Akika Nakamichi

We propose a simple model for the origin of pink noise (or 1/f fluctuation) based on the beat of cooperative waves. These cooperative waves arise spontaneously in a system with synchronization, resonance, and infrared di…

An Analysis of Elephants' Movement Data in Sub-Saharan Africa Using Clustering

2021-11-05 · Gregory Glatzer, Prasenjit Mitra, Johnson Kinyua

Understanding the movement of animals is crucial to conservation efforts. Past research often focuses on factors affecting movement, rather than locations of interest that animals return to or habitat. We explore the use…

Clustering

The impact of physiological stress conditions on protein structure and trypsin inhibition of serine protease inhibitor Kazal type 1 (SPINK1) and its N34S variant

2020-12-18 · Ina Buchholz, Felix Nagel, Annelie Klein, Preshit R. Wagh 외

One of the most common mutations in the serine protease inhibitor Kazal type 1 (SPINK1) gene is the N34S variant which is strongly associated with chronic pancreatitis. Although it is assumed that N34S mutation constitut…

PInKS: Preconditioned Commonsense Inference with Minimal Supervision

2022-06-16 · Ehsan Qasemi, Piyush Khanna, Qiang Ning, Muhao Chen

Reasoning with preconditions such as "glass can be used for drinking water unless the glass is shattered" remains an open problem for language models. The main challenge lies in the scarcity of preconditions data and the…

Informativeness