paper-with-me

홈 › Papers

If there's a Trigger Warning, then where's the Trigger? Investigating Trigger Warnings at the Passage Level

2024-04-15 · Matti Wiegmann, Jennifer Rakete, Magdalena Wolska, Benno Stein, Martin Potthast

Trigger warnings are labels that preface documents with sensitive content if this content could be perceived as harmful by certain groups of readers. Since warnings about a document intuitively need to be shown before reading it, authors usually assign trigger warnings at the document level. What parts of their writing prompted them to assign a warning, however, remains unclear. We investigate for the first time the feasibility of identifying the triggering passages of a document, both manually and computationally. We create a dataset of 4,135 English passages, each annotated with one of eight common trigger warnings. In a large-scale evaluation, we then systematically evaluate the effectiveness of fine-tuned and few-shot classifiers, and their generalizability. We find that trigger annotation belongs to the group of subjective annotation tasks in NLP, and that automatic trigger classification remains challenging but feasible.

📄 PDF Abstract BibTeX arXiv:2404.09615

Code (1)

mattiwe/passage-level-trigger-warnings 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Trigger Warnings: Bootstrapping a Violence Detector for FanFiction

2022-09-09 · Magdalena Wolska, Christopher Schröder, Ole Borchardt, Benno Stein 외

We present the first dataset and evaluation results on a newly defined computational task of trigger warning assignment. Labeled corpus data has been compiled from narrative works hosted on Archive of Our Own (AO3), a we…

Binary Classification

TWeddit : A Dataset of Triggering Stories Predominantly Shared by Women on Reddit

2026-01-16 · Shirlene Rose Bandela, Sanjeev Parthasarathy, Vaibhav Garg arxiv

Warning: This paper may contain examples and topics that may be disturbing to some readers, especially survivors of miscarriage and sexual violence. People affected by abortion, miscarriage, or sexual violence often shar…

LSTM recurrent neural network assisted aircraft stall prediction for enhanced situational awareness

2020-12-09 · Tahsin Sejat Saniat, Tahiat Goni, Shaikat M. Galib

Since the dawn of mankind's introduction to powered flights, there have been multiple incidents which can be attributed to aircraft stalls. Most modern-day aircraft are equipped with advanced warning systems to warn the …

Deactivating Refusal Triggers: Understanding and Mitigating Overrefusal in Safety Alignment

2026-03-12 · Zhiyu Xue, Zimo Qi, Guangliang Liu, Bocheng Chen 외 arxiv

Safety alignment aims to ensure that large language models (LLMs) refuse harmful requests by post-training on harmful queries paired with refusal answers. Although safety alignment is widely adopted in industry, the over…

UNICORN: A Unified Backdoor Trigger Inversion Framework

2023-04-05 · Zhenting Wang, Kai Mei, Juan Zhai, Shiqing Ma

The backdoor attack, where the adversary uses inputs stamped with triggers (e.g., a patch) to activate pre-planted malicious behaviors, is a severe threat to Deep Neural Network (DNN) models. Trigger inversion is an effe…

Backdoor Attack