paper-with-me

홈 › Papers

Automatically Exposing Problems with Neural Dialog Models

2021-09-14 · EMNLP 2021 11 · Dian Yu, Kenji Sagae

Neural dialog models are known to suffer from problems such as generating unsafe and inconsistent responses. Even though these problems are crucial and prevalent, they are mostly manually identified by model designers through interactions. Recently, some research instructs crowdworkers to goad the bots into triggering such problems. However, humans leverage superficial clues such as hate speech, while leaving systematic problems undercover. In this paper, we propose two methods including reinforcement learning to automatically trigger a dialog model into generating problematic responses. We show the effect of our methods in exposing safety and contradiction issues with state-of-the-art dialog models.

📄 PDF Abstract BibTeX arXiv:2109.06950

Code (1)

diandyu/trigger 공식 구현 pytorch

Tasks

reinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

On a Chatbot Conducting Dialogue-in-Dialogue

2019-09-01 · WS 2019 9 · Boris Galitsky, Dmitry Ilvovsky, Elizaveta Goncharova

We demo a chatbot that delivers content in the form of virtual dialogues automatically produced from plain texts extracted and selected from documents. This virtual dialogue content is provided in the form of answers der…

ChatbotForm

Observing Dialogue in Therapy: Categorizing and Forecasting Behavioral Codes

2019-06-30 · ACL 2019 7 · Jie Cao, Michael Tanana, Zac E. Imel, Eric Poitras 외

Automatically analyzing dialogue can help understand and guide behavior in domains such as counseling, where interactions are largely mediated by conversation. In this paper, we study modeling behavioral codes used to as…

On a Chatbot Providing Virtual Dialogues

2019-09-01 · RANLP 2019 9 · Boris Galitsky, Dmitry Ilvovsky, Elizaveta Goncharova

We present a chatbot that delivers content in the form of virtual dialogues automatically produced from the plain texts that are extracted and selected from the documents. This virtual dialogue content is provided in the…

ChatbotForm

Extracting relevant information from physician-patient dialogues for automated clinical note taking

2019-11-01 · WS 2019 11 · Serena Jeblee, Faiza Khan Khattak, Noah Crampton, Muhammad Mamdani 외

We present a system for automatically extracting pertinent medical information from dialogues between clinicians and patients. The system parses each dialogue and extracts entities such as medications and symptoms, using…

MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Models

2025-09-18 · Siyu Yan, Long Zeng, Xuecheng Wu, Chengcheng Han 외 arxiv

As large language models~(LLMs) become widely adopted, ensuring their alignment with human values is crucial to prevent jailbreaks where adversaries manipulate models to produce harmful content. While most defenses targe…

Red Teaming