Exploring Human-LLM Conversations: Mental Models and the Originator of Toxicity
This study explores real-world human interactions with large language models (LLMs) in diverse, unconstrained settings in contrast to most prior research focusing on ethically trimmed models like ChatGPT for specific tasks. We aim to understand the originator of toxicity. Our findings show that although LLMs are rightfully accused of providing toxic content, it is mostly demanded or at least provoked by humans who actively seek such content. Our manual analysis of hundreds of conversations judged as toxic by APIs commercial vendors, also raises questions with respect to current practices of what user requests are refused to answer. Furthermore, we conjecture based on multiple empirical indicators that humans exhibit a change of their mental model, switching from the mindset of interacting with a machine more towards interacting with a human.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Is Your Toxicity My Toxicity? Exploring the Impact of Rater Identity on Toxicity Annotation
Machine learning models are commonly used to detect toxicity in online conversations. These models are trained on datasets annotated by human raters. We explore how raters' self-described identities impact how they annot…
Revisiting Contextual Toxicity Detection in Conversations
Understanding toxicity in user conversations is undoubtedly an important problem. Addressing "covert" or implicit cases of toxicity is particularly hard and requires context. Very few previous studies have analysed the i…
Data AugmentationToxic Comment ClassificationSay ‘YES’ to Positivity: Detecting Toxic Language in Workplace Communications
Workplace communication (e.g. email, chat, etc.) is a central part of enterprise productivity. Healthy conversations are crucial for creating an inclusive environment and maintaining harmony in an organization. Toxic com…
Exploring the Impact of Personality Traits on LLM Bias and Toxicity
With the different roles that AI is expected to play in human life, imbuing large language models (LLMs) with different personalities has attracted increasing research interests. While the "personification" enhances huma…
Text GenerationToxicity Detection: Does Context Really Matter?
Moderation is crucial to promoting healthy on-line discussions. Although several `toxicity' detection datasets and models have been published, most of them ignore the context of the posts, implicitly assuming that commen…