paper-with-me

Papers

When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity

2025-09-14 · Shiyao Cui, Xijia Feng, Yingkang Wang, Junxiao Yang, Zhexin Zhang, Biplab Sikdar, Hongning Wang, Han Qiu, Minlie Huang arxiv

Emojis are globally used non-verbal cues in digital communication, and extensive research has examined how large language models (LLMs) understand and utilize emojis across contexts. While usually associated with friendliness or playfulness, it is observed that emojis may trigger toxic content generation in LLMs. Motivated by such a observation, we aim to investigate: (1) whether emojis can clearly enhance the toxicity generation in LLMs and (2) how to interpret this phenomenon. We begin with a comprehensive exploration of emoji-triggered LLM toxicity generation by automating the construction of prompts with emojis to subtly express toxic intent. Experiments across 5 mainstream languages on 7 famous LLMs along with jailbreak tasks demonstrate that prompts with emojis could easily induce toxicity generation. To understand this phenomenon, we conduct model-level interpretations spanning semantic cognition, sequence generation and tokenization, suggesting that emojis can act as a heterogeneous semantic channel to bypass the safety mechanisms. To pursue deeper insights, we further probe the pre-training corpus and uncover potential correlation between the emoji-related data polution with the toxicity generation behaviors. Supplementary materials provide our implementation code and data. (Warning: This paper contains potentially sensitive contents)

📄 PDF Abstract BibTeX arXiv:2509.11141

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Quantitative Analysis of Comparison of Emoji Sentiment: Taiwan Mandarin Users and English Users

2022-11-01 · ROCLING 2022 11 · Fang-Yu Chang

Emojis have become essential components in our digital communication. Emojis, especially smiley face emojis and heart emojis, are considered the ones conveying more emotions. In this paper, two functions of emoji usages …

Language ModelingLanguage Modelling

SmileyNet -- Towards the Prediction of the Lottery by Reading Tea Leaves with AI

2024-07-31 · Andreas Birk

We introduce SmileyNet, a novel neural network with psychic abilities. It is inspired by the fact that a positive mood can lead to improved cognitive capabilities including classification tasks. The network is hence pres…

Task Adaptive Pretraining of Transformers for Hostility Detection

2021-01-09 · Tathagata Raha, Sayar Ghosh Roy, Ujwal Narayan, Zubair Abid 외

Identifying adverse and hostile content on the web and more particularly, on social media, has become a problem of paramount interest in recent years. With their ever increasing popularity, fine-tuning of pretrained Tran…

Binary ClassificationClassificationGeneral ClassificationMulti-Label Classification+1

Sentiment of Emojis

2015-09-25 · Petra Kralj Novak, Jasmina Smailović, Borut Sluban, Igor Mozetič

There is a new generation of emoticons, called emojis, that is increasingly being used in mobile communications and social media. In the past two years, over ten billion emojis were used on Twitter. Emojis are Unicode gr…

Sentiment Analysis

Investigating the Influence of Users Personality on the Ambiguous Emoji Perception

2022-07-01 · NAACL (Emoji) 2022 7 · Olga Iarygina

Emojis are an integral part of Internet communication nowadays. Even though, they are supposed to make the text clearer and less dubious, some emojis are ambiguous and can be interpreted in different ways. One of the fac…