paper-with-me

홈 › Papers

A 106K Multi-Topic Multilingual Conversational User Dataset with Emoticons

2025-02-26 · Heng Er Metilda Chee, Jiayin Wang, Zhiqiang Guo, Weizhi Ma, Qinglang Guo, Min Zhang

Instant messaging has become a predominant form of communication, with texts and emoticons enabling users to express emotions and ideas efficiently. Emoticons, in particular, have gained significant traction as a medium for conveying sentiments and information, leading to the growing importance of emoticon retrieval and recommendation systems. However, one of the key challenges in this area has been the absence of datasets that capture both the temporal dynamics and user-specific interactions with emoticons, limiting the progress of personalized user modeling and recommendation approaches. To address this, we introduce the emoticon dataset, a comprehensive resource that includes time-based data along with anonymous user identifiers across different conversations. As the largest publicly accessible emoticon dataset to date, it comprises 22K unique users, 370K emoticons, and 8.3M messages. The data was collected from a widely-used messaging platform across 67 conversations and 720 hours of crawling. Strict privacy and safety checks were applied to ensure the integrity of both text and image data. Spanning across 10 distinct domains, the emoticon dataset provides rich insights into temporal, multilingual, and cross-domain behaviors, which were previously unavailable in other emoticon-based datasets. Our in-depth experiments, both quantitative and qualitative, demonstrate the dataset's potential in modeling user behavior and personalized recommendation systems, opening up new possibilities for research in personalized retrieval and conversational AI. The dataset is freely accessible.

📄 PDF Abstract BibTeX arXiv:2502.19108

Code (0)

등록된 구현이 없습니다.

Tasks

Recommendation SystemsRetrieval

Similar Papers 제목 키워드 기반

Investigating Thematic Patterns and User Preferences in LLM Interactions using BERTopic

2025-10-08 · Abhay Bhandarkar, Gaurav Mishra, Khushi Juchani, Harsh Singhal arxiv

This study applies BERTopic, a transformer-based topic modeling technique, to the lmsys-chat-1m dataset, a multilingual conversational corpus built from head-to-head evaluations of large language models (LLMs). Each user…

OleSpeech-IV: A Large-Scale Multispeaker and Multilingual Conversational Speech Dataset with Diverse Topics

2025-09-04 · Wei Chu, Yuanzhe Dong, Ke Tan, Dong Han 외 arxiv

OleSpeech-IV dataset is a large-scale multispeaker and multilingual conversational speech dataset with diverse topics. The audio content comes from publicly-available English podcasts, talk shows, teleconferences, and ot…

TruthBot: An Automated Conversational Tool for Intent Learning, Curated Information Presenting, and Fake News Alerting

2021-01-31 · Ankur Gupta, Yash Varun, Prarthana Das, Nithya Muttineni 외

We present TruthBot, an all-in-one multilingual conversational chatbot designed for seeking truth (trustworthy and verified information) on specific topics. It helps users to obtain information specific to certain topics…

Chatbot

Multilingual Datasets for Custom Input Extraction and Explanation Requests Parsing in Conversational XAI Systems

2025-08-20 · Qianli Wang, Tatiana Anikina, Nils Feldhus, Simon Ostermann 외 arxiv

Conversational explainable artificial intelligence (ConvXAI) systems based on large language models (LLMs) have garnered considerable attention for their ability to enhance user comprehension through dialogue-based expla…

Intent Recognition

Topic-based Evaluation for Conversational Bots

2018-01-11 · Fenfei Guo, Angeliki Metallinou, Chandra Khatri, Anirudh Raju 외

Dialog evaluation is a challenging problem, especially for non task-oriented dialogs where conversational success is not well-defined. We propose to evaluate dialog quality using topic-based metrics that describe the abi…

DiversityTopic Classification