paper-with-me

홈 › Papers

Clustering Vietnamese Conversations From Facebook Page To Build Training Dataset For Chatbot

2021-12-31 · Trieu Hai Nguyen, Thi-Kim-Ngoan Pham, Thi-Hong-Minh Bui, Thanh-Quynh-Chau Nguyen

The biggest challenge of building chatbots is training data. The required data must be realistic and large enough to train chatbots. We create a tool to get actual training data from Facebook messenger of a Facebook page. After text preprocessing steps, the newly obtained dataset generates FVnC and Sample dataset. We use the Retraining of BERT for Vietnamese (PhoBERT) to extract features of our text data. K-Means and DBSCAN clustering algorithms are used for clustering tasks based on output embeddings from PhoBERT$_{base}$. We apply V-measure score and Silhouette score to evaluate the performance of clustering algorithms. We also demonstrate the efficiency of PhoBERT compared to other models in feature extraction on the Sample dataset and wiki dataset. A GridSearch algorithm that combines both clustering evaluations is also proposed to find optimal parameters. Thanks to clustering such a number of conversations, we save a lot of time and effort to build data and storylines for training chatbot.

📄 PDF Abstract BibTeX arXiv:2112.15338

Code (1)

trieuntu/conversation_clustering 공식 구현 pytorch

Tasks

ChatbotClustering

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Residual Connection 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
WordPiece 설명 없음

Similar Papers 제목 키워드 기반

Dialogue Act Segmentation for Vietnamese Human-Human Conversational Texts

2017-08-16 · Thi Lan Ngo, Khac Linh Pham, Minh Son Cao, Son Bao Pham 외

Dialog act identification plays an important role in understanding conversations. It has been widely applied in many fields such as dialogue systems, automatic machine translation, automatic speech recognition, and espec…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)BIG-bench Machine LearningDeep Learning+4

HSD Shared Task in VLSP Campaign 2019:Hate Speech Detection for Social Good

2020-07-13 · Xuan-Son Vu, Thanh Vu, Mai-Vu Tran, Thanh Le-Cong 외

The paper describes the organisation of the "HateSpeech Detection" (HSD) task at the VLSP workshop 2019 on detecting the fine-grained presence of hate speech in Vietnamese textual items (i.e., messages) extracted from Fa…

General ClassificationHate Speech DetectionMulti-class Classificationvalid

VAIS Hate Speech Detection System: A Deep Learning based Approach for System Combination

2019-10-12 · Thai Binh Nguyen, Quang Minh Nguyen, Thu Hien Nguyen, Ngoc Phuong Pham 외

Nowadays, Social network sites (SNSs) such as Facebook, Twitter are common places where people show their opinions, sentiments and share information with others. However, some people use SNSs to post abuse and harassment…

Hate Speech Detection

Prioritizing Original News on Facebook

2021-02-16 · Xiuyan Ni, Shujian Bu, Igor L. Markov

This work outlines how we prioritize original news, a critical indicator of news quality. By examining the landscape and life-cycle of news posts on our social media platform, we identify challenges of building and deplo…

Clustering

VSoLSCSum: Building a Vietnamese Sentence-Comment Dataset for Social Context Summarization

2016-12-01 · WS 2016 12 · Minh-Tien Nguyen, Dac Viet Lai, Phong-Khac Do, Duc-Vu Tran 외

This paper presents VSoLSCSum, a Vietnamese linked sentence-comment dataset, which was manually created to treat the lack of standard corpora for social context summarization in Vietnamese. The dataset was collected thro…

Learning-To-RankSentence