paper-with-me

Papers

MegaChat: A Synthetic Persian Q&A Dataset for High-Quality Sales Chatbot Evaluation

2025-11-28 · Mahdi Rahmani, AmirHossein Saffari, Reyhane Rahmani arxiv

Small and medium-sized enterprises (SMEs) in Iran increasingly leverage Telegram for sales, where real-time engagement is essential for conversion. However, developing AI-driven chatbots for this purpose requires large, high-quality question-and-answer (Q&A) datasets, which are typically expensive and resource-intensive to produce, especially for low-resource languages like Persian. In this paper, we introduce MegaChat, the first fully synthetic Persian Q&A dataset designed to evaluate intelligent sales chatbots in Telegram-based e-commerce. We propose a novel, automated multi-agent architecture that generates persona-aware Q&A pairs by collecting data from active Telegram shopping channels. The system employs specialized agents for question generation, validation, and refinement, ensuring the production of realistic and diverse conversational data. To evaluate answer generation, we compare three classic retrieval-augmented generation (RAG) models with our advanced agentic system, which features multi-query retrieval, reranking, and persona-aligned response synthesis. Using GPT-5.1 for evaluation across six quality dimensions, our results show that the agentic architecture outperformed traditional RAG models in 4 out of 5 diverse channels, demonstrating its ability to generate scalable, high-quality datasets without relying on expensive human annotation or complex fine-tuning. MegaChat provides SMEs with an efficient, cost-effective solution for building intelligent customer engagement systems in specialized commercial domains, enabling advancements in multilingual conversational AI for low-resource languages. Download: https://github.com/MegaChat-Tech/MegaChat-DataSet

📄 PDF Abstract BibTeX arXiv:2511.23397

Code (0)

등록된 구현이 없습니다.

Tasks

Question GenerationAnswer Generation

Similar Papers 제목 키워드 기반

Persian Pixel: A large-scale synthetic OCR dataset for Persian language

2026-07-22 · Pouria Mahdi, Haq Nawaz Malik arxiv

Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages despite Persian being spoken by more than 110 million people across multiple countries. This gap arises fr…

Synthetic Data Generation

Matina: A Large-Scale 73B Token Persian Text Corpus

2025-02-13 · Sara Bourbour Hosseinbeigi, Fatemeh Taherinezhad, Heshaam Faili, Hamed Baghbani 외

Text corpora are essential for training models used in tasks like summarization, translation, and large language models (LLMs). While various efforts have been made to collect monolingual and multilingual datasets in man…

ArticlesDiversity

Improving Direct Persian-English Speech-to-Speech Translation with Discrete Units and Synthetic Parallel Data

2025-11-16 · Sina Rashidi, Hossein Sameti arxiv

Direct speech-to-speech translation (S2ST), in which all components are trained jointly, is an attractive alternative to cascaded systems because it offers a simpler pipeline and lower inference latency. However, direct …

Speech-to-Speech Translation

ELAB: Extensive LLM Alignment Benchmark in Persian Language

2025-04-17 · Zahra Pourbahman, Fatemeh Rajabi, Mohammadhossein Sadeghi, Omid Ghahroodi 외

This paper presents a comprehensive evaluation framework for aligning Persian Large Language Models (LLMs) with critical ethical dimensions, including safety, fairness, and social norms. It addresses the gaps in existing…

FairnessRed Teaming

ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis

2025-10-12 · Mohammad Javad Ranjbar Kalahroodi, Heshaam Faili, Azadeh Shakery arxiv

Persian remains substantially underrepresented in open speech-text resources, limiting progress in multi-speaker text-to-speech (TTS), speech-language modelling, and low-resource speech processing. We introduce ParsVoice…

Text-To-Speech SynthesisSpeaker IdentificationLanguage Modelling