paper-with-me

홈 › Papers

MCFEND: A Multi-source Benchmark Dataset for Chinese Fake News Detection

2024-03-14 · Yupeng Li, Haorui He, Jin Bai, Dacheng Wen

The prevalence of fake news across various online sources has had a significant influence on the public. Existing Chinese fake news detection datasets are limited to news sourced solely from Weibo. However, fake news originating from multiple sources exhibits diversity in various aspects, including its content and social context. Methods trained on purely one single news source can hardly be applicable to real-world scenarios. Our pilot experiment demonstrates that the F1 score of the state-of-the-art method that learns from a large Chinese fake news detection dataset, Weibo-21, drops significantly from 0.943 to 0.470 when the test data is changed to multi-source news data, failing to identify more than one-third of the multi-source fake news. To address this limitation, we constructed the first multi-source benchmark dataset for Chinese fake news detection, termed MCFEND, which is composed of news we collected from diverse sources such as social platforms, messaging apps, and traditional online news outlets. Notably, such news has been fact-checked by 14 authoritative fact-checking agencies worldwide. In addition, various existing Chinese fake news detection methods are thoroughly evaluated on our proposed dataset in cross-source, multi-source, and unseen source ways. MCFEND, as a benchmark dataset, aims to advance Chinese fake news detection approaches in real-world scenarios.

📄 PDF Abstract BibTeX arXiv:2403.09092

Code (0)

등록된 구현이 없습니다.

Tasks

Fact CheckingFake News Detection

Similar Papers 제목 키워드 기반

C-Pack: Packed Resources For General Chinese Embeddings

2023-09-14 · Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff 외

We introduce C-Pack, a package of resources that significantly advance the field of general Chinese embeddings. C-Pack includes three critical resources. 1) C-MTEB is a comprehensive benchmark for Chinese text embeddings…

MTEB Benchmark

CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation

2024-07-01 · Yuxuan Wang, Yijun Liu, Fei Yu, Chen Huang 외

Despite the rapid development of Chinese vision-language models (VLMs), most existing Chinese vision-language (VL) datasets are constructed on Western-centric images from existing English VL datasets. The cultural bias i…

Image-text RetrievalQuestion AnsweringText RetrievalVisual Grounding+1

An Improved Traditional Chinese Evaluation Suite for Foundation Model

2024-03-04 · Zhi-Rui Tam, Ya-Ting Pai, Yen-Wei Lee, Jun-Da Chen 외

We present TMMLU+, a new benchmark designed for Traditional Chinese language understanding. TMMLU+ is a multi-choice question-answering dataset with 66 subjects from elementary to professional level. It is six times larg…

Multiple-choiceQuestion Answering

RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios

2025-09-22 · Fei Zhao, Chengqiang Lu, Yufan Shen, Qimeng Wang 외 arxiv

While various multimodal multi-image evaluation datasets have been emerged, but these datasets are primarily based on English, and there has yet to be a Chinese multi-image dataset. To fill this gap, we introduce RealBen…

CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation

2024-01-02 · Quan Tu, Shilong Fan, Zihang Tian, Rui Yan

Recently, the advent of large language models (LLMs) has revolutionized generative agents. Among them, Role-Playing Conversational Agents (RPCAs) attract considerable attention due to their ability to emotionally engage …