paper-with-me

홈 › Papers

SynCred-Bench: Benchmarking Synthetic Credibility in AI-Generated Visual Misinformation

2026-06-02 · Junxiao Yang, Minghao Zhang, Xiaoce Wang, Haoran Liu, Shiyao Cui, Hongning Wang, Minlie Huang arxiv

Recent generative models can now produce visual artifacts with realistic embedded text and layouts, creating a new misinformation threat: synthetic credibility. We introduce SYNCRED-Bench, a benchmark of 600 AI-generated misinformation images balanced across six credible-form categories and seven fine-grained circulation styles, together with FP450, a real-image negative set for measuring false positives. Extensive evaluation shows that existing systems remain unreliable: under a 5% false-positive-rate constraint, 15 MLLMs achieve only 10.5% true positive rate (TPR), open-source AIGC detectors achieve less than 5%, and commercial APIs reach 57.6%. Human annotators also struggled to identify synthetic credibility, reaching only 63% TPR. These findings establish synthetic credibility as a severe and underexplored visual misinformation challenge, and provide a benchmark for developing detectors that reason beyond superficial credibility cues.

📄 PDF Abstract BibTeX arXiv:2606.03348

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On the use of automatically generated synthetic image datasets for benchmarking face recognition

2021-06-08 · Laurent Colbois, Tiago de Freitas Pereira, Sébastien Marcel

The availability of large-scale face datasets has been key in the progress of face recognition. However, due to licensing issues or copyright infringement, some datasets are not available anymore (e.g. MS-Celeb-1M). Rece…

BenchmarkingFace Recognition

Is Seeing Believing? Evaluating Human Sensitivity to Synthetic Video

2026-03-14 · David Wegmann, Emil Stevnsborg, Søren Knudsen, Luca Rossi 외 arxiv

Advances in machine learning have enabled the creation of realistic synthetic videos known as deepfakes. As deepfakes proliferate, concerns about rapid spread of disinformation and manipulation of public perception are m…

Deep FinResearch Bench: Evaluating AI's Ability to Conduct Professional Financial Investment Research

2026-04-22 · Mirazul Haque, Antony Papadimitriou, Samuel Mensah, Zhiqiang Ma 외 arxiv

We introduce Deep FinResearch Bench, a practical and comprehensive evaluation framework for deep research (DR) agents in financial investment research. The benchmark assesses three dimensions of report quality: qualitati…

Bridging the Gap Between Ideal and Real-world Evaluation: Benchmarking AI-Generated Image Detection in Challenging Scenarios

2025-09-11 · Chunxiao Li, Xiaoxiao Wang, Meiling Li, Boming Miao 외 arxiv

With the rapid advancement of generative models, highly realistic image synthesis has posed new challenges to digital security and media credibility. Although AI-generated image detection methods have partially addressed…

Few-Shot Learning

RecAD: Towards A Unified Library for Recommender Attack and Defense

2023-09-09 · Changsheng Wang, Jianbai Ye, Wenjie Wang, Chongming Gao 외

In recent years, recommender systems have become a ubiquitous part of our daily lives, while they suffer from a high risk of being attacked due to the growing commercial and social values. Despite significant research pr…

BenchmarkingRecommendation Systems