paper-with-me

Papers

MCA-Bench: A Multimodal Benchmark for Evaluating CAPTCHA Robustness Against VLM-based Attacks

2025-06-06 · Zonglin Wu, Yule Xue, Xin Wei, Yiren Song

As automated attack techniques rapidly advance, CAPTCHAs remain a critical defense mechanism against malicious bots. However, existing CAPTCHA schemes encompass a diverse range of modalities -- from static distorted text and obfuscated images to interactive clicks, sliding puzzles, and logic-based questions -- yet the community still lacks a unified, large-scale, multimodal benchmark to rigorously evaluate their security robustness. To address this gap, we introduce MCA-Bench, a comprehensive and reproducible benchmarking suite that integrates heterogeneous CAPTCHA types into a single evaluation protocol. Leveraging a shared vision-language model backbone, we fine-tune specialized cracking agents for each CAPTCHA category, enabling consistent, cross-modal assessments. Extensive experiments reveal that MCA-Bench effectively maps the vulnerability spectrum of modern CAPTCHA designs under varied attack settings, and crucially offers the first quantitative analysis of how challenge complexity, interaction depth, and model solvability interrelate. Based on these findings, we propose three actionable design principles and identify key open challenges, laying the groundwork for systematic CAPTCHA hardening, fair benchmarking, and broader community collaboration. Datasets and code are available online.

📄 PDF Abstract BibTeX arXiv:2506.05982

Code (2)

noheadwuzonglin/mca-bench 공식 구현 pytorch
MindSpore-scientific-2/code-12/tree/main/MCA mindspore

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents

2025-05-30 · Yaxin Luo, Zhaoyi Li, Jiacheng Liu, Jiacheng Cui 외

CAPTCHAs have been a critical bottleneck for deploying web agents in real-world applications, often blocking them from completing end-to-end automation tasks. While modern multimodal LLM agents have demonstrated impressi…

BenchmarkingBlockingMultimodal ReasoningVisual Reasoning

Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent Defense

2026-02-09 · Jiacheng Liu, Yaxin Luo, Jiacheng Cui, Xinyi Shang 외 arxiv

The rapid evolution of GUI-enabled agents has rendered traditional CAPTCHAs obsolete. While previous benchmarks like OpenCaptchaWorld established a baseline for evaluating multimodal agents, recent advancements in reason…

Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine Differentiation

2025-10-04 · Arina Kharlamova, Bowei He, Chen Ma, Xue Liu arxiv

Online services rely on CAPTCHAs as a first line of defense against automated abuse, yet recent advances in multi-modal large language models (MLLMs) have eroded the effectiveness of conventional designs that focus on te…

Spatial Reasoning

CAPTURE: A Benchmark and Evaluation for LVLMs in CAPTCHA Resolving

2025-12-12 · Jianyi Zhang, Ziyin Zhou, Xu Ji, Shizhao Liu 외 arxiv

Benefiting from strong and efficient multi-modal alignment strategies, Large Visual Language Models (LVLMs) are able to simulate human visual and reasoning capabilities, such as solving CAPTCHAs. However, existing benchm…

Breaking reCAPTCHAv2

2024-09-13 · Andreas Plesner, Tobias Vontobel, Roger Wattenhofer

Our work examines the efficacy of employing advanced machine learning methods to solve captchas from Google's reCAPTCHAv2 system. We evaluate the effectiveness of automated systems in solving captchas by utilizing advanc…

Image SegmentationSemantic Segmentation