TextSleuth: A New Dataset and Baseline for Scene Text Manipulation Detection
With the rise of digital content on social media and the advancement of image editing tools, tampering with scene text has become a serious concern. Scene text manipulation detection (STMD) is a kind of image manipulation detection (IMD) with focus on the tampering of scene text pixels, which is crucial for image content integrity and media forensics. In this paper, we present TextSleuth, a novel benchmark dataset specifically designed for STMD, by integrating three public datasets with newly introduced manipulation and annotations. We introduce professional edits on the Total-Text dataset (~1K images) with four levels of manipulated region perceptibility, and a large synthetic manipulation set (858K images) on the SynthText dataset, as well the integration of the Tampered-IC13 dataset (378 images). We established a new STMD baseline based on TextSleuth using MMFusion-IML, the state-of-the-art image manipulation detection model. We performed extensive experiments, reporting the AUC from ROC analysis and the balanced accuracy (bACC) metrics to maintain a balanced performance evaluation. The MMFusion-IML baseline achieves 0.641 AUC and 0.588 bACC on the Total-Text subset. In comparison, it achieves 0.89 AUC and 0.8272 bACC on the Tampered-IC13 subset. This showcases the real-world STMD challenges reflected in our new dataset. TextSleuth is a valuable resource for future research in scene text manipulation detection and forensics. The dataset is available at https://github.com/abhineet-pandey/Text-Sleuth.
Code (1)
Tasks
Image ManipulationImage Manipulation DetectionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
TextSleuth: Towards Explainable Tampered Text Detection
Recently, tampered text detection has attracted increasing attention due to its essential role in information security. Although existing methods can detect the tampered text region, the interpretation of such detection …
Domain GeneralizationOptical Character Recognition (OCR)Text DetectionComprehensive Visual Question Answering on Point Clouds through Compositional Scene Manipulation
Visual Question Answering on 3D Point Cloud (VQA-3D) is an emerging yet challenging field that aims at answering various types of textual questions given an entire point cloud scene. To tackle this problem, we propose th…
Common Sense ReasoningQuestion AnsweringScene UnderstandingVisual Question Answering+2Image Manipulation via Multi-Hop Instructions -- A New Dataset and Weakly-Supervised Neuro-Symbolic Approach
We are interested in image manipulation via natural language text -- a task that is useful for multiple AI applications but requires complex reasoning over multi-modal spaces. We extend recently proposed Neuro Symbolic C…
Image ManipulationQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)Splat-MOVER: Multi-Stage, Open-Vocabulary Robotic Manipulation via Editable Gaussian Splatting
We present Splat-MOVER, a modular robotics stack for open-vocabulary robotic manipulation, which leverages the editability of Gaussian Splatting (GSplat) scene representations to enable multi-stage manipulation tasks. Sp…
Grasp GenerationSimulated Gaussian ManipulationGenerating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
Learning to use tools or objects in common scenes, particularly handling them in various ways as instructed, is a key challenge for developing interactive robots. Training models to generate such manipulation traject…
valid