Spotting tell-tale visual artifacts in face swapping videos: strengths and pitfalls of CNN detectors
Face swapping manipulations in video streams represents an increasing threat in remote video communications, due to advances in automated and real-time tools. Recent literature proposes to characterize and exploit visual artifacts introduced in video frames by swapping algorithms when dealing with challenging physical scenes, such as face occlusions. This paper investigates the effectiveness of this approach by benchmarking CNN-based data-driven models on two data corpora (including a newly collected one) and analyzing generalization capabilities with respect to different acquisition sources and swapping algorithms. The results confirm excellent performance of general-purpose CNN architectures when operating within the same data source, but a significant difficulty in robustly characterizing occlusion-based visual cues across datasets. This highlights the need for specialized detection strategies to deal with such artifacts.
Code (0)
등록된 구현이 없습니다.
Tasks
BenchmarkingFace SwappingSimilar Papers 제목 키워드 기반
Designing an AI-Driven Talent Intelligence Solution: Exploring Big Data to extend the TOE Framework
AI has the potential to improve approaches to talent management enabling dynamic provisions through implementing advanced automation. This study aims to identify the new requirements for developing AI-oriented artifacts …
ManagementTaleDiffusion: Multi-Character Story Generation with Dialogue Rendering
Text-to-story visualization is challenging due to the need for consistent interaction among multiple characters across frames. Existing methods struggle with character consistency, leading to artifact generation and inac…
Story VisualizationStory GenerationFlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration
LLM-based evolution has emerged as a promising way to improve agents by refining non-parametric artifacts, but its wall-clock cost remains a major bottleneck. We identify that this cost comes from synchronized stage exec…
Seeing wake words: Audio-visual Keyword Spotting
The goal of this work is to automatically determine whether and when a word of interest is spoken by a talking face, with or without the audio. We propose a zero-shot method suitable for in the wild videos. Our key contr…
Keyword SpottingLip ReadingVisual Keyword SpottingTowards Robust Visual Information Extraction in Real World: New Dataset and Novel Solution
Visual information extraction (VIE) has attracted considerable attention recently owing to its various advanced applications such as document understanding, automatic marking and intelligent education. Most existing work…
3D Feature Matchingdocument understandingText DetectionText Spotting